Introducing Fikr VLA-1

Robots that see, understand, and act

Fikr AI is a new AI company building vision-language-action models that give real-world robots general-purpose intelligence - from perception to dexterous manipulation.

> pick up the cube > planning trajectory...
Fikr VLA-1 executing a natural-language instruction
The model

One foundation model, from pixels to motor commands

Fikr VLA-1 is a vision-language-action foundation model that runs directly on the robot. It fuses camera streams and natural-language instructions into a single representation, then decodes smooth, low-latency action trajectories - no task-specific programming required.

50 Hz
Closed-loop control frequency on embedded hardware
150+
Distinct manipulation tasks from a single checkpoint
<60 ms
End-to-end perception-to-action latency
3
Robot embodiments supported: arms, humanoids, mobile bases
Capabilities

Built for the messiness of the real world

Trained on large-scale robot demonstrations and web-scale vision-language data, Fikr VLA-1 generalizes to objects, scenes, and instructions it has never encountered.

Open-world perception

Recognizes and localizes everyday objects in cluttered scenes, tracking them through occlusion and changing lighting.

Language-guided tasks

Follows free-form spoken or written instructions, decomposing multi-step requests into executable subtasks.

Dexterous manipulation

Performs fine-grained, bi-manual tasks - folding, zipping, pouring, insertion - with human-like smoothness.

Rapid adaptation

Adapts to new tasks and embodiments with 50-100 demonstrations, fine-tuning in hours instead of months.

How it works

Vision. Language. Action.

A single end-to-end architecture replaces the traditional perception-planning-control stack.

01 - VISION

See the scene

Multi-view camera streams are encoded into rich spatial tokens that capture objects, geometry, and affordances in real time.

02 - LANGUAGE

Understand the goal

A language backbone grounds the instruction in the scene, reasoning about intent, constraints, and task order.

03 - ACTION

Act with precision

A diffusion-based action decoder outputs continuous joint trajectories at 50 Hz, closing the loop on every frame.

Benchmarks

State-of-the-art task success

Average success rate across our internal evaluation suite of 150 real-world manipulation tasks, including unseen objects and novel instructions.

Fikr VLA-1 on-robot, single checkpoint
87%
Fikr VLA-1 Lite edge-optimized variant
79%
Best open-source VLA baseline
61%
Task-specific behavior cloning
48%

Internal evaluation, June 2026. Success averaged over 10 rollouts per task on physical hardware.

Specifications

Model details

Model family
Fikr VLA-1 and Fikr VLA-1 Lite
Inputs
RGB camera streams (up to 4 views), proprioception, natural-language instructions
Outputs
Continuous action trajectories (joint positions or end-effector poses)
Control frequency
Up to 50 Hz closed-loop
Deployment
On-device (embedded GPU) or cloud-assisted hybrid
Embodiments
Single and dual robot arms, humanoids, mobile manipulators
Fine-tuning
Adapts to new tasks with 50-100 teleoperated demonstrations

Bring intelligence to your robots

We are partnering with a small group of robotics teams for early access to Fikr VLA-1. Tell us about your platform and use case.

Contact us