Fikr AI is a new AI company building vision-language-action models that give real-world robots general-purpose intelligence - from perception to dexterous manipulation.
Fikr VLA-1 is a vision-language-action foundation model that runs directly on the robot. It fuses camera streams and natural-language instructions into a single representation, then decodes smooth, low-latency action trajectories - no task-specific programming required.
Trained on large-scale robot demonstrations and web-scale vision-language data, Fikr VLA-1 generalizes to objects, scenes, and instructions it has never encountered.
Recognizes and localizes everyday objects in cluttered scenes, tracking them through occlusion and changing lighting.
Follows free-form spoken or written instructions, decomposing multi-step requests into executable subtasks.
Performs fine-grained, bi-manual tasks - folding, zipping, pouring, insertion - with human-like smoothness.
Adapts to new tasks and embodiments with 50-100 demonstrations, fine-tuning in hours instead of months.
A single end-to-end architecture replaces the traditional perception-planning-control stack.
Multi-view camera streams are encoded into rich spatial tokens that capture objects, geometry, and affordances in real time.
A language backbone grounds the instruction in the scene, reasoning about intent, constraints, and task order.
A diffusion-based action decoder outputs continuous joint trajectories at 50 Hz, closing the loop on every frame.
Average success rate across our internal evaluation suite of 150 real-world manipulation tasks, including unseen objects and novel instructions.
Internal evaluation, June 2026. Success averaged over 10 rollouts per task on physical hardware.
We are partnering with a small group of robotics teams for early access to Fikr VLA-1. Tell us about your platform and use case.
Contact us