Vision-Language-Action (VLA) / Robotics Foundation Models Jobs

Filter

My recent searches
Filter by:
Budget
to
to
to
Type
Skills
Languages
    Job State
    1 jobs found

    I am expanding an existing embodied-AI stack and now want to double-down on Vision-Language-Action modelling. The core goal is to design, train and scale a full pipeline that takes raw multimodal data, learns a joint representation and closes the loop all the way to real-time action on a physical robot. You will start from large, messy datasets (images, video clips, proprioception, language annotations) that already sit on our cluster. The job is to craft a new architecture in PyTorch, schedule distributed training, and iterate until the model achieves reliable closed-loop visuomotor reasoning in simulation and on hardware. Robust multimodal representation learning, a world-model/JEPA component, temporal memory, and predictive control need to come together in a single, maintainable codeba...

    $168 Average bid
    $168 Avg Bid
    61 bids

    Recommended Articles Just for You