How can we adapt generalist policies at deployment time to go beyond their pre-training?
Vision-language-action (VLA) models learn broad skill repertoires from pretraining, giving us powerful priors over plausible robot behaviors. The standard strategy is to just prompt these models with new, desired task goals. However, as tasks scale in complexity and horizon, a successful policy must implicitly decompose goals into atomic, executable behaviors, and ground each behavior in a skill the robot can actually pull off in its current context. Often, directly prompting a VLA with a challenging goal fails.
Ask a VLA to move the hammer to the plate then grasp the mushroom and its action distribution collapses into the wrong modes — it moves the mushroom to the plate and grabs the spoon instead.
A VLM can break the goal into sensible sub-instructions, but has no idea which commands actually work on this robot. Saying pick up the hammer results in the VLA grabbing the spoon. Plausible ≠ grounded.
The only way to know which prompts actually make progress is to try them on the robot and learn from the result. Our method runs RL over the prompt space, learning a grounded mapping between "semantic" actions and physical behaviors, achieving both decomposition and grounding.
SARL runs RL over a VLA's language prompts, lifting learning from robot actions to the semantic level.
Rather than viewing a VLA as a policy to be statically prompted, we view it as a semantically controllable action prior that can be dynamically guided throughout deployment. This motivates a simple but powerful transformation of the RL problem: instead of learning over the robot action space \(\mathcal{A}_{\mathrm{robot}}\) (joint positions, end-effector deltas, etc.), we learn over a semantic action space \(\mathcal{A}_{\mathrm{sem}}\) — the space of language commands — and deploy the VLA as a transformation between the two. We call the resulting problem a semantic MDP.
Concretely, at each step SARL picks a semantic action (a prompt) \(\ell\), the VLA turns it into robot actions, and the environment transitions and returns a reward. SARL learns a semantic Q-function \(Q_{\mathrm{sem}}(s, \ell)\) via temporal-difference backups that estimates how effective each prompt is at making task progress, and acts by sampling from the softmax over these Q-values. Finally, naively searching over all possible prompts is intractable, so SARL uses a VLM to propose a small set of candidate semantic actions from the current observation.
Task suite. We evaluate SARL on complex, long-horizon WidowX tasks (above), plus ten tasks on the simulated LIBERO-LONG benchmark. These tasks are multi-step and require composing skills seen during pretraining (Bridge V2 for real, LIBERO-90 for sim) — the base policies achieve near-0% success on most of them, making them a clean test of deployment-time adaptation.
SARL improves the base VLA's success rate from near 0% under the task prompt up to 60–80% after only 60–120 online episodes, in both the real world and simulation. It beats action-space RL methods (DSRL, Residual RL) that are fundamentally limited by the base policy's behavior under a single fixed prompt, and it beats an in-context-learning VLM baseline that proposes plausible commands but struggles to ground them in physical behavior.
Real-world WidowX. Across all four tasks, SARL achieves the best improvement of generalist-policy behavior in deployment. Each data point represents 10 evaluations.
Simulated LIBERO-LONG. On long-horizon LIBERO-LONG tasks, SARL outperforms DSRL, Residual RL, and the ICL VLM. It successfully adapts the policy on four tasks and matches performance on another that is already near-solved. Each point represents 64 evaluations with standard error over 3 seeds.
Is SARL simply discovering a better task prompt? No. We evaluated the base policy under a wide range of task prompts, and on every long-horizon task no single prompt achieves any success — including the exact prompts SARL optimizes over. SARL succeeds by modulating the prompt within an episode, decomposing the goal into achievable stages.
| Real-world Task | VLM Best-of-10 a | Prompt Cache b |
|---|---|---|
| Task 1 (long-horizon) | 0% | 0 / 100 |
| Task 2 (long-horizon) | 0% | 0 / 100 |
| Task 3 (long-horizon) | 0% | 0 / 100 |
| Task 4 (short-horizon) | 50% | 26 / 100 |
a VLA success rate under the best of 10 task-prompt rephrasings generated by a VLM.
b Number of successful VLA rollouts under task prompts sampled from the set of prompts that SARL optimizes over.
Base policy on diverse task prompts. On all long-horizon tasks, no prompt alone achieves any success.
SARL vs. zero-shot prompting. The zero-shot base policy fails on long-horizon tasks; SARL solves them through learned prompting, sequencing skills covered under the pretraining distribution.
Like SARL, an in-context-learning VLM can decompose a complex goal into sub-instructions — but its commands are not grounded in what the VLA can actually execute. Over many episodes of experience, SARL learns the behavior each prompt induces, and this grounding is critical for performant adaptation.
SARL vs. VLM prompting. The VLM picks commands that are semantically plausible but ungrounded: it references an out-of-distribution “hammer” by name (making the VLA grab a spoon) or picks the wrong verb (drop), prematurely releasing the sushi. SARL learns to select the prompts that actually work.
On long-horizon tasks where no fixed prompt yields any success (Insight 1), action-space RL methods (DSRL, Residual RL) cannot bootstrap at all. In contrast, SARL bootstraps by stitching together behaviors induced by different commands.
To further show that SARL's performance gains come from carefully modulating the VLA's prompt, we run an additional Residual Policy experiment where, instead of using a fixed task prompt, the ICL VLM changes the prompt within each episode.
VLM-prompted Residual RL. Even when the Residual policy's prompt is modulated within each episode by the ICL VLM, it fails to match SARL on our hardest tasks. This baseline succeeds only when the VLM happens to pick task-solving prompts—it cannot improve the prompt-selection strategy itself.
On short-horizon tasks (for example, Task 4), action-space RL remains powerful. We therefore view the shape of exploration induced by each of these methods—SARL and action-space RL—as complementary. We are excited by this finding, since it suggests that semantic exploration is a new lever the community can pull to design algorithms for efficient adaptation.
Watch SARL learn each task online, sped up 30x.
@misc{bhatia2026sarl,
author = {Bhatia, Jagdeep Singh and Wagenmaker, Andrew and Chen, William and Levine, Sergey},
title = {Adapting Generalist Robot Policies with Semantic Reinforcement Learning},
year = {2026},
}