Adapting Generalist Robot Policies with
Semantic Reinforcement Learning

University of California, Berkeley
SARL teaser: RL over semantic actions (language prompts).

TL;DR We propose optimizing the prompt inputs of generalist robot policies with reinforcement learning, enabling efficient real-robot adaptation on complex & long-horizon tasks, where existing methods for improving robot behavior in deployment struggle.

The Challenge

How can we adapt generalist policies at deployment time to go beyond their pre-training?

Vision-language-action (VLA) models learn broad skill repertoires from pretraining, giving us powerful priors over plausible robot behaviors. The standard strategy is to just prompt these models with new, desired task goals. However, as tasks scale in complexity and horizon, a successful policy must implicitly decompose goals into atomic, executable behaviors, and ground each behavior in a skill the robot can actually pull off in its current context. Often, directly prompting a VLA with a challenging goal fails.

Why prompting fails

Long-horizon prompts are out of support

Ask a VLA to move the hammer to the plate then grasp the mushroom and its action distribution collapses into the wrong modes — it moves the mushroom to the plate and grabs the spoon instead.

Why a VLM isn't enough

VLMs decompose, but don't ground

A VLM can break the goal into sensible sub-instructions, but has no idea which commands actually work on this robot. Saying pick up the hammer results in the VLA grabbing the spoon. Plausible ≠ grounded.

Why we need RL

RL grounds language in behavior

The only way to know which prompts actually make progress is to try them on the robot and learn from the result. Our method runs RL over the prompt space, learning a grounded mapping between "semantic" actions and physical behaviors, achieving both decomposition and grounding.

Semantic Action RL (SARL)

SARL runs RL over a VLA's language prompts, lifting learning from robot actions to the semantic level.

Rather than viewing a VLA as a policy to be statically prompted, we view it as a semantically controllable action prior that can be dynamically guided throughout deployment. This motivates a simple but powerful transformation of the RL problem: instead of learning over the robot action space \(\mathcal{A}_{\mathrm{robot}}\) (joint positions, end-effector deltas, etc.), we learn over a semantic action space \(\mathcal{A}_{\mathrm{sem}}\) — the space of language commands — and deploy the VLA as a transformation between the two. We call the resulting problem a semantic MDP.

Concretely, at each step SARL picks a semantic action (a prompt) \(\ell\), the VLA turns it into robot actions, and the environment transitions and returns a reward. SARL learns a semantic Q-function \(Q_{\mathrm{sem}}(s, \ell)\) via temporal-difference backups that estimates how effective each prompt is at making task progress, and acts by sampling from the softmax over these Q-values. Finally, naively searching over all possible prompts is intractable, so SARL uses a VLM to propose a small set of candidate semantic actions from the current observation.

Algorithm 1  ·  Semantic Action RL (SARL)
  1. Initialize the semantic Q-function \(Q_{\mathrm{sem}}(s, \ell)\) and an empty replay buffer \(\mathcal{B}\).
  2. For each step \(t = 1, 2, 3, \ldots\) :
    1. Query the VLM for a small set of candidate prompts \(\mathcal{A}_{\mathrm{sem}}\) from the current observation and task goal.
    2. Sample a prompt \(\ell\) from \(\mathrm{softmax}(Q_{\mathrm{sem}}(s, \cdot))\) — prompts with higher Q-values are more likely.
    3. The VLA executes \(\ell\) as robot actions; observe reward \(r\) and next state \(s'\).
    4. Add the transition \((s, \ell, r, s')\) to the buffer \(\mathcal{B}\).
    5. Update \(Q_{\mathrm{sem}}\) toward \(r + V_{\mathrm{sem}}(s')\) with TD backups over \(\mathcal{B}\).

Tasks & Rollouts

Four long-horizon WidowX tasks.

Task suite. We evaluate SARL on complex, long-horizon WidowX tasks (above), plus ten tasks on the simulated LIBERO-LONG benchmark. These tasks are multi-step and require composing skills seen during pretraining (Bridge V2 for real, LIBERO-90 for sim) — the base policies achieve near-0% success on most of them, making them a clean test of deployment-time adaptation.

Task 1 — Move the hammer to the plate then grasp the mushroom.
SARL (Ours)
ICL VLM
DSRL
Residual RL
Zero-shot Base
Task 2 — Move the yellow sushi to the pot and pick up the spoon.
SARL (Ours)
ICL VLM
DSRL
Residual RL
Zero-shot Base
Task 3 — Move the banana to the pot on the right and sushi to the bowl.
SARL (Ours)
ICL VLM
DSRL
Residual RL
Zero-shot Base
Task 4 — Move the pot lid to the towel.
SARL (Ours)
ICL VLM
DSRL
Residual RL
Zero-shot Base

Results

SARL improves the base VLA's success rate from near 0% under the task prompt up to 60–80% after only 60–120 online episodes, in both the real world and simulation. It beats action-space RL methods (DSRL, Residual RL) that are fundamentally limited by the base policy's behavior under a single fixed prompt, and it beats an in-context-learning VLM baseline that proposes plausible commands but struggles to ground them in physical behavior.

Real-world WidowX results across four tasks.

Real-world WidowX. Across all four tasks, SARL achieves the best improvement of generalist-policy behavior in deployment. Each data point represents 10 evaluations.

Simulated LIBERO-LONG results across ten tasks.

Simulated LIBERO-LONG. On long-horizon LIBERO-LONG tasks, SARL outperforms DSRL, Residual RL, and the ICL VLM. It successfully adapts the policy on four tasks and matches performance on another that is already near-solved. Each point represents 64 evaluations with standard error over 3 seeds.

Insight 1: SARL decomposes complex goals — it doesn't just optimize the task prompt

Is SARL simply discovering a better task prompt? No. We evaluated the base policy under a wide range of task prompts, and on every long-horizon task no single prompt achieves any success — including the exact prompts SARL optimizes over. SARL succeeds by modulating the prompt within an episode, decomposing the goal into achievable stages.

Real-world Task VLM Best-of-10 a Prompt Cache b
Task 1 (long-horizon)0%0 / 100
Task 2 (long-horizon)0%0 / 100
Task 3 (long-horizon)0%0 / 100
Task 4 (short-horizon)50%26 / 100

a VLA success rate under the best of 10 task-prompt rephrasings generated by a VLM.
b Number of successful VLA rollouts under task prompts sampled from the set of prompts that SARL optimizes over.

Base policy on diverse task prompts. On all long-horizon tasks, no prompt alone achieves any success.

SARL vs. zero-shot base policy.

SARL vs. zero-shot prompting. The zero-shot base policy fails on long-horizon tasks; SARL solves them through learned prompting, sequencing skills covered under the pretraining distribution.

Insight 2: Grounding instructions is as important as decomposition

Like SARL, an in-context-learning VLM can decompose a complex goal into sub-instructions — but its commands are not grounded in what the VLA can actually execute. Over many episodes of experience, SARL learns the behavior each prompt induces, and this grounding is critical for performant adaptation.

SARL vs. in-context-learning VLM.

SARL vs. VLM prompting. The VLM picks commands that are semantically plausible but ungrounded: it references an out-of-distribution “hammer” by name (making the VLA grab a spoon) or picks the wrong verb (drop), prematurely releasing the sushi. SARL learns to select the prompts that actually work.

Insight 3: Semantic exploration unlocks adaptation where action-space RL can't learn

On long-horizon tasks where no fixed prompt yields any success (Insight 1), action-space RL methods (DSRL, Residual RL) cannot bootstrap at all. In contrast, SARL bootstraps by stitching together behaviors induced by different commands.

To further show that SARL's performance gains come from carefully modulating the VLA's prompt, we run an additional Residual Policy experiment where, instead of using a fixed task prompt, the ICL VLM changes the prompt within each episode.

Residual RL with VLM-modulated prompts vs. SARL.

VLM-prompted Residual RL. Even when the Residual policy's prompt is modulated within each episode by the ICL VLM, it fails to match SARL on our hardest tasks. This baseline succeeds only when the VLM happens to pick task-solving prompts—it cannot improve the prompt-selection strategy itself.

On short-horizon tasks (for example, Task 4), action-space RL remains powerful. We therefore view the shape of exploration induced by each of these methods—SARL and action-space RL—as complementary. We are excited by this finding, since it suggests that semantic exploration is a new lever the community can pull to design algorithms for efficient adaptation.

Training Timelapses (SARL)

Watch SARL learn each task online, sped up 30x.

Task 1 — Move the hammer to the plate then grasp the mushroom
Task 2 — Move the yellow sushi to the pot and pick up the spoon
Task 3 — Move the banana to the pot on the right and sushi to the bowl.
Task 4 — Move the pot lid to the towel

BibTeX

@misc{bhatia2026sarl,
  author    = {Bhatia, Jagdeep Singh and Wagenmaker, Andrew and Chen, William and Levine, Sergey},
  title     = {Adapting Generalist Robot Policies with Semantic Reinforcement Learning},
  year      = {2026},
}