Sample paths, not a single mode
The actor can keep the multimodal behavior of latent generative policies without exposing a closed-form endpoint density.
Can an RL actor be expressive, soft, cheap, and better?
1Guangzhou University · 2Zhejiang University · 3Southeast University
Overview
SoftGAC keeps the multimodal power of latent generative policies, replaces endpoint entropy estimates with analytical path-wise soft regularization, and uses a single stochastic forward pass to keep action generation cheap while aiming for higher returns, not just faster inference.
In maximum entropy reinforcement learning, SoftGAC is positioned against diffusion policy, flow policy, and flow matching policy actors that rely on high-NFE generative sampling.
Diffusion, score, and flow policies are attractive because their action distributions can be multimodal and highly non-Gaussian. That expressiveness is exactly what a diagonal Gaussian actor struggles to provide in contact-rich, high-dimensional control.
But MaxEnt RL also wants a real soft objective, and online actor-critic wants action generation to stay cheap during both environment interaction and policy updates. Existing generative actors often give up one demand in this bargain: softness becomes a bound, estimator, or proxy; cheapness becomes inference-only; or performance drops when the sampler is shortened.
The actor is expressive because it is an implicit sampler: it starts from noise, follows a short latent bridge, and outputs an action. Many internal paths can land in different action modes.
The difficulty is softness: the endpoint action density is hidden, so the usual SAC entropy term is not directly available. SoftGAC lifts the soft objective to the sampled path, while preserving the SAC endpoint target in the marginal action distribution.
The actor can keep the multimodal behavior of latent generative policies without exposing a closed-form endpoint density.
Regularize the generated path against a high-entropy reference bridge, instead of estimating endpoint entropy after the fact.
The path-space lift preserves the MaxEnt endpoint target; the extra regularization shapes how the action is reached.
Think of the reference bridge as a broad, high-entropy way to move through latent space. The actor can spend control energy to bend those paths toward actions that the critic values.
A small budget keeps many routes alive; a larger budget concentrates probability near high-value modes. This is the visual intuition behind being expressive and soft at the same time.
The second obstacle is computation, but the goal is not merely to make sampling faster. Diffusion and flow actors usually generate one action by repeatedly calling a shared sampler, so high NFE increases latency, actor-update cost, activation memory, and the depth of backpropagation through the sampler. One-step and few-step reductions can lower sampling cost, but they often buy that speed by accepting weaker returns.
We argue against this old bargain, and believe that careful actor-structure design can make a faster actor stronger, matching or even outperforming high-NFE diffusion policies with a single sampled forward pass and a parameter budget comparable to strong actor baselines. The comparison is therefore not a larger network hiding cost in parameters, but many black-box sampler calls versus one structured white-box bridge call.
The experiments compare SoftGAC against diffusion and flow-matching actor-critic baselines under a unified JAX codebase, shared main critic update, comparable actor parameter budgets, and 8 seeds.
Across the full 12-task suite, the one-pass bridge actor is not merely competitive with high-NFE policies. The hard locomotion and obstacle tasks make the compute-return gap especially clear.
SoftGAC ties the four goals to one structural choice: a short stochastic bridge actor. The path keeps generation expressive, the path-space objective keeps softness principled, one sampled pass keeps compute cheap, and the experiments show the result can be better rather than merely faster.
Ke He1, Le He1, Shunpu Tang2, Yafei Wang3, and Lisheng Fan1. Code is available with the paper.
1Guangzhou University · 2Zhejiang University · 3Southeast University
@article{he2026generative,
title = {Generative Actor-Critic with Soft Bridge Policies},
author = {He, Ke and He, Le and Tang, Shunpu and Wang, Yafei and Fan, Lisheng},
journal = {arXiv preprint arXiv:2605.08733},
year = {2026},
archivePrefix = {arXiv},
eprint = {2605.08733},
url = {https://arxiv.org/abs/2605.08733}
}