Generative Actor-Critic with Soft Bridge Policies

Can an RL actor be expressive, soft, cheap, and better?

Ke He1 · Le He1 · Shunpu Tang2 · Yafei Wang3 · Lisheng Fan1

1Guangzhou University · 2Zhejiang University · 3Southeast University

Overview

SoftGAC keeps the multimodal power of latent generative policies, replaces endpoint entropy estimates with analytical path-wise soft regularization, and uses a single stochastic forward pass to keep action generation cheap while aiming for higher returns, not just faster inference.

In maximum entropy reinforcement learning, SoftGAC is positioned against diffusion policy, flow policy, and flow matching policy actors that rely on high-NFE generative sampling.

Expressive actor multimodal stochastic bridge policy
Principled softness exact path KL as control energy
Cheap by design comparable params, one sampled call
Hard-task gains stronger returns where high-NFE actors struggle
Humanoid Run
Dog Run
H1 Hurdle
H1 Run

The old bargain is too cramped.

Diffusion, score, and flow policies are attractive because their action distributions can be multimodal and highly non-Gaussian. That expressiveness is exactly what a diagonal Gaussian actor struggles to provide in contact-rich, high-dimensional control.

But MaxEnt RL also wants a real soft objective, and online actor-critic wants action generation to stay cheap during both environment interaction and policy updates. Existing generative actors often give up one demand in this bargain: softness becomes a bound, estimator, or proxy; cheapness becomes inference-only; or performance drops when the sampler is shortened.

A soft objective for expressive implicit actors.

The actor is expressive because it is an implicit sampler: it starts from noise, follows a short latent bridge, and outputs an action. Many internal paths can land in different action modes.

The difficulty is softness: the endpoint action density is hidden, so the usual SAC entropy term is not directly available. SoftGAC lifts the soft objective to the sampled path, while preserving the SAC endpoint target in the marginal action distribution.

Expressive

Sample paths, not a single mode

The actor can keep the multimodal behavior of latent generative policies without exposing a closed-form endpoint density.

Soft

Move entropy to the route

Regularize the generated path against a high-entropy reference bridge, instead of estimating endpoint entropy after the fact.

Principled

Keep SAC's endpoint goal

The path-space lift preserves the MaxEnt endpoint target; the extra regularization shapes how the action is reached.

SoftGAC bridge overview showing actor and reference latent paths with local KL terms that sum to control energy.
The critic still scores the terminal action. The soft penalty guides the route used by the structured bridge to reach that action.

Value-guided paths without losing stochasticity.

Think of the reference bridge as a broad, high-entropy way to move through latent space. The actor can spend control energy to bend those paths toward actions that the critic values.

A small budget keeps many routes alive; a larger budget concentrates probability near high-value modes. This is the visual intuition behind being expressive and soft at the same time.

2D bridge visualization showing critic modes, intermediate latent bridge densities, and terminal policy densities under different control-energy budgets.
In a controlled 2D critic, bridge densities move from the high-entropy reference toward value modes as the control-energy budget increases.

A faster actor could have been stronger all along.

The second obstacle is computation, but the goal is not merely to make sampling faster. Diffusion and flow actors usually generate one action by repeatedly calling a shared sampler, so high NFE increases latency, actor-update cost, activation memory, and the depth of backpropagation through the sampler. One-step and few-step reductions can lower sampling cost, but they often buy that speed by accepting weaker returns.

We argue against this old bargain, and believe that careful actor-structure design can make a faster actor stronger, matching or even outperforming high-NFE diffusion policies with a single sampled forward pass and a parameter budget comparable to strong actor baselines. The comparison is therefore not a larger network hiding cost in parameters, but many black-box sampler calls versus one structured white-box bridge call.

Cartoon comparison of diffusion or flow actors repeatedly calling a shared black-box sampler versus SoftGAC using one structured white-box bridge actor with local KL costs.
The actor budget is comparable: diffusion and flow policies reuse one shared sampler many times as NFE grows, while SoftGAC spends the budget once in a structured actor whose bridge transitions expose local KL costs.

Low cost is not bought by surrendering performance.

The experiments compare SoftGAC against diffusion and flow-matching actor-critic baselines under a unified JAX codebase, shared main critic update, comparable actor parameter budgets, and 8 seeds.

Across the full 12-task suite, the one-pass bridge actor is not merely competitive with high-NFE policies. The hard locomotion and obstacle tasks make the compute-return gap especially clear.

Full IQM learning curves on twelve benchmark tasks comparing SoftGAC with FLAC, DIME, FlowRL, QSM, QVPO, and CrossQ-SAC.
Across all 12 benchmark tasks, the full learning curves show that SoftGAC's gains are not caused by a small hard-task subset.
Bar charts showing per-action inference time and actor parameter count by domain.
Per-action latency stays in the low-latency one-pass regime, far below the high-NFE diffusion baselines, while actor parameters remain comparable.
Ablation curves comparing SoftGAC with and without the soft path regularizer.
Removing the path-space soft regularizer hurts consistently, showing that the gains are not only from the bridge architecture.
Compute-return tradeoff scatter plots for all benchmark tasks.
The Pareto structure is visible task by task: SoftGAC remains close to low-NFE latency while reaching higher or competitive IQM returns.
  • Expressiveness asks whether the actor can represent the action distribution the task actually needs.
  • Softness asks whether stochasticity is part of the objective rather than a post-hoc exploration trick.
  • Cheapness asks whether computation stays small during both training and deployment.
  • Better asks whether lower cost can come with higher returns, not just matching the slower actor.

SoftGAC ties the four goals to one structural choice: a short stochastic bridge actor. The path keeps generation expressive, the path-space objective keeps softness principled, one sampled pass keeps compute cheap, and the experiments show the result can be better rather than merely faster.

Generative Actor-Critic with Soft Bridge Policies

Ke He1, Le He1, Shunpu Tang2, Yafei Wang3, and Lisheng Fan1. Code is available with the paper.

1Guangzhou University · 2Zhejiang University · 3Southeast University

Citation

BibTeX

@article{he2026generative,
  title         = {Generative Actor-Critic with Soft Bridge Policies},
  author        = {He, Ke and He, Le and Tang, Shunpu and Wang, Yafei and Fan, Lisheng},
  journal       = {arXiv preprint arXiv:2605.08733},
  year          = {2026},
  archivePrefix = {arXiv},
  eprint        = {2605.08733},
  url           = {https://arxiv.org/abs/2605.08733}
}