Inference Guide#
This guide covers how to use trained Synora models for inference and deployment.
Overview#
Synora agents load from a checkpoint, run in eval() mode, and take the same
tensors they were trained on. Each model’s page documents the observation
layout it expects; this guide covers the mechanics around it.
Loading Trained Models#
from synora import DreamerAgent
# Load from checkpoint
agent = DreamerAgent.from_pretrained("path/to/checkpoint")
agent.eval()
Basic Inference#
Dreamer#
import torch
from synora import DreamerAgent
agent = DreamerAgent.from_pretrained("dreamer_checkpoint")
# Single step prediction. Observations are float tensors in [0, 1] shaped
# [batch, channels, height, width]; actions are [batch, action_dim].
obs = torch.rand(1, 3, 64, 64)
action = torch.tensor([[0.1, 0.2, 0.3, 0.4, 0.5, 0.6]])
with torch.no_grad():
next_obs, reward = agent.predict({"obs": obs, "action": action})
JEPA#
I-JEPA is evaluated with a frozen encoder, so inference means extracting representations rather than rolling out a policy. Load the EMA target-encoder from a training checkpoint and average-pool its patch tokens, exactly as the paper’s linear-evaluation protocol does:
import torch
from synora.training.eval_jepa import load_jepa_encoder, make_eval_transforms
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
encoder = load_jepa_encoder("results/jepa/jepa_run-latest.pth.tar", device)
transform = make_eval_transforms(crop_size=224)
images = torch.stack([transform(pil_image) for pil_image in batch]).to(device)
with torch.no_grad():
representations = encoder(images).mean(dim=1) # [batch, embed_dim]
To reproduce the paper’s ImageNet linear-probe number, use
synora.training.eval_jepa.jepa_linear_probe instead, which trains the
linear head on top of these features.
Rollout and Imagination#
Generate imagined trajectories:
# Dreamer imagination
from synora import DreamerAgent
agent = DreamerAgent.from_pretrained("dreamer_checkpoint")
initial_obs = torch.randn(3, 64, 64)
horizon = 10
imagined_trajectory = agent.imagine_rollout(initial_obs, horizon)
# Returns dict with imagined observations, actions, rewards
Batch Inference#
Process multiple inputs efficiently:
batch_size = 32
obs_batch = torch.randn(batch_size, 3, 64, 64)
action_batch = torch.randn(batch_size, 6)
with torch.no_grad():
predictions = agent.predict_batch({"obs": obs_batch, "action": action_batch})
GPU Acceleration#
Move to GPU for faster inference:
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
agent = agent.to(device)
inputs = {key: value.to(device) for key, value in inputs.items()}
with torch.no_grad():
output = agent.predict(inputs)
Real-time Inference#
For interactive applications:
import torch
from synora import DreamerAgent
class InferenceServer:
def __init__(self, device="cuda"):
self.device = torch.device(device)
self.agent = DreamerAgent.from_pretrained("checkpoint").to(self.device).eval()
def predict(self, obs, action):
inputs = {
"obs": obs.to(self.device),
"action": action.to(self.device),
}
with torch.no_grad():
return self.agent.predict(inputs)
server = InferenceServer()
Performance Optimization#
The Efficient Inference and Deployment covers this in depth. In short:
import torch
from synora.inference import optimize_for_inference, make_stepper
# eval mode, no grad, bf16 where supported, torch.compile with CUDA graphs
actor = optimize_for_inference(agent.dreamer.actor, precision="auto", compile=True)
# or drive the whole observe -> act loop through a uniform, batchable stepper
stepper = make_stepper(agent)
state = stepper.init_state(batch_size=1)
with torch.inference_mode():
state = stepper.observe(state, obs, prev_action)
action = stepper.act(state)
Prefer torch.compile (via optimize_for_inference or synora.maybe_compile)
over torch.jit.script, which is in maintenance mode upstream. Always run inference under
torch.inference_mode(): it skips autograd bookkeeping that no_grad still
performs.
Exporting Models#
Every model class in the top-level synora namespace has a deployment-oriented export() method, and synora.export_model(module, ...) exports any other nn.Module with the same options. High-level wrapper agents such as Dreamer and PlaNet use the same exporter for their contained modules:
model.export("model.pt2", format="exported_program", example_inputs=example_inputs)
agent.export("agent_actor.onnx", format="onnx")
Format |
Alias |
Output |
|---|---|---|
|
|
|
|
|
AOTInductor package, runnable from C++ without the model’s Python code |
|
— |
ONNX graph for ONNX Runtime, TensorRT conversion, or other production runtimes |
|
|
TorchScript |
|
|
Torch-TensorRT module, compiled through the dynamo IR (requires |
Load any artifact back with synora.load_exported(path), and check it against
the eager module with synora.verify_export(module, path, example_inputs).
For whole deployment bundles (weights, config, artifacts and a manifest), see
Efficient Inference and Deployment.
Dreamer exports its deterministic actor by default. The exported Dreamer actor
accepts concatenated latent features with shape [batch, stoch_size + deter_size]
and returns actions:
import torch
from synora import DreamerAgent
agent = DreamerAgent(env="cartpole_balance")
agent.export("dreamer_actor.onnx", format="onnx")
agent.export("dreamer_actor.pt", format="torchscript")
features = torch.zeros(1, agent.args.stoch_size + agent.args.deter_size)
agent.export(
"dreamer_actor_dynamic.onnx",
format="onnx",
example_inputs=features,
input_names=["features"],
output_names=["actions"],
dynamic_axes={"features": {0: "batch"}, "actions": {0: "batch"}},
)
Export individual components by passing target when the agent provides more
than one deployable module:
agent.export("dreamer_encoder.onnx", format="onnx", target="obs_encoder")
agent.export("dreamer_reward.pt", format="torchscript", target="reward_model")
For any lower-level torch.nn.Module model, pass example_inputs explicitly if Synora cannot infer a safe default:
import torch
import synora
genie = synora.create_model("genie-small", image_size=32)
video = torch.randn(1, 3, genie.num_frames, genie.image_size, genie.image_size)
genie.export("genie_small.onnx", format="onnx", example_inputs=video)
vit = synora.VisionTransformer(img_size=[224])
images = torch.randn(1, 3, 224, 224)
vit.export("vit.onnx", format="onnx", example_inputs=images)
Agents that contain multiple deployable modules accept either short target names such as "obs_encoder" or fully qualified paths such as "dreamer.obs_encoder". JEPA exports a ViT encoder target by default, while lower-level JEPA VisionTransformer modules have .export() themselves.
TensorRT export requires torch-tensorrt in the deployment environment:
agent.export("dreamer_actor_trt.ep", format="tensorrt")
Note that the exported Dreamer actor above uses ActionDecoder(deter=True),
which estimates the mode from 100 random samples. For a deterministic,
single-pass policy, export DreamerStepper(agent).step_module() instead (see
Efficient Inference and Deployment).
Integration Examples#
With Gym Environments#
import synora
from synora import DreamerAgent
env = synora.make_env("Pendulum-v1", backend="gym")
agent = DreamerAgent.from_pretrained("pendulum_checkpoint")
obs, _ = env.reset()
done = False
while not done:
action = agent.act(obs) # Get action from agent
obs, reward, done, _, _ = env.step(action)
With Custom Environments#
class CustomEnv:
def step(self, action):
# Your environment logic
return obs, reward, done
env = CustomEnv()
agent = DreamerAgent.from_pretrained("custom_checkpoint")
for episode in range(10):
obs = env.reset()
total_reward = 0
while True:
with torch.no_grad():
next_obs_pred, reward_pred = agent.predict({"obs": obs, "action": action})
# Use predictions for planning/control
action = agent.plan(obs, next_obs_pred, reward_pred)
obs, reward, done = env.step(action)
total_reward += reward
if done:
break
print(f"Episode {episode}: {total_reward}")
Troubleshooting#
Memory Issues#
Use smaller batch sizes
Enable gradient checkpointing
Clear cache:
torch.cuda.empty_cache()
Speed Issues#
Move to GPU
Use
optimize_for_inference(..., compile=True)(CUDA graphs) and bf16Batch inputs when possible; for Genie generation pass
use_cache=TrueMeasure with
synora.inference.benchmark_stepbefore and after each change
Accuracy Issues#
Ensure inputs are normalized the same way as during training
Check model loading
Verify input shapes match training