← Blog

Image-to-3D on NVIDIA L4 and Blackwell: Shipping Hunyuan3D 2.1 on Cloud Run GPUs

Every 3D generation lane on three.ws runs on NVIDIA silicon. Moving our realism lane from Hunyuan3D 2.0 to 2.1 cost us two days on problems that were invisible from the model card and from the platform docs alike. This is the field report: what failed, what the logs actually said, and what we changed. If you are putting a large diffusion-plus-mesh pipeline on managed GPU infrastructure, two of these will probably bite you too.

The short version

  1. The fleet
  2. Why 2.1 was worth the trouble
  3. Wall 1: memory-backed /tmp
  4. Wall 2: Blackwell needs sm_120
  5. The quota surprise
  6. Designing for a one-line rollback
  7. Spending the GPU budget
  8. A checklist for your own port

The fleet

three.ws runs two distinct NVIDIA layers. The first is a free hosted lane: one NVIDIA_API_KEY unlocks TRELLIS for text-to-3D, FLUX.1-schnell for images, the Llama and Nemotron chat lineup, embeddings, reranking, safety, and speech. That layer is documented model by model in NVIDIA models on three.ws and it is not what this post is about.

The second is our self-hosted GPU fleet: twelve Cloud Run services, each a FastAPI worker wrapping one model, autoscaled per lane. Eleven run on nvidia-l4. One runs on nvidia-rtx-pro-6000.

LaneWorkerGPUCPU / memory
Image to 3D (PBR)model-hunyuan3d-21-rtxnvidia-rtx-pro-600020 / 80 Gi
Image to 3D (fallback)model-hunyuan3dnvidia-l48 / 32 Gi
Text to 3Dmodel-trellisnvidia-l48 / 32 Gi
Image to 3D (fast)model-triposg, model-triposrnvidia-l48 / 32 Gi, 4 / 16 Gi
Texture synthesistexturenvidia-l48 / 32 Gi
Auto-riggingrig, unirignvidia-l44 / 16 Gi
Text to motionmodel-text2motionnvidia-l44 / 16 Gi
Video to scenemodel-video2scenenvidia-l48 / 32 Gi

The L4 is a genuinely excellent default for this workload: 24 GiB of VRAM, a low power envelope, and wide availability. Nine of these lanes never needed anything else. The tenth did, and finding out why took longer than it should have.

Why 2.1 was worth the trouble

Hunyuan3D 2.0 gives you a mesh with a single baked diffuse texture. It reads fine in a product thumbnail and it reads like plastic under real lighting, because a diffuse map carries no information about how a surface responds to light.

Hunyuan3D 2.1 runs a shape DiT for geometry, then a separate PBR paint pass: multiview PBR diffusion, RealESRGAN super-resolution, a texture bake, and an inpaint step to close seams. It exports a GLB with a true physically-based material set:

PBR maps are the single biggest realism lever available in a generated asset, and they are why the same model dropped into an AR scene under a real room's lighting stops looking like a toy. That was the whole motivation. Everything below is what stood between the motivation and the deploy.

Wall 1: memory-backed /tmp

The 2.1 service on an L4 never finished loading. Not slowly. Not intermittently. Every single cold start ended the same way:

Container terminated on signal 9

Signal 9 is SIGKILL. On Cloud Run that almost always means the instance exceeded its memory limit and the platform killed it. Which was confusing, because the model fits comfortably in 24 GiB of VRAM. The problem was never VRAM.

The arithmetic nobody does until it fails

Our worker stages weights from a bucket into /tmp, then loads them. Two numbers matter:

On Cloud Run, /tmp is a tmpfs. It is not disk. Every byte written there is a byte of the instance's memory allocation, and it stays allocated until you delete the file. So the peak is not 14 GiB and it is not 18 GiB. It is both at once, against the L4 tier's ceiling of 32 GiB:

18 GiB  staged weights, resident in tmpfs
+ 14 GiB  model loaded into process memory
--------
  32 GiB  peak, against a 32 GiB limit
  → SIGKILL, every time

This is worth internalizing as a general rule, because it is not specific to us or to this model: on any platform where /tmp is memory-backed, staging a large artifact before loading it doubles your peak. The same code on a VM with a real disk works fine, which is exactly why it survives local testing and dies in production.

Why the health check looked green

The part that cost us the most time

Our worker opens its HTTP port immediately and loads the pipeline in a background task, so cold starts do not fail the platform's startup probe. That is a deliberate and common design. The consequence is that GET /health answers 200 while the model is still loading, which means a service that can never finish loading still looks alive.

The fix is not to remove the background load. It is to make readiness a distinct field that only flips when the pipeline is actually usable, and to surface the failure rather than swallowing it:

{
  "ok": true,
  "model": "hunyuan3d-2.1",
  "gpu_available": true,
  "gpu_name": "NVIDIA L4",
  "pipeline_loaded": true,
  "ready": true,
  "load_error": null
}

ok is liveness: the process is up. ready is the one your router should gate on. load_error carries a sanitized string so a failed load is a diagnosis instead of a mystery. If you take one operational idea from this post, take this one: a boolean that means "the port is open" and a boolean that means "this instance can do work" are different booleans, and conflating them turns a five-minute fix into a two-day hunt.

The two real fixes

There are exactly two ways out, and they are worth knowing both:

  1. Stage incrementally. Fetch one weight subtree, call from_pretrained on it, delete the staged copy, move to the next. The staged and resident copies never coexist at full size, and peak memory drops to roughly the resident footprint plus one subtree. This is the right fix if you are pinned to a memory-constrained tier.
  2. Move to a tier whose floor clears the peak. The platform minimums for nvidia-rtx-pro-6000 on Cloud Run are 20 CPU and 80 GiB, which clears a 32 GiB peak with room to spare and stops the problem being a problem.

We took the second, because we wanted the faster card anyway and because of the quota situation below. The first remains the correct fix for anyone who wants 2.1 on an L4, and it is a contained change: it lives entirely inside the staging function.

Wall 2: Blackwell needs sm_120

Switching GPU type should be a one-line change to a deploy config. It was not, and the reason is a good thing to have in your head before you plan a Blackwell migration.

RTX PRO 6000 is Blackwell, and Blackwell is compute capability 12.0. Our L4 image was built on CUDA 12.4 with torch 2.5 (cu124). The cu124 wheels predate that architecture and ship no sm_120 kernels at all. There is no runtime flag, no fallback path, and no JIT rescue that makes prebuilt cu124 binaries emit Blackwell code. You rebuild or you do not run.

The RTX image is the same application code on a different foundation:

L4 imageRTX PRO 6000 image
ArchitectureAda Lovelace, sm_89Blackwell, sm_120
CUDA12.412.8
PyTorch2.5 (cu124)2.7.1 (cu128)
Extensions built for8.98.9;12.0

Note the last row. We compile the custom CUDA extensions for both architectures, not just the one we deploy to:

ENV TORCH_CUDA_ARCH_LIST="8.9;12.0"

It costs build minutes and a little image size. It buys something worth much more: one image that boots on either GPU type. That turns "which card is available in this region today" from a rebuild into a deploy flag, and it turns a rollback into a redeploy rather than a re-architecture. If you are building images for a fleet that spans GPU generations, build fat and pick at deploy time.

The same discipline applies one level up. Hunyuan3D 2.0 and 2.1 have mutually incompatible Python stacks (torch 2.3 / cu121 with hy3dgen against torch 2.5 / cu124 with hy3dshape and hy3dpaint). Rather than fight that, we run them as two separate services. Trying to unify them in one image would have cost days and produced something more fragile than two clean containers.

The quota surprise

This one is not code, and it is the finding most likely to save someone a planning cycle.

The intuitive assumption is that the small, cheap, mature GPU is the easy one to get, and the big new one is the scarce one you have to beg for. In us-central1, for us, it was the reverse. Our L4 quota was 3 GPUs, shared across the entire fleet and permanently pinned at that ceiling. Every new L4 lane was a fight with the lanes we already had. The RTX PRO 6000 quota in the same region was granted at a far higher number.

Practical takeaway

Check your granted quota per GPU type, per region before you choose a tier on price or spec. The larger card can be the more available card, and "available" beats "theoretically cheaper" every time you are trying to ship. Note also that a granted quota number and what deploy-time enforcement lets you run can differ, so confirm with an actual deploy rather than a dashboard reading.

Our RTX service runs with min instances and max instances both set to 1, held warm. Nothing about that is elegant, but for a lane where a cold start means loading 14 GiB before the first token of work, a warm instance is the difference between a 20-second response and a 4-minute one. Warm capacity is the cheapest latency optimization available to anyone running large models behind a request path.

Designing for a one-line rollback

The best decision in this whole migration was made before any of it started: every image-to-3D worker speaks the same wire contract. Same request shape, same response shape, same task-polling semantics, regardless of which model or GPU is behind it.

POST /infer     → 202 { "task_id": "...", "status": "queued" }
GET  /tasks/:id → { "status": "done", "result_gcs_url": "...", "elapsed_ms": 224140 }
GET  /health    → { "ok": true, "ready": true, "load_error": null }

Because of that, rolling the realism lane back from 2.1 to 2.0 is repointing one environment variable at a different service URL. No rebuild, no code change, no redeploy of the caller. When you are moving a production lane onto new hardware, the ability to undo it in thirty seconds is what makes it reasonable to try at all.

Two smaller decisions carried more weight than they looked like they would:

Spending the GPU budget

With the memory ceiling gone, the interesting question becomes how to spend the compute. The 2.1 budget splits between the shape DiT (inference steps and marching-cubes octree resolution) and the paint pass (how many views, and at what resolution each view diffuses). We expose three tiers:

TierShape stepsOctreePaint viewsPaint resolution
draft302566512
standard503846512
high (default)505126768

The non-obvious choice is that the multiview count stays at 6 across all three tiers. Views are the expensive axis: each additional view is another full diffusion pass held in VRAM alongside the shape DiT, the DINOv2 reference encoder, and RealESRGAN. Holding views constant and spending the extra budget on octree resolution and per-view resolution buys sharper geometry and crisper PBR maps without pushing VRAM into swap-or-die territory. The texture atlas is pinned high at load time regardless of tier (2048 render, 4096 texture), because atlas resolution is cheap relative to what it contributes.

The elapsed_ms in the response above is real: a high-tier generation is a multi-minute job. That is the correct trade for us. GPU time is the cheapest input in this pipeline and a user's opinion of the result is the most expensive output, so quality wins by default and the fast lanes exist for people who explicitly ask for one.

A checklist for your own port

If you are about to put a large image-to-3D or diffusion pipeline on managed GPU infrastructure, this is the list we wish we had started with:

  1. Find out whether /tmp is memory-backed. If it is, add your staged size to your resident size and compare that sum, not either number alone, to the tier's memory limit.
  2. Separate liveness from readiness. Return both, gate routing on readiness, and surface the load error in the payload.
  3. Read your granted GPU quota per type and per region before you pick a tier. Then verify it with a real deploy.
  4. Match the CUDA toolkit to the target architecture. Blackwell is sm_120 and needs cu128 wheels; no flag substitutes for the rebuild.
  5. Set TORCH_CUDA_ARCH_LIST to every architecture you might deploy to, not just today's. Fat images make GPU choice a deploy-time decision.
  6. Persist task state outside the instance the moment you enable autoscaling.
  7. Validate remote fetches on every redirect hop, not only the submitted URL.
  8. Keep one wire contract across model versions so rollback is an environment variable.
  9. Keep the old lane deployed while the new one earns its place. A warm fallback costs less than an outage.

One licensing note, because it is easy to skip and expensive to skip: generative 3D checkpoints ship under a wide spread of terms, and several popular ones carry non-commercial or otherwise restricted licenses that differ from the permissive license on the surrounding code. Read the license on the exact checkpoint you intend to deploy, including any super-resolution or encoder models pulled in as dependencies, before it reaches production.


Where this runs

Everything above is in production behind the forge, the text and image to 3D surface on three.ws. The free lane needs no key and no account; the paid lanes route to the self-hosted fleet described here. Generated models drop straight into AR, the avatar studio, and the <agent-3d> web component for embedding on any site.

three.ws joined the NVIDIA Inception program in July 2026. It is a startup program rather than a partnership or an investment, and in practice it means GPU capacity and engineering access on the exact constraint this post is about. The work of relaxing that constraint is ongoing, and we will keep publishing the numbers as we get them.

Related reading


← All posts