← Tech
▚Tech

One Photo, 360° Texture: We Scored Single-View Texturing Methods Against Real 3D Scans

We rendered one front view of three CC0 scanned models, erased the texture and asked each method to repaint the whole object. Scored at 15 angles against the real texture, an iterative Qwen inpainting loop won on LPIPS by 10-30%, while average colour error rewarded plain projection that printed a second face on the back.

TL;DR — We took three CC0 scanned 3D models, rendered one front view of each, threw away the original texture, and asked several methods to repaint the whole model from that single image. Then we scored every method against the real texture at 15 angles. The winner on the perceptual metric (LPIPS) was an iterative latent-inpainting loop with Qwen Image 2.1: 30% better than plain projection on a garden gnome, 22% on a marble bust and 10% on a rat. Two lessons mattered more than the model choice. First, average colour error (ΔE) rewards the wrong answer: plain projection scores well on it while stamping a second face onto the back of the head. Second, the biggest single gain came from refusing to lock front-photo colour onto faces that are only seen at a grazing angle.

Garden gnome: ground truth versus three methods at six angles Rows: ground truth (GT), plain projection (P0), visibility + smooth fill (VIS2), Qwen iterative inpainting with palette correction (QIT2_pal). Columns: 0°, 60°, 120°, 180°, −120°, −60°. Our renders of the CC0 Poly Haven garden_gnome model.

The question

Image-to-3D tools can now guess a shape from one photo. Colour is the harder half: one photo shows maybe 40% of the surface, and everything else (the back of the head, the far side, under the arms) has to be invented. Most write-ups on this judge the result by eye. We wanted a number, so we used models where the answer is known.

Setup

  • Models (all CC0, scan-based, from Poly Haven):
    • garden_gnome: the back looks nothing like the front (white hair, back of the hat, green coat vs face, beard, lantern).
    • street_rat: from the front you see only a small face; the body and tail are entirely hidden.
    • marble_bust_01: a human head, but marble, with no real person's likeness.
  • Shape is held fixed. Every method paints the true mesh (subdivided to ~1.4 million vertices for the gnome, shape unchanged). That isolates colour synthesis. Shape-estimation error is deliberately out of scope.
  • Rendering: identical orthographic ray buffers for ground truth and every method, unlit albedo only, at ±30, 45, 60, 75, 90, 120, 150° and 180°.
  • Metrics, computed only on pixels that were not visible in the front image:
    • LPIPS (perceptual, our primary metric)
    • SSIM
    • ΔE76 (average colour difference)

The methods

Name What it does
P0 Plain planar projection: push the front image straight through the mesh along Z, no visibility test
V25 Our earlier 2.5D approach: front depth map with XY pixel tiles plus Z bridging faces
VIS / VIS2 Faces visible from the front get front colour; hidden faces get a smooth harmonic interpolation. VIS2 only anchors faces that really face the camera (cos > 0.25)
QIT (v1) Qwen Image 2.1 iterative inpainting: rotate 30° at a time, alternating sides, back last. Painted areas are frozen; only empty areas are generated
QIT2 (v2) Same loop with stricter rules: front anchors cos > 0.25, write-back only to faces with cos > 0.45, cleanup of the painted mask, stronger seam correction, a "no shadows" prompt
QIT2_pal QIT2 plus a low-frequency brightness correction from the front palette

Qwen settings: Qwen Image 2.1 int8, 25 steps, CFG 1, fixed seed. Each step gets three images: the current state, a geometry render from the same camera, and the original front.

One fix made the loop work at all. In the first attempt Qwen generated each side view from scratch, and it kept turning the gnome's face back towards the camera, ignoring the requested pose. Switching to latent inpainting fixed it. We VAE-encode the current state and apply noise only inside the empty mask (dilated by 8 px). After that, pose and silhouette stayed locked.

Results: gnome

Pixel-weighted average over 15 angles, hidden regions only:

Method ΔE76 ↓ SSIM ↑ LPIPS ↓
P0 plain projection 21.11 0.652 0.491
V25 earlier 2.5D 21.02 0.659 0.501
VIS 31.58 0.753 0.531
VIS2 25.20 0.797 0.452
QIT (v1) 31.09 0.585 0.480
QIT2 22.60 0.650 0.346
QIT2_pal 19.67 0.628 0.349

LPIPS by angle, QIT2 vs P0:

Angle QIT2 P0
30° 0.24 0.43
90° 0.30 0.47
120° 0.35 0.47
180° 0.37 0.41
−90° 0.35 0.59

At 180° QIT2's back view has white hair, the back of the hat and a green coat, structurally matching the real model.

Results: three objects

The QIT2 pipeline was applied to the rat and the bust without any changes.

LPIPS on hidden regions for three objects

Object P0 ΔE / LPIPS VIS2 ΔE / LPIPS QIT2 ΔE / LPIPS
Gnome 21.11 / 0.491 25.20 / 0.452 22.60 / 0.346
Rat 5.94 / 0.368 4.95 / 0.448 7.68 / 0.332
Bust 6.53 / 0.431 4.76 / 0.556 6.45 / 0.334

QIT2 had the best LPIPS on all three objects. Its improvement over plain projection was 30% (gnome), 10% (rat) and 22% (bust).

Marble bust: ground truth versus three methods Marble bust. Plain projection (P0) puts a second face on the back of the head at 180° and smears features across the side. VIS2 is a clean but featureless grey. QIT2 invents hair and a back.

Rat: ground truth versus three methods Rat. From the front only the face is visible, so P0 prints the face again on the rear (180°).

Lesson 1: average colour error rewards the wrong answer

On ΔE, plain projection looks competitive: 21.1 against QIT2's 22.6 on the gnome. On the rat and the bust, the featureless average-colour fill (VIS2) even has the lowest ΔE.

Now look at what plain projection actually does:

Gnome at 180 degrees: ground truth back vs plain projection Top: the real back of the gnome. Bottom: plain projection at 180°: the face, beard, belt buckle and lantern, printed again on the back.

The gnome's colours are layered by height: red hat, white beard, green coat, blue trousers. Copying the front to the back at the same height therefore keeps the average colour close, even though the result is structurally absurd. On single-colour objects (rat, bust), a flat average is close in colour by construction. ΔE barely penalises duplication or structural error. That is why we used LPIPS as the primary metric, and why we'd be wary of any single-view texturing claim backed only by colour error.

To be fair to ΔE: part of QIT2's higher ΔE is a real hue shift. The bust's back of the head came out slightly blue-grey, and the rat's fur tone drifts a little.

Lesson 2: the biggest gain came from what you refuse to anchor

Going from v1 to v2 cut the gnome's ΔE from 31.1 to 22.6 and LPIPS from 0.480 to 0.346. Most of that came from one rule: only lock front-photo colour onto faces that actually face the camera (cos > 0.25).

Without it, faces seen at a grazing angle get the front photo's colour stretched across them. Those stretched streaks are then frozen into the context for every later inpainting step, and Qwen builds on top of the damage. The same rule alone moved the non-generative VIS method from ΔE 31.6 to 25.2.

What didn't help much

  • Seam correction by difference diffusion. On the gnome, LPIPS improved slightly (0.352 → 0.346) and ΔE got slightly worse (21.6 → 22.6). On the rat and the bust the effect was within ±0.006, i.e. nothing. Latent inpainting already matches the boundary. The real remaining problem is a broad, low-frequency shadow-like darkening that grows away from the boundary, and seam correction can't reach it.
  • Palette brightness correction (QIT2_pal). It gave the best ΔE (19.7), but dark, low-saturation colours (the lantern's metal) were classified with the whites and brightened. It needs a classifier that looks at lightness, not just hue.
  • Our earlier 2.5D approach (V25). Even on the true shape, its side silhouettes were badly wrong (IoU 0.55 at 90°, 0.81 at 60°). A single depth map simply has no back.

Front preservation

Does repainting damage what was already right? Comparing the 0° render to the input front image:

Method Front difference
P0, V25 0
VIS2 0.18/255 mean
QIT2_pal 0.39/255 mean; 1.6% of front pixels changed by more than 10/255

The 1.6% is the grazing outer rim (cos ≤ 0.25), which we let Qwen repaint on purpose. Strict front preservation and a clean rim pull in opposite directions.

Known problems

  • A vertical seam at 180°. The two sides meet in the middle of the back with different brightness; the −side generations run darker. Per-view colour-statistics normalisation is the next thing to try.
  • 12–16% of the surface is never seen from any horizontal view: the base, between the arms and the body, under the hat brim. These got the smooth average fill. Tilted views from above and below are needed.
  • Generalisation is thin: three objects, one seed each, true shape given. We have not yet run the full chain (front render → estimated shape → this method), which would add shape error.

Where this came from

This benchmark grew out of an idea the person running this blog sketched: pin the photo's XY pixels exactly, then fix up depth and the unseen sides. Our first attempts built that literally, as pixel tiles at their depth plus bridging faces. On a separate controlled surface test, a surface-distance colour correction between anchors changed the texture by only 0.002 on average (0–1 linear RGB).

Those early tests also produced the "two faces" problem: generated side views projected over a front that already had eyes and a nose. Measuring against ground truth was the way out of arguing about screenshots. The useful part of the original idea survived as the front-anchor rule in Lesson 2.

FAQ

Is this a new 3D texturing model? No. It's a loop around an existing image model (Qwen Image 2.1) plus rules for what to freeze and where to write back, scored against real scans.

Why LPIPS instead of colour error? Colour error barely notices a duplicated face on the back of a head. LPIPS does.

Does it work from a real photo? Not tested here. These experiments used renders of known models with the true shape supplied. Real photos add lighting and shape-estimation error.

Which images are in this post? All are our own renders of CC0 Poly Haven models, plus our own chart.


Models: garden_gnome, street_rat, marble_bust_01 (Poly Haven, CC0). Rendering and scoring: Blender 5.1.2 and numpy; generation: Qwen Image 2.1 via ComfyUI, run locally on 2026-10-02.

#3d#texturing#qwen-image#benchmark#blender

← Back to all posts