All articlesStandard
Fitting Photography Guidelines: Capture Standards for AI Imagery
What a fitting shoot has to deliver before a generation model can use it — the geometry, the light, the poses and the files.

A generation model does not improve a photograph — it inherits it. Feed it a fitting shoot where the lens changed between looks, the key light drifted, or the model stepped off the floor mark, and every one of those inconsistencies is carried faithfully into the output. The standard below is what a fitting shoot has to deliver before it can be used as input: measured once, taped to the floor, and unchanged for the whole batch.
This is the capture specification we issue to partner studios shooting fitting imagery for AI-generated PDP sets. It is deliberately rigid. Every rule in it exists because breaking it produced a visible defect downstream.
What the inputs actually decide
A fitting shoot is not a campaign. Nothing here is meant to be beautiful — it is meant to be measurable. The generation model reads three things from every frame:
- The silhouette. Where the garment ends and the body begins, in every view.
- The material. True colour, weave, sheen and how the fabric falls.
- The geometry. Where the camera stood, how high, and through what focal length.
Get the first two right and the model knows what the garment is. Get the third right and it knows how to place that garment on any body it generates later. A batch that drifts on geometry mid-shoot produces a range whose images will not sit together on a category grid — the failure shows up not in one image, but across twenty.
Nothing here is meant to be beautiful. It is meant to be measurable.
Contrast: the bodysuit and the background
The fitting model always wears a base bodysuit under the garment, and the bodysuit tone must contrast with the product — never match it.
- Light garment — dark bodysuit. Accepted.
- Dark garment — light bodysuit. Accepted.
- Similar tones. Not accepted. The garment edge cannot be separated from the body, and the silhouette the model extracts will be wrong.
The bodysuit is the only base layer, and its job is to disappear. Clean tonal separation is what lets the model isolate the exact garment shape rather than guessing at it.
The background runs on the same logic and one extra rule: it never changes. One seamless, neutral, evenly washed colour for the entire programme. The photographer checks before shooting each look that the garment does not blend into it — and when something does blend, the fix is the bodysuit and the styling, never a different backdrop.
The set: one straight axis
Background, model, all three cameras and the key light sit on a single straight line. Distances are measured once, taped on the floor, and never adjusted.
Light: flat, frontal, nothing behind
One wide frontal diffused source aimed at the model, plus an even diffused wash on the background. That is the whole lighting plan.
Flat, even light gives true fabric colour and texture and no drama — which is exactly what is wanted here. Hard shadows, backlight, coloured gels and on-camera flash are all disallowed. Every one of them writes lighting information into the frame that the generation model then has to be told to ignore.
Cameras: height, tilt and lenses
All three cameras are perfectly level. Zero tilt. A tilted camera bends verticals and warps the garment silhouette, and no amount of downstream correction puts that back.
Nothing wider than 100 mm
Wide lenses add perspective distortion that transfers directly into the generations and warps garment proportions. The rule is absolute:
- Camera A — full body. 100 mm minimum, 150–200 mm recommended for flatter perspective and easier staging.
- Cameras B and C — close-ups. Fixed 100 mm on both.
- Anything from 35 to 85 mm. Not accepted anywhere in the setup.
Pick one focal length for camera A and keep it for the whole programme. Changing focal length between batches changes perspective, and that breaks input consistency just as thoroughly as moving the camera would.
Once it is aligned, nothing moves
Tripod height, tilt, distance and focus distance stay fixed for the entire shoot. No reframing, no zooming, no stepping closer for a better look at a detail. If an item has to be re-shot later, the same geometry is rebuilt from the logged measurements before the item goes back on the mark — and any change of geometry mid-batch makes the inputs inconsistent, which fails QC for the whole batch rather than the one frame.
Poses: eight views, ten frames
A strict A-pose captured from eight views — front, back, left, right, and three-quarters front and back on both sides — plus two relaxed poses that show how the fabric behaves in motion. The model rotates on the floor mark; the cameras never move.
The two relaxed poses come last and are shot on camera A only. They are natural stances in movement, and their whole job is to show drape, folds and fabric behaviour — the things a static A-pose cannot tell you.
That makes ten poses per look: camera A shoots all ten, and the close-up camera shoots the eight A-pose views. Eighteen frames per SKU.
Framing: 80% fill
The subject occupies about 80% of the frame on every camera — the model for camera A, the product for the close-up cameras.
Files: resolution, format and naming
Deliver lossy WebP at quality 95, converted directly from the camera originals — and nothing else done to them.
- Long edge. 5000 px minimum on the longest side of every frame, on all three cameras.
- File size. 1.3 MB minimum per frame. A file that comes in under 1.3 MB at quality 95 is usually underexposed or out of focus, so treat it as a flag and check the frame before it ships.
- Format. Lossy WebP, quality 95, sRGB, with a consistent camera profile across the batch.
- Post-production. None. No cropping, retouching, sharpening, filters or AI enhancement — all of it destroys information the generation model needs.
Every frame is named SKU_N.webp: the product SKU, then a two-digit frame number that encodes the pose, the view and the camera. So 159051388_05.webp is SKU 159051388, back view, full body.
| View | Camera A — full body | Camera B or C — close-up |
|---|---|---|
| Front | SKU_01.webp | SKU_11.webp |
| 3/4 right front | SKU_02.webp | SKU_12.webp |
| Right | SKU_03.webp | SKU_13.webp |
| 3/4 right back | SKU_04.webp | SKU_14.webp |
| Back | SKU_05.webp | SKU_15.webp |
| 3/4 left back | SKU_06.webp | SKU_16.webp |
| Left | SKU_07.webp | SKU_17.webp |
| 3/4 left front | SKU_08.webp | SKU_18.webp |
| Relaxed 1 — front | SKU_09.webp | — |
| Relaxed 2 — 3/4 | SKU_10.webp | — |
Close-ups come from camera B for above-waist products and camera C for below-waist ones. The view order is identical to the full-body series, shifted by ten.
The details that fail a batch
These are the ones that get overlooked, and each of them moves colour accuracy or generation quality on its own.
- Model consistency. The same fitting model — or one with the same measurements — within a product range.
- White balance. Custom WB set from a grey card at the start of each session. Never auto.
- Exposure. Fixed manual settings, identical across the batch, shutter at 1/125 s or faster.
- Aperture. f/8 or narrower on all three cameras, so the whole garment stays sharp.
- Camera sync. All three cameras fire together on each pose, by radio trigger or tethered capture.
- Garment prep and fit. Size matched to the model so it sits visually correct, then steamed, lint-rolled and with tags tucked away.
- No extras. No accessories, jewellery or visible logos unless they are part of the product.
- Hair and face. Hair tied back and away from necklines and shoulders.
- Sharp focus. Fabric texture crisp at 100% zoom, with no motion blur.
Before you deliver
Run this over every batch before it is sent. A batch that fails any line comes back for a re-shoot, which is far more expensive than the ten minutes this takes.
- Bodysuit and background both contrast with the garment.
- Three cameras — one full body, two close-ups.
- Lenses: A at 100 mm or longer, B and C at 100 mm.
- Diffused frontal light, no shadows and no backlight.
- Model height within range — women 175–180 cm, men 180–185 cm.
- A-pose across eight views, plus the two relaxed poses.
- Garment size matches the model, with correct visual fit.
- Subject fills roughly 80% of the frame.
- Camera and model positions unchanged since the set was aligned.
- WebP lossy 95, sRGB, 5000 px or more, 1.3 MB or more.
- Fixed white balance, manual exposure, f/8 or narrower.
- No retouching, crops or filters.
- Files named
SKU_N.webp.
Everything above is upstream work, and it is the cheapest work in the pipeline. A batch shot to this standard generates cleanly the first time; a batch that drifts costs a re-shoot plus the studio day that produced it. If you are weighing that cost against what studio photography runs today, we broke the numbers down in the real cost of a fashion photoshoot — and if you are starting from packshots rather than a fitting shoot, flat lay versus on-model covers what changes.
Questions or edge cases are worth raising before the shoot, not after it.
FAQ
02- Why does fitting photography for AI need its own standard?
- Because a generation model inherits the geometry of its input. A lens change, a moved light or a camera nudged between looks all transfer into the output as inconsistency. Locking the set once and shooting the whole batch against it is what makes generations reproducible.
- What is the minimum focal length for this kind of capture?
- 100 mm on every camera. Anything wider introduces perspective distortion that warps garment proportions, and that distortion is carried straight into the generated image. 150–200 mm is preferred on the full-body camera.
- How many frames does one product need?
- Eighteen. The full-body camera shoots ten poses — eight A-pose views plus two relaxed stances — and the close-up camera shoots the same eight A-pose views of the product.
- Why deliver WebP rather than raw camera files?
- Lossy WebP at quality 95 preserves everything the generation model reads while staying a practical size to move in volume. What matters more than the container is that nothing is done to the frame afterwards — no cropping, retouching, sharpening or AI enhancement.

