How to use and evaluate “Candid Urban Fashion Portrait on a European Pedestrian Bridge”
This couple prompt is designed as an editable starting point, not a guaranteed one-click result. The notes below translate its JSON into a readable visual plan, identify the controls that matter, and show how to diagnose a weak generation without simply adding more text.
Prompt-specific blueprint
Scene and atmosphere
- Location
- A pedestrian bridge in the heart of a grand European-style city, with an elegant historic cityscape in the background. Large ornate stone buildings, classical façades, tall windows, decorative architectural details, and a busy street below are all clearly visible. Several city buses and cars pass…
- Time
- Late afternoon
- Mood
- Refined, candid, metropolitan, calm, confident, natural, and authentic.
- Subject Action
- The woman stands casually in the center of the pedestrian bridge with both arms extended outward, lightly resting her hands on the dark metal railing. Her legs are positioned naturally with one slightly crossed in front of the other, creating an effortless street-style pose.
- Environment Detail
- The setting should feel like a real urban travel moment in a grand European city. Preserve the sense of scale from the architecture and the street activity below, with moving buses and cars adding authentic city life and realism.
Visual treatment
- Photography
- Ultra-realistic natural smartphone photography with a candid urban fashion portrait feel. The image must look like a real smartphone photo captured in natural conditions, not a studio fashion shoot, CGI render, illustration, or overproduced commercial campaign.
- Camera
- Vertical 3:4 smartphone composition with a natural handheld feel, slight smartphone lens softness, realistic perspective, and subtle depth of field.
- Camera Perspective
- A natural eye-level to slightly elevated candid portrait perspective appropriate for a smartphone travel-fashion photo, maintaining realistic spatial relationships and authentic proportions.
- Lighting
- Soft late-afternoon daylight creates gentle warm highlights on the buildings and her hair. The sky is pale and softly illuminated, giving the scene a refined metropolitan atmosphere without becoming overly dramatic.
- Rendering
- RAW smartphone photography aesthetic with realistic skin pores, individual hair strands, accurate hands and fingers, natural body proportions, detailed denim and leather textures, realistic architectural detail, subtle depth of field, natural exposure, and realistic background motion.
- Skin Rendering
- Preserve authentic human skin texture with pores, natural tonal variation, and realistic detail. No skin smoothing, no plastic skin, no over-retouching, and no artificial glow.
Composition and wardrobe
- People Count Label
- one uploaded reference person
- People Count
- single
- Reference Instruction
- Use one uploaded reference person as the sole main subject and exclusive identity source.
- Framing
- Vertical 3:4 candid urban fashion portrait with the woman standing centrally on the pedestrian bridge, clearly visible from head to boots, while still allowing the historic European cityscape and the busy street below to remain an important part of the frame.
- Wardrobe · Top
- A cropped dark charcoal denim jacket with visible stitching, silver buttons, and a structured relaxed fit.
- Wardrobe · Bottom
- A short flowing black pleated skirt.
- Wardrobe · Footwear
- Tall black knee-high leather boots with a simple pointed-toe silhouette.
Recommended workflow
- 01
Prepare the input
Choose reference images you have permission to use. For Candid Urban Fashion Portrait on a European Pedestrian Bridge, prioritize a clear subject and avoid unrelated people or private details in the frame.
- 02
Lock the visual plan
Confirm subject count, framing, scene, and lighting before adding finishing language. Resolve conflicts such as close-up versus full-body composition.
- 03
Adapt for one model
Start with ChatGPT Image. Keep a baseline generation and document the exact model version, references, aspect ratio, and visible settings.
- 04
Revise from evidence
Name the visible failure, change one instruction block, and compare it against the baseline. Keep only revisions that improve the intended criterion.
Model adaptation notes
ChatGPT Image
Keep the JSON structure, attach references in the same conversation, and use follow-up edits for one local correction at a time.
Gemini
Provide the JSON with clearly labeled reference images. Restate any ignored constraint in plain language without duplicating the whole prompt.
A compatibility label means the visual intent can be adapted; it does not promise identical output across providers, versions, or settings.
Prompt-specific quality checklist
- Preserve every important visual and narrative element from the original prompt without simplifying the scene.
- Keep the subject clearly recognizable as the same real person from the reference photograph.
- The generated face should match the reference photograph extremely closely, including very small facial details.
- Do not beautify, glamourize, or idealize the person beyond the reference.
- Do not make the person younger than in the reference photograph.
- Do not intentionally make the body sexier or more muscular if the original prompt does not explicitly request it.
- If the reference person is slim, keep her slim; if she is fuller, keep her fuller; if she is fit, keep her fit.
Failures to check before publishing
- pasted face
- copied face pose
- frozen reference expression
- duplicated head angle
- copied reference gaze
- mismatched gaze direction
- reference pose lock
Why the editable controls matter
Gender
Required identity direction for the reference subject.
People count
Default is a single uploaded reference person.
Identity or subject drift
Reduce competing style language, improve the reference, and separate stable traits from pose, gaze, and expression.
Composition breaks
Check whether crop, shot distance, aspect ratio, subject count, and required objects can all be satisfied in one frame.
Artificial lighting or texture
Name a motivated light source and believable material behavior before adding grain, grading, sharpness, or resolution terms.







Comments
Share the result you got, attach an image, and reply to other people's results.
You need to sign in to comment.