Model-neutral workflow

How to turn an image into a text prompt

A useful prompt is not a caption and not a pile of style words. It is an editable record of the visual decisions that make a reference image work.

Reference image moving through visual analysis into a structured text prompt
The translation works in three stages: observe the reference, decide which visual facts matter, then arrange those facts as instructions a new image can follow.

First, separate visual prompting from OCR.

“Image to text” can describe two different tasks. Optical character recognition extracts letters already printed inside an image: a receipt total, a sign, a label, or a paragraph. Image-to-prompt writing does something else. It translates the scene itself into instructions for generating, recreating, or editing an image.

If a photograph contains a red umbrella in a concrete corridor, OCR may return no text at all. A visual prompt should capture the umbrella, the corridor, the low camera angle, the wet reflections, the blue-hour light, and the visual tension created by one warm color inside a cool scene.

Use OCR when the words inside the image are the output. Use an image-to-text prompt when the appearance, composition, atmosphere, or design language is the output.

Decide what the prompt must do.

The same reference can produce several correct prompts because the intended job changes what deserves emphasis. Before describing anything, choose one of four outcomes.

Recreate

Preserve the subject, framing, lighting, palette, materials, and atmosphere as closely as practical.

Restyle

Keep content and composition, but replace the medium, rendering treatment, color behavior, or finish.

Edit

State the exact change and the elements that must remain stable. Local edits need a preserve list, not a new scene description.

Borrow a system

Reuse only selected traits such as lighting, typography space, product staging, or camera language for a new subject.

This decision prevents a common failure: copying every visible detail when only the composition matters, or writing a vague style label when the user actually needs a close recreation.

Use a six-pass visual reading.

Read the image from stable facts to interpretive details. The order below keeps the prompt coherent and makes later editing easier.

  1. Subject and action. Name the primary subject first. Add count, age group, form, clothing, pose, expression, or movement only when visually important.
  2. Setting and relationships. Describe where the subject is, what surrounds it, and which objects affect the story or balance of the frame.
  3. Composition and camera. Record shot size, viewpoint, orientation, subject placement, depth, crop, focal plane, and negative space.
  4. Light and color. Identify the light source, softness, direction, contrast, time of day, dominant hues, accent colors, and whether shadows are open or dense.
  5. Material and style. Name surfaces, texture, medium, rendering treatment, grain, edge quality, and final-use context. The image prompt style guide explains how to replace vague labels with visible traits.
  6. Constraints. State what must stay, what may change, and what should be absent. Use negative language sparingly and only for failures you can anticipate.
Subject

translucent ribbed-glass vessel with a narrow neck

Setting

frosted white studio surface with an empty editorial backdrop

Frame

low three-quarter camera angle, close product crop, generous space above

Light

large soft source from the left, long feathered shadow, restrained highlights

Finish

quiet luxury product photography, tactile glass detail, clean color separation

For a reusable order and optional fields, continue with what a good image prompt should include.

Turn observations into one working prompt.

Suppose the reference shows a single person under a red umbrella in a brutalist corridor at blue hour. A weak caption identifies the objects but discards the visual logic. A useful prompt connects those objects to framing, light, surface, and mood.

Caption

A person holding a red umbrella in a concrete hallway at night.

Working prompt

cinematic full-body photograph of a solitary figure holding a vivid red umbrella in a monumental raw-concrete corridor, centered one-point perspective, low eye-level camera, blue-hour ambient light, wet pavement reflecting cool cyan tones, repeating rectangular openings fading into haze, restrained palette with one red accent, realistic concrete texture, quiet suspense, wide 16:9 composition

The longer version is valuable because each phrase controls a visible dimension. Remove “quiet suspense” and the geometry still survives. Remove “centered one-point perspective” and the entire image structure can change.

A template you can edit

[medium or image type] of [primary subject + defining attributes], [action or pose], in [setting + important objects], [shot size + viewpoint + subject placement], [light source + contrast], [dominant palette + accent], [materials + surface detail], [style or production context], [mood], [constraints or aspect ratio]

Use the template as a diagnostic tool, not a form that must be filled completely. If a field does not change the picture, leave it out.

Adapt the prompt only after the visual reading is sound.

A model-neutral description gives you a stable source of truth. From there, adjust syntax and control language for the destination rather than rewriting the visual analysis from scratch.

  • Midjourney: keep the visual phrase compact, place parameters at the end, and distinguish an image prompt from a style reference. See the Midjourney reference workflow.
  • FLUX: favor direct positive descriptions and explicit reference roles. Avoid long lists of contradictory adjectives.
  • GPT Image or conversational editors: write the desired result as a clear production brief, then state what must remain unchanged.
  • JSON workflows: separate reusable fields such as subject, composition, lighting, and constraints without assuming the model requires JSON. See the image-to-JSON prompt guide.

Model syntax cannot repair a weak observation. Get the subject, frame, light, and material relationships right before adding parameters or structured fields.

Remove details that sound precise but do no work.

Adjective stacking

“Beautiful, stunning, epic, amazing” communicates approval, not visual direction. Replace it with light, composition, material, or mood evidence.

Contradictory camera cues

A close-up, full-body view, aerial angle, and eye-level portrait cannot all define the same frame. Choose one hierarchy.

Copying incidental defects

Compression noise, accidental clutter, or a cropped edge may not belong in the desired result. Preserve decisions, not every artifact.

Mixing content and edit instructions

For edits, say what changes and what stays. Re-describing the whole image can invite unnecessary drift.

Check the prompt before you use it.

  • Can a reader identify the main subject in the first clause?
  • Does the prompt explain the frame rather than only naming objects?
  • Are light and color described through observable behavior?
  • Does each style word have a visible meaning?
  • Are the details consistent with one another?
  • For an edit, is the change separate from the preserve list?
  • Could you remove any phrase without changing the likely result? If yes, remove it.

When you want more comparisons, study the annotated AI image prompt examples. When you are ready to work from your own reference, generate a prompt from your image and edit it using this checklist.

Common questions

Should an image prompt describe every visible object?

No. Describe the objects that establish identity, composition, scale, story, or visual balance. Incidental details can distract from the main instruction.

How long should an image-to-text prompt be?

Long enough to preserve the visual decisions you care about, and no longer. A simple product image may need one dense sentence; a controlled scene or edit may need sections for content, composition, style, and constraints.

Can one prompt work in every image model?

The visual description can travel well, but parameters and reference controls differ. Keep a model-neutral source prompt, then create small model-specific versions.

Apply the method to your own reference.

Generate a draft, then keep only the visual details that serve your intended recreation, edit, or style study.

Try the image-to-text prompt generator