Generate at or near the resolution the model was trained for, then enlarge afterward if you need something bigger. Going straight to a very large size usually gives you a worse picture, not a more detailed one, and it costs more to produce. This is the opposite of how cameras work, which is why it catches people out.
If you want the short version: pick the aspect ratio you need, take the offered size closest to the model's native resolution, and upscale as a second step.
Why does generating larger not give more detail?
An image model learned what a picture looks like at a particular size. Ask for something much larger and it does not add finer detail, it repeats the composition it knows. You get two horizons, a doubled subject, an extra limb, a face that appears twice at different scales.
The reason is that the model has no concept of the whole frame at a size it has never seen. It works on local neighborhoods of pixels, and at a much larger canvas those neighborhoods each look like a complete scene to it. So it draws a complete scene in each of them.
This is the single most common cause of the duplicated-subject artifact people blame on their prompt. No rewording fixes it, because the prompt was never the problem.
Models differ in where that limit sits, and newer ones are trained across multiple resolutions and tolerate more range. The principle holds even so: there is a size this model is comfortable at, and far past it quality falls off rather than improving.
What is a model's native resolution?
Native resolution is the size, or set of sizes, a model was trained on. Earlier open image models were commonly trained around 512x512. Later ones moved to roughly 1024x1024, and several modern models train across a set of aspect ratios at a similar total pixel count.
That jump was not free. The SDXL paper by Podell and colleagues describes the model as using "a three times larger UNet backbone" than previous versions of Stable Diffusion, trained across multiple aspect ratios. Supporting a higher native resolution meant a substantially bigger model, which is the clearest evidence that resolution is a property a model is built for rather than a number you can raise from the outside.
What matters is not the exact number but the total pixel budget. A model comfortable at 1024x1024 is usually comfortable at 1152x896 or 896x1152, because those are about the same number of pixels rearranged. It is far less comfortable at 2048x2048, which is four times the area.
If a tool offers you a fixed list of sizes, that list usually reflects what its models handle well. A form offering 512x512, 768x768, 1024x1024, 1024x768, and 768x1024 is offering one pixel budget in a few shapes plus a smaller option, and staying inside it is the safe default.
Does aspect ratio change what you get?
Yes, and more than people expect. Aspect ratio changes composition as well as crop. A tall frame pushes the model toward a subject filling it, a portrait, a standing figure, a tree. A wide frame pushes it toward context: landscape, room, horizon.
That means aspect ratio is a compositional instruction as much as a technical setting. If you keep getting a close portrait when you wanted a scene, switching from square to landscape often does more than adding "wide shot" to the prompt.
It also means you should choose the ratio you need before iterating on the prompt. Getting a prompt right at square and then switching to landscape gives you a different picture, and some of your prompt work will not carry over.
When should you upscale instead?
Almost always, if you need something big. Generate at a size the model handles, then enlarge with a tool built for it. Upscalers are trained specifically to add plausible detail while enlarging, which is a narrower and much better-defined job than generating a coherent scene at an unfamiliar size.
The practical sequence:
Pick your aspect ratio.
Generate at the offered size closest to the model's native budget.
Generate several and choose one you are happy with at that size.
Upscale the chosen image to the size you need.
Doing the choosing before the upscaling matters. Upscaling is a separate step with its own cost and time, and running it on images you are going to discard wastes both.
For print, work backward from the physical size and the printer's requirements rather than picking a pixel count that sounds large. For screens, you almost never need what people assume: a full-width web hero is smaller than most people generate.
Why do services cap the size?
Because generation is memory-bound, and the ceiling is a property of the machine rather than a policy choice. Memory use scales with the number of pixels, and past a point the job does not run slower, it fails.
This is why a service will offer a fixed menu of sizes and refuse anything above it, and why the cap can differ between two services running the same model. They are telling you about their hardware.
It is also why generating large is more expensive wherever you do it. If you are running locally, you feel it as your own memory limit. If you are using a hosted service, you feel it as price or as a refusal.
How should you evaluate resolution support?
If image size matters to your work, these are the questions that decide whether a service fits:
What sizes are offered, and which aspect ratios? A service offering only square is a poor fit for banners or posters.
Is there a hard pixel ceiling, and does the interface tell you before you generate? Finding out through a failed job is a bad experience and, depending on billing, may still cost you.
Does it name the model? Without knowing the model, you cannot reason about its native resolution.
Is upscaling available, or are you exporting to another tool? A workflow that leaves the service is fine, it just needs to be a decision rather than a surprise.
How are large generations billed? Per image regardless of size, or scaled with pixels? This changes how you experiment.
What to do when a picture comes out doubled
If you get a duplicated subject or a second horizon, work through this before touching the prompt:
Drop to a smaller offered size and regenerate. If it resolves, the issue was resolution.
Change the aspect ratio toward the shape the composition wants.
Generate a batch. Some seeds produce the artifact and others do not at the same settings.
Only then adjust the prompt.
Most people do these in reverse order and spend the afternoon rewriting a prompt that was fine.
The tradeoff nobody mentions
Working at a smaller generation size makes you faster in a way that outweighs the resolution you gave up. You see more variations per unit of time and money, so you find the good composition sooner. Then you spend the expensive step once, on the picture you already know you want.
Generating everything at maximum size from the start feels thorough. In practice it means fewer attempts, slower feedback, and a higher bill for the same result.
Where Carpathian fits
Carpathian AI offers image generation at 512x512, 768x768, 1024x1024, and the portrait and landscape variants 768x1024 and 1024x768, which is one pixel budget in the shapes most work needs plus a faster small option. The size menu is checked against what the serving machine can hold, so a size it cannot run is refused in the form rather than after you wait for a failed job. Generation needs a free account, and each account carries an allowance before anything is billed.
What Resolution Should I Generate AI Images At? | Carpathian