Roger

Bitcoin · Macro · AI · Freedom Tech
← All writing

Latest

The 17.3-Gigabyte Question

Two image models, one graphics card: Qwen-Image-2.1 against Z-Image Turbo, 30 prompts each. Qwen needed 17.3 GB of weights on a 12 GB card and was 2.65 times slower — but the interesting part is where the time actually goes, and which model wins on sharpness versus tone.

9 Oct 2026 2,133 words · 9 min Also on Nostr as a long-form note
The 17.3-Gigabyte Question

Same prompts, same seed, same graphics card. One model needed 17.3 GB of weights on a card with 12. It finished the job, and it finished it differently.


On 9 October 2026, two image models ran the same thirty prompts through the same graphics card. One was Z-Image Turbo, a lean model that fits comfortably into 12 GB of VRAM. The other was Qwen-Image-2.1, whose three weight files together measure 17.29 GB, five gigabytes more than the card has.

The lean model averaged 14.3 seconds per image. The large one averaged 37.8 — 2.65 times slower.

That number is almost the least interesting thing in this comparison. Because per sampling step, the large model was only about a third slower. The rest of the gap comes from somewhere else entirely, and where it comes from changes which model you should actually use.

What was measured, and how

Thirty prompts, six groups of five: Bitcoin, people, nature, vehicles, futuristic scenes, architecture. Each prompt went through both models unchanged. Same seed (777), same resolution (1024×1024), same sampler (euler, simple), same CFG of 1.0. No style words like "photorealistic" were added, since those would measure how each model reacts to style words rather than the models themselves.

Qwen ran 25 steps, the value in its own official workflow. Z-Image ran 8, the value in its own. Forcing one to the other's step count would be a different experiment, and a less honest one: these are the settings each model ships with.

The hardware: an Intel i7-6700K with 62 GB of RAM and an NVIDIA RTX 3060 with 12 GB of VRAM, driver 595.71.05. ComfyUI at main commit 08ff3c11b3a8, file version 0.39.0.

Memory was released before every single run, so both models started from the same state. 60 runs, 60 successes, no failures.

The measurement problem nobody mentions

ComfyUI returns no timestamps in its execution history. Every field I checked came back None. The obvious approach, starting a stopwatch and waiting for the image, would have lumped two completely different things into one number: the time spent moving 17 GB of weights from disk into memory, and the time spent actually computing an image.

Those are not the same cost, and they do not scale the same way. For a comparison like this one, the split is the finding. So the measurement reads ComfyUI's WebSocket alongside the job and separates three phases: loading (job start to first sampler event), sampling (first to last sampler event), and the remainder (VAE decode plus writing the file).

The three add up to the total in every run. For Qwen: 7.5 + 27.3 + 2.8 = 37.6 against a measured mean of 37.8. The two-tenths difference comes from averaging the phases separately, since the phases are not equally populated in every run. For Z-Image: 5.6 + 6.7 + 1.8 = 14.1 against 14.3.

One caveat belongs in the text rather than a footnote: the clock runs on the machine issuing the job, not on the GPU, so network and queue time are inside these numbers. Over a local network against a warm server that distance is small, but it is not zero.

Where the time actually goes

Time per image, split into loading, sampling and decode
Time per image, split into loading, sampling and decode
                       Qwen-Image-2.1      Z-Image Turbo
steps                  25                  8
mean total             37.8 s              14.3 s
  loading              7.5 s               5.6 s
  sampling             27.3 s              6.7 s
  decode + save        2.8 s               1.8 s
seconds per step       1.09 s              0.84 s
Seconds per sampling step
Seconds per sampling step

Read the last line first. Per sampling step, Qwen is 30 percent slower than Z-Image. Not 165 percent slower. Thirty percent.

Multiply by steps and the gap widens: 25 × 1.09 = 27.3 seconds of sampling against 8 × 0.84 = 6.7. The factor of four lives entirely in the step count.

Then watch the total collapse back to 2.65. Roughly 7.5 seconds of Qwen's run and 5.6 seconds of Z-Image's run go to work that does not scale with steps at all: moving weights off the disk. Z-Image's total is half fixed cost. Qwen's is a fifth. The larger model has more absolute overhead and a smaller proportion of it.

There is a second pattern in the Qwen numbers that only appears because the phases were separated. Its sampling time is almost perfectly constant across all 30 runs: 27.00, 27.05, 27.11 seconds in the first three, a spread of 0.11 seconds. Its loading time is not, ranging from 7.46 to 8.29 seconds. The variability in Qwen's total time comes from the memory system, not from the computation.

That is what you would expect from a model that does not fit. Z-Image's weights stay resident. Qwen's get paged in, partly evicted, and paged back. How long that takes depends on what else the kernel is doing with the same page cache.

Sharper and louder

Speed is the easy half. The harder question is whether the extra 23.5 seconds buy anything.

To answer it I measured every one of the 60 images the same way: Laplacian variance for sharpness, mean absolute Laplacian for high-frequency noise, and the share of near-black and near-white pixels for blown-out regions.

Sharpness has a trap. Measured across the whole frame, Laplacian variance also measures how large the subject is, so a close-up of a coin will always outscore a landscape. Sharpness was therefore measured on an identical 512-pixel crop from the centre of every image, and that is the number quoted below.

                       Qwen-Image-2.1      Z-Image Turbo
sharpness (centre)     1243                1101
noise                  17.0                10.6
contrast               48.0                51.9
near-black pixels      2.66 %              0.95 %
near-white pixels      0.07 %              0.03 %

Qwen produces the sharper image in 20 of 30 subjects. Z-Image wins the other ten. That is a real advantage, but not a rout, and the pattern behind it is more useful than the count.

Qwen's wins cluster in subjects with manufactured detail: the Bitcoin coin (6961 against 2448, nearly threefold), the classic car, the bicycle, the regional train. Z-Image's wins cluster in textured, organic or high-frequency subjects: breaking waves on volcanic rock (5170 against 1896), the station hall's steel roof, the library interior.

Noise runs the other way, and decisively. Z-Image is the quieter image in 28 of 30 subjects. Qwen's noise figure is 60 percent higher on average, and its worst case is worse: up to 28 percent near-black pixels in a single image, against Z-Image's worst of 7.7 percent. In dark scenes, Qwen is more likely to produce blocked-up shadows that have to be rescued in post.

The honest summary is therefore not that the bigger model is better. It is that the bigger model resolves structure better and renders tone worse, and which of those matters depends on what the image is for.

Sharpness against noise, one point per generation
Sharpness against noise, one point per generation

What the numbers cannot see

A measurement suite tells you what it was built to measure, and it is worth stating plainly what mine does not.

Looking at all 30 pairs with my own eyes, Qwen produced two images where the head of the subject was cropped out of frame entirely: the craftsman holding a hammer, and the humanoid robot. Both prompts had asked for a full figure. Z-Image framed both correctly.

Every pair is shown below. Top row is Z-Image Turbo, bottom row Qwen-Image-2.1. Same prompt, same seed in each column.

Bitcoin. Coin, miner, chart, hardware wallet, node.

Bitcoin: top Z-Image, bottom Qwen
Bitcoin: top Z-Image, bottom Qwen

People. Elderly man, young woman, craftsman, child, group.

People: top Z-Image, bottom Qwen
People: top Z-Image, bottom Qwen

Nature. Fog forest, summit, surf, dunes, red fox.

Nature: top Z-Image, bottom Qwen
Nature: top Z-Image, bottom Qwen

Vehicles. Classic car, motorcycle, lorry, regional train, bicycle.

Vehicles: top Z-Image, bottom Qwen
Vehicles: top Z-Image, bottom Qwen

Futuristic. Future city, robot, cockpit, portal, cyberpunk alley.

Futuristic: top Z-Image, bottom Qwen
Futuristic: top Z-Image, bottom Qwen

Architecture. Church, glass house, bridge, library, station hall.

Architecture: top Z-Image, bottom Qwen
Architecture: top Z-Image, bottom Qwen

No Laplacian variance catches that. A headless craftsman can be perfectly sharp. Any evaluation reporting only sharpness and noise would have scored those two images as Qwen wins, and they are not wins but failures.

There is also a text problem in both models, and it is not equal. Both put unreadable lettering on the rim of the Bitcoin coin. Both render neon signage in the cyberpunk alley as plausible-looking nonsense. Neither can be trusted to spell anything, and the standard mitigation, keeping text out of the prompt, costs you subjects that legitimately contain writing.

Then the deepest caveat: this comparison holds the sampler, seed and resolution fixed, which is the right way to isolate the models and the wrong way to answer "which produces better images in practice." In practice you would tune each model, and a 25-step setting that suits Qwen may not be the setting that suits it best at 12 GB.

The trap: an architecture string

Getting Qwen-Image-2.1 to run at all took longer than the comparison did, and the reason is worth recording because it will catch the next person.

The efficient way to run a large model on a small card is GGUF, a quantised format that shrinks weights substantially. A community GGUF build of Qwen-Image-2.1 was already on the machine, 4.3 GB instead of 7.26.

It did not load. The error was Unknown architecture: 'qwen_image21'.

Reading the file's own header explains it. The GGUF declares general.architecture = 'qwen_image21'. The GGUF loader in ComfyUI keeps a list of architecture names it recognises, and that list contains qwen_image, without the "21". The loader is not broken — it simply predates the 2.1 architecture. The newest release of that loader, checked the same day, still does not list it.

The fix was to abandon GGUF for this model and use the safetensors build the official workflow ships with: 17.29 GB instead of 4.3, which is what forced the offloading and produced the timing pattern this article is about. A cheaper file would have been faster. It was also unusable.

One detail worth knowing if you go down this road: the tensor names in a safetensors file do not appear in file order. Layer 19 sits at byte 4.58 billion; layer 2 sits at 4.77 billion. Anyone trying to infer download progress from which tensor name fails to load is reading noise. I did exactly that, briefly, before checking the offsets and correcting myself.

Where this leaves you

The practical answer divides cleanly, and it divides on something other than quality.

Z-Image Turbo costs 14.3 seconds and produces a quieter, more conventionally attractive image. It fits in the card. It does not thrash memory. Its worst case is better than Qwen's worst case. For iteration, for trying six framings of a product shot before lunch, it is not close.

Qwen-Image-2.1 costs 37.8 seconds and resolves manufactured detail that Z-Image smears. On the Bitcoin coin it produced nearly three times the edge contrast. On text-adjacent, mechanical and architectural subjects, that is visible. It also occasionally forgets to include the head.

The steel-man for the opposite view is real and I want to state it properly. If you are running this comparison on a card that fits both models, 24 GB or more, the entire offloading penalty disappears. Qwen's loading time would drop toward Z-Image's, its total would fall closer to its 27.3 seconds of pure sampling, and the 2.65× ratio would compress toward the 4× step-count factor with a much smaller fixed cost underneath. The gap measured here is partly a property of a 12 GB card, not of the model.

And there is a second honest position: for most people most of the time, neither model's output needs to be examined at this resolution. If you are generating a mood image for a post, 14.3 seconds of Z-Image is not worse than 37.8 seconds of Qwen. It is just faster, and the difference is invisible at the size it will be seen.

The number that should change your plan

Here is the arithmetic that matters, and it is arithmetic rather than a prediction.

Qwen spends 7.5 of its 37.8 seconds, one image in five, moving weights it cannot keep resident. That cost does not fall when you lower the step count, and it does not fall when you shrink the image. It only falls when the weights fit.

The right way to read the 2.65× is therefore not as a statement about Qwen being slow. It is as a statement about a mismatch: 17.29 GB of model on 12 GB of memory — paying a toll on every single image for a constraint that a larger card would remove entirely.

Until then, the honest description of this pair is not faster versus better. One model fits and the other does not. The one that does not is worth the wait exactly when the subject has edges a smaller model cannot hold.