RESEARCH NOTE 000 2026.09.22
Testing prompt improvements through images.
As we make Daily Prompts, we record what changed and what happened. That includes useful results and comparisons that leave questions open.
This introductory note examines two existing production records. Display dates are not generation dates.
Browse Daily Prompts ↗Changed condition · September 8 entry
Define when the image should have no text.
An earlier Gemini browser test using the same fictional bookshop image. Look at the lower-right whitespace. Open either image to inspect it in full.


- What changed
- We added a condition specifying zero text information when the input contains no lettering.
- What we observed
- For this input, the three invented lines in the earlier output did not appear in the later one. Composition and rendering also changed.
- What remains unknown
- This was a small historical test with one input, not an established causal effect. It does not establish reliability for our current internal generator or other prompt languages.
Same-prompt rerun · September 16 entry
A better result does not always mean a better prompt.
The exact same prompt was submitted twice to Codex internal image generation. The saved review judged one city heading in the first result to be misspelled.


- What changed
- The prompt was unchanged. The record explicitly identifies identical submitted text.
- What we observed
- The second image was selected. The images and lettering varied despite identical instructions.
- What remains unknown
- We do not count this as a prompt-structure improvement. Two images cannot establish the amount of variation or a success rate.
METHOD
Learn within ordinary production.
- Define checks for subject, composition, text and input preservation before generation.
- Use at most two candidates per day and two generation calls per candidate, counting failures. Stop early when sufficient.
- When revising, change one structural element and distinguish that comparison from an unchanged-prompt rerun.
- Check relevant later work and retain counterexamples. One successful image does not justify a general rule.
RECORD AUDIT
Start with what the records actually contain.
We inspected 16 entries with internal-generation manifests among September 13–29. This production inventory includes future display dates; it is not a chronological quality trend.
The records contain 47 successful images and one invocation with no output: at least 48 calls. One candidate exceeded the call limit. September 30 is excluded because its ledger references one missing generation manifest.
Selection is not a quality pass. Standardized ratings are incomplete, so quality improvement and first-pass success rates are not calculated. Monetary cost is unknown.
Counts and image checksums ↗The next question
Does the no-text condition reduce unwanted lettering across different subjects? We will record relevant ordinary production results, including regressions and no change.
No extra images are generated for research. We prepare another note when roughly ten newly evaluable candidates provide new findings.