Task: Investigate local image generation via ComfyUI for AI agents
Investigate local image generation via ComfyUI for AI agents
Test whether image generation can be offloaded from Codex/ChatGPT to a local ComfyUI instance, leaving task formulation, prompt engineering, and vision review to the agent, in order to reduce token consumption and freezes when working with images.
Context
When working in local Codex via subscription, stability issues are regularly observed in long chats, especially when the agent starts actively working with images and layouts.
Practical observation:
"Codex is also vibes-coded. Chats constantly freeze. If it stalls, that's it. No errors, nothing. Just a long wait, then reconnect. You start a new chat, and it goes on. Until it freezes again."
And separately:
"As soon as it starts working with pictures, it freezes really fast."
Beyond freezes, image generation/processing inside an AI session noticeably consumes context/tokens and can make iterative design work expensive and unstable.
Hypothesis
Separate reasoning/vision from actual image generation.
Leave the AI agent responsible for:
- understanding the task and UI/visual requirements;
- preparing and refining prompts;
- selecting generation parameters/variants;
- analyzing the resulting images via vision;
- comparing with requirements/references;
- forming the next iteration;
- deciding whether the result is sufficient.
Perform the heavy generation itself locally via ComfyUI and suitable local image models/workflows.
Proposed cycle:
requirements / current UI
↓
AI agent
↓
prompt + generation parameters
↓
local ComfyUI
↓
image(s)
↓
AI vision review
↓
prompt / parameter correction
↺
In other words, use the expensive multimodal model primarily where it is most useful: for task formulation, reasoning, and result verification, while offloading mass pixel generation to a specialized local pipeline.
What to Investigate
- Set up a reproducible local ComfyUI instance.
- Select a minimal set of models/workflows for UI mockups, illustrations, and other typical visual development tasks.
- Define a simple programmatic interface through which the coding agent can run workflows and retrieve results without manual work in the ComfyUI UI.
- Test whether the agent can independently execute the
prompt → generate → inspect → adjust → generatecycle. - Test working with multiple candidates and selecting the best result.
- Determine which intermediate data actually needs to be returned to the AI context and what can be stored locally to avoid bloating the session.
- Investigate whether the image-generation workflow can be isolated from the main coding chat so that visual iterations do not degrade the stability of the primary context.
- Document real failure modes: ComfyUI freezes, VRAM/RAM shortages, long generation jobs, incorrect prompts, vision-review issues, file accumulation, etc.
- (Note: missing 9th item in original source numbering skip or just standard list)
Compare with Current Workflow
Run several identical visual tasks using two approaches:
A. Current Approach
AI/Codex/ChatGPT handles image generation independently within its multimodal session.
B. Local Pipeline
AI creates the assignment and verifies the result, while generation is performed by ComfyUI.
For each scenario, track:
- wall-clock time to an acceptable result;
- number of iterations;
- token/context consumption, to the extent it can be measured;
- stability of the main AI session;
- number of freezes/reconnects/restarts;
- quality of the final result;
- amount of manual human intervention;
- local GPU/VRAM/RAM costs;
- how easily the agent can independently diagnose a failed generation.
Success Criteria
Obtain a working experimental pipeline where a coding agent can invoke local ComfyUI, receive an image, visually inspect it independently, and trigger the next iteration if necessary.
After several real-world tasks, answer the following questions:
- Does this reduce AI context/token consumption?
- Does the coding chat become more stable?
- Does the full cycle to an acceptable result speed up or slow down?
- Is the quality of local models sufficient for real product work?
- Which types of visual tasks are rational to perform locally, and which are more advantageous to leave to an external multimodal/image model?
- Is it worth establishing this pipeline as a permanent tool/skill for AI agents?
Important Limitation
Do not assume it is pre-proven that Codex freezes are specifically caused by token consumption or model-side image generation. This is an observed correlation that the experiment should help separate from other causes. The goal is to test the practical effectiveness of the alternative workflow, not to pre-explain the internal cause of Codex freezes.