← Back to blog

Why Run ComfyUI Locally for Character Cards

ComfyUI turns a local GPU into a portrait and sceneart factory for SillyTavern. Instead of burning cloud credits per image or fighting inconsistent outputs…

Published
  • comfyui
  • image-generation
  • local-llm
  • character-cards
  • sillytavern

SillyTavern itself doesn’t generate images — it consumes them. ComfyUI sits upstream as the generator, and the character card is the delivery format.

  • Cost: unlimited generations after hardware purchase, no per-image API fees.
  • Consistency: IPAdapter and LoRA let you lock a face across dozens of expressions.
  • Privacy: prompts and outputs never leave your machine, which matters for original characters.
  • Integration: ComfyUI exposes an API endpoint that SillyTavern’s image-generation extension can call directly.

The tradeoff is setup time and VRAM. A 12 GB card handles SDXL comfortably; 8 GB works with quantized checkpoints and tiled upscaling.

Hardware and Software Baseline

ComponentMinimumComfortable
GPU VRAM8 GB12–16 GB
RAM16 GB32 GB
Storage60 GB SSD200 GB NVMe
Python3.10–3.123.11

Install ComfyUI via the portable build or a git clone with a virtual environment. Add ComfyUI Manager first — it resolves custom nodes without manual dependency hunting. You will need:

  • ComfyUI_IPAdapter_plus for reference-image conditioning
  • ComfyUI-Impact-Pack for face detailing and masking
  • rgthree-comfy for cleaner graph routing

Choosing Models for Portrait Work

Two families dominate local character work in 2026:

  • SDXL-derived checkpoints — best balance of quality, speed, and LoRA ecosystem. Use a photoreal or semi-realistic finetune for portraits.
  • Flux-based models — superior prompt adherence and text rendering, heavier VRAM cost. Worth it for detailed scene art with props and signage.

Layer a face-refiner pass (a second sampler run at low denoise on a masked face region) rather than relying on the base model alone. It fixes the small-eye and asymmetric-jaw artifacts that break character consistency.

The Core Portrait Workflow

Build this graph once and save it as a template.

  1. Load Checkpoint → base model.
  2. CLIP Text Encode → positive prompt with a fixed character descriptor block.
  3. CLIP Text Encode → negative prompt (anatomy errors, watermark, extra limbs).
  4. Empty Latent Image → 832×1216 for a portrait orientation.
  5. KSampler → 28–32 steps, DPM++ 2M Karras or the model’s recommended sampler, CFG 5–7.
  6. VAE Decode → image.
  7. FaceDetailer (Impact Pack) → refine the face region.
  8. Save Image with a prefix like iris_portrait_.

Keep the positive prompt’s character block identical across every generation. Only swap the expression, pose, and lighting tokens. That single discipline does more for consistency than any node.

The Fixed Descriptor Block

Write one paragraph describing the character and paste it verbatim into every prompt:

1girl, silver bob cut, amber eyes, freckles across nose,
leather aviator jacket, painter's smock, warm rim lighting,
upper body portrait, sharp focus, detailed skin texture

Then append variable tokens: smiling, arms crossed, looking away, rain-soaked street background.

Locking Identity with IPAdapter

Text prompts drift. IPAdapter anchors the face to a reference image.

  • Generate one strong “hero” portrait first.
  • Feed it into IPAdapter Unified Loader with a face-focused preset.
  • Weight around 0.6–0.8; higher values flatten pose variety.
  • Combine with a low-weight character LoRA (0.4–0.6) if you have trained one.

For a full expression sheet, run the same graph across a batch of emotion tokens with IPAdapter active. You get a coherent set instead of twelve unrelated people.

Scene Art and Backgrounds

Scene art follows the same graph with three changes:

  • Switch to landscape latents (1344×768).
  • Drop IPAdapter weight to 0.3–0.4, or disable it if the character isn’t in frame.
  • Add environment tokens and a style anchor (cinematic, moody, golden hour).

Generate backgrounds separately from character sprites. Compositing them later in an image editor gives you far more control than prompting both at once.

Wiring ComfyUI into SillyTavern

SillyTavern’s image-generation extension supports a ComfyUI backend.

  1. Enable API mode in ComfyUI (--listen if SillyTavern runs on another device).
  2. Export your workflow in API format from the ComfyUI menu.
  3. In SillyTavern, select ComfyUI as the image source and paste the API JSON.
  4. Map the prompt node and seed node so SillyTavern can inject text per message.

Now character cards can carry portrait references, and scene prompts can fire inline during a chat.

Practical Tips

  • Seed locking: fix the seed while iterating on a prompt, then free it for final batches.
  • Batch size 4, pick 1: faster than tuning a single image to perfection.
  • Upscale last: run a 1.5× latent upscale, then a detail pass. Don’t upscale before FaceDetailer.
  • Name files systematically: character_expression_seed.png keeps large card libraries navigable.

For a ready-made example of the output quality this pipeline targets, see the Iris the Pixel Painter card — a character built around exactly this kind of locally generated portrait and scene set.

Conclusion

A local ComfyUI pipeline gives you unlimited, consistent, private art for every character card you run in SillyTavern. The setup cost is a weekend; the payoff is a reusable asset library you fully own. Once your graph is stable, browse the MiniTavern character card market for new cards to feed it, and manage your library on the go with the MiniTavern iOS and Android apps or the MiniTavern Chrome extension.

More guides you might like

How World Info Actually Fires

World Info (also called World Book or lorebook) is the mechanism that keeps longrunning roleplay coherent. Instead of stuffing every fact into the characte…

  • world-book
  • world-info
  • lorebook
  • sillytavern
Read more