Get Started

How to Build a PNGTuber and VTuber Stream Kit With AI

How to Build a PNGTuber and VTuber Stream Kit With AI - AI Art Create Blog Post
PNGTuber avatar generator
VTuber assets
Twitch emotes
Recraft V4
OBS

September 18th, 2026 β€’ 8 min read

Last updated at September 18th, 2026

A usable PNGTuber kit needs two matching avatar states, not one polished illustration. Start with a neutral mouth-closed portrait, derive the speaking state from that reference, then build expressions and channel graphics around the same palette.

AI Art Create generates the visual assets. Microphone detection, state switching, and Live2D or VRM rigging still happen in dedicated streaming software.

What belongs in a complete PNGTuber stream kit?

A starter PNGTuber kit contains a silent avatar, a speaking avatar, several emotional states, and a small set of channel graphics. Each asset must remain readable at its real display size and feel like the same character.

A practical first kit includes:

  • Mouth-closed neutral avatar
  • Mouth-open speaking avatar
  • Happy, surprised, annoyed, and sleepy reactions
  • Four to eight Twitch or Discord emotes
  • Subscriber badge or channel icon concepts
  • Starting Soon, Be Right Back, and Just Chatting backgrounds

Do not generate everything on day one. The silent and speaking states determine whether the core illusion works.

How should you design the base PNGTuber avatar?

Design the base avatar for fast recognition. A clear silhouette, stable head position, readable eyes, and a simple background matter more than intricate clothing texture.

For a portrait generator such as FLUX.2, use a front-facing or slight three-quarter crop:

Original PNGTuber character, chest-up portrait, centered.
Round face, dark green eyes, short copper hair, black cat-ear headset.
Navy hoodie with one mint drawstring, clean cel-shaded illustration.
Neutral friendly expression, mouth closed, even lighting.
Plain transparent-style green background, no text, no extra objects.

FLUX.2 costs 6 credits per request in AI Art Create and accepts a source image. Generate a small set, approve one portrait, and use that image for every later state.

How do you create matching talking and silent states?

Create the speaking state by changing only the mouth and a small amount of facial energy. If the head angle or crop changes, the avatar will appear to jump whenever the microphone activates.

Use a controlled edit prompt:

Preserve the character, crop, head position, eyes, hair, headset,
hoodie, palette, lighting, and background.
Change only the mouth to an open speaking shape.
Keep the jaw and face proportions unchanged.

Compare the two images at the same dimensions. Toggle between them quickly. Watch the outer head silhouette, shoulders, headset, and eye position. Small differences become obvious during repeated switching.

Step 1: Prepare the image files

Crop both states to the same canvas size. Align the eyes and shoulders. Remove the background if your streaming layout requires transparency.

Transparent raster output is not guaranteed across every model in the current catalog, so plan for a background-removal pass when needed.

Step 2: Configure reactive switching

Import the silent and speaking images into a PNGTuber application such as Veadotube Mini, or use another OBS-compatible reactive avatar setup. Assign microphone thresholds inside that software.

Reaction timing depends on the streaming application, audio buffer, microphone processing, and computer. AI Art Create is not part of the live switching loop after the files are exported.

Step 3: Test during a real stream scene

Test while the game, alerts, browser sources, and voice processing are running. A threshold that behaves well on an empty desktop may chatter or lag when noise suppression and a busy scene are active.

Check three situations: normal speech, quiet endings, and laughter. Tune the microphone threshold until the speaking state does not flicker between words.

How do you make Twitch emotes that remain readable?

Design Twitch emotes for the smallest display size first. Twitch commonly displays emotes at 28 by 28, 56 by 56, and 112 by 112 pixels. Fine fingers, thin outlines, and small text usually disappear at 28 pixels.

Recraft V4 is the current catalog model built for editable SVG output and flat illustration. It costs 6 credits per request. Ask for one expression, a thick silhouette, limited colors, and no text:

Vector chat emote of the same copper-haired cat-ear character laughing.
Head and hands only, thick clean outline, five flat colors,
exaggerated readable expression, centered, no text, no background.

Open the result in Illustrator or Figma when you need to correct curves or enforce exact brand colors. Recraft supports explicit color instructions, which is useful when the avatar and channel graphics must share a palette.

What should you generate for stream scenes?

Stream scenes should leave room for gameplay, chat, alerts, and the avatar. A beautiful full-frame illustration can still fail as an overlay if every corner is visually busy.

For a Starting Soon scene, identify where the title and countdown will sit. For Just Chatting, reserve a large calm area for the host and chat panel. For gameplay, frame decorative detail around the perimeter instead of the center.

Use a 16:9 canvas for common desktop streaming layouts. Generate backgrounds separately from text so spelling and future schedule changes do not require rebuilding the art.

Can AI Art Create generate a Live2D or VRM model?

AI Art Create does not currently generate rigged Live2D models, 3D meshes, or VRM files. It creates source images, vector graphics, and short video assets that can support a VTuber or PNGTuber workflow.

That limitation matters. VRoid Studio, Live2D Cubism, and specialist rigging tools solve different problems. Use AI Art Create to establish the character direction and produce supporting assets, then move into rigging software or work with an artist when deformable animation is required.

How do you add short reaction animation?

Animate an approved reaction still with a narrow image-to-video prompt. Wan 2.2 creates 5 second clips for 45 credits. Kling offers higher-fidelity 5 or 10 second motion for 100 credits.

Request one action, such as a blink and small smile. Avoid camera movement if the clip will loop beside a game window. Exported video may need trimming and loop cleanup in an editor.

Start with the two states that matter

Build the silent and speaking portraits before expanding the kit. Open the PNGTuber and VTuber asset workflow to compare recommended models. Your first generation is free, with no credit card required.

A

Author: AI Art Create

AI Image & Video Generation Expert

AI generation expert and content creator at AI Art Create. Passionate about the future of AI-powered image and video creation.