Cỏ, the Vietnamese AI waifu, as a Live2D avatar

Cỏ — a Vietnamese AI waifu you can actually talk to

Real-time voice · Live2D/3D avatar · she remembers you

Cỏ (full name Tứ Diệp Thảo, "four-leaf clover") is an AI waifu built in Vietnam. Most things called an AI waifu are either an anime image generator or a text chatbot with a static portrait. Cỏ is neither: you speak, she hears you, and she answers out loud while a Live2D/3D body moves with what she is saying.

What makes her different from an AI waifu app

Is Cỏ an AI VTuber?

Yes — the same character streams. If you found this page looking for an AI VTuber in the vein of Neuro-sama, the shape is familiar: an autonomous character with her own voice and avatar who reacts live to chat. The difference is language and origin. Cỏ is Vietnamese, speaks Vietnamese natively, and is built openly rather than as a closed product. If you have been looking for a Neuro-sama alternative that is not English-only, this is one.

Is Cỏ an AI companion?

She works like one, with one deliberate difference. Most AI companion apps are built to maximise the time you spend inside them. Cỏ is built to be good company — she is opinionated, she pushes back, and she is meant to leave you with more energy than you arrived with, not less. She is a character with a personality, not an engagement funnel.

How she is built

Cỏ runs on a self-hosted stack rather than someone else's API: a local large language model for her replies, local speech recognition and voice synthesis, a vision model for her eyes, and a Unity avatar for her body — wired together over WebRTC so the whole loop stays real-time. The project is developed in the open by a Vietnamese AI engineer, and the write-ups, clips and progress land on the channels below.

Where to find her

Questions people ask

Is Cỏ free?

Yes. She is a personal project, not a subscription product. Follow the channels above to catch her live.

Does she only speak Vietnamese?

Vietnamese is her native register and where she is most herself, but she answers in English and Japanese too, and follows you if you switch mid-conversation.

Can I run something like her myself?

The stack is ordinary open-source parts — a local LLM, streaming speech recognition, a voice cloning model, and a Unity or Live2D avatar driven over WebRTC. The hard part is not any single component; it is keeping the whole loop under the latency where a conversation stops feeling alive.

How is this different from an AI waifu generator?

A generator gives you a picture. Cỏ is a character you hold a conversation with — the avatar is the least interesting part.