How the video is made
Two steps run one after another. First your text becomes speech, with a female or a male voice. Then a lip-sync model animates the portrait to that audio, so the mouth, jaw and small head movements follow the words. English, Russian and Persian were tested; other languages may work too.
It is used for greetings, a short announcement from an avatar, a fun message from a pet, or to give a voice to a restored family photo.
Before you start
- Use a clear, front-facing face; profiles and covered mouths do not animate well.
- The text can be up to 200 characters, about 20 seconds of speech.
- Only use photos of people who agree to it, and never to put words in a real person's mouth to mislead.
- The video takes a few minutes; it arrives in the chat when ready.




