Video avatars are digital facsimiles of human presenters who can deliver information as a real person would on video, but everything is powered by artificial intelligence. The technology behind video avatars accessible to developers of online learning has improved dramatically in recent years so that modern avatars can mimic natural movements, use lifelike facial expressions, and even adjust their tone to match the mood of the text they’re reading. This means that viewers can find these avatars surprisingly engaging and relatable—it almost feels as if you’re watching a real presenter, even though it’s all generated by a computer.
The first step in creating a video with a video avatar is writing the script, the content of the video presentation. For developers of online learning, this is typically teaching material, but could also be welcome messages, or explainer videos to help learners navigate the material, or establish what they are expected to do. The next step is to select the avatar. Many platforms offer pre-made avatars with a variety of appearances, ages, and personalities, that can be selected to fit your brand or audience for your learning programe. Some advanced platforms allow the creation of a custom avatar based on a real person by uploading a photo or providing a short video of yourself. This way, the presenter can actually look and sound like the developer or content author.
After choosing the look of the avatar, the next step is to set the voice. Most tools have a wide selection of voices, each with their own characteristics—such as pitch, speed, and accent—so that the voice can match the tone desired for the audience. Some platforms use sophisticated artificial voices that sound highly realistic, including regional accents or even gender. More sophisticated systems can also accept a recording a real human voice for an extra personal touch, either as audio or as part of a submitted video and the avatar will lip-sync to match your audio.
After this, the script can be uploaded to the platform. The AI ensures the avatar’s lips, facial muscles, and even head nods move in sync with the voice, just as a real presenter would. The process also allows the AI to infuse some natural human elements, such as blinking, smiling, or pausing for effect, which helps make the delivery feel more believable. Some platforms will also permit the addition of gestures, like waving or pointing, to make the message even clearer. You also manually add pauses to improve pacing or specify pronounciation where the automatic voice generation gets it wrong. Increasingly, avatar creation tools incorporate tools to translate the video into another language. Finally, they may offer the facility to generate captions to add accessibility.
Once the video is generated, it can be reviewed to see how it looks and sounds. If there are mistakes or required edits, there is no need to start over from scratch. The script can be edited, or the appearance or voice of the avatar can be changed, and the platform will create a new video automatically. This flexibility is one of the main reasons video avatars are so popular—they’re fast, easy, and adaptable.
As part of this programme, a video avatar was created. A short video was created and uploaded to the selected tool, in this case, Heygen (but other tools are available!):

Now see what happened when the avatar generator is asked to create the video
Press the play button to play the video
And then asked to regenerate the video in French with captions:
Press the play button to play the video
In the event of accessibility issues, you can download the course text from here
As always, human scrutiny is needed, as the avatar generator has no concept of whether its results are accurate. Some personal experiments with French translation proved quite successful, but German less so.