Generating clips using voice clones and virtual avatars
By using artificial intelligence to generate virtual avatars and voice cloning technology, combined with templates and personalized generation, the problems of low efficiency in user privacy protection and content creation in video conferencing are solved, and personalized virtual avatars and privacy protection are generated efficiently.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-08
- Publication Date
- 2026-04-10
AI Technical Summary
Existing technologies struggle to effectively protect user appearance and voice privacy and generate personalized virtual avatars in video conferencing, leading to user privacy leaks and low content creation efficiency.
By using artificial intelligence to generate virtual avatars and voice cloning technology, combined with templates and personalized avatar generation, and utilizing TTS models and video generation models, virtual avatars that resemble the user's appearance and voice are generated, and security is ensured through password and biometric verification.
It achieves personalization and consistency of virtual avatars, improves the efficiency of video content creation, protects user privacy, and reduces the consumption of computing and network resources.
Smart Images

Figure CN121838731A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates generally to digital content generation, and more specifically to clip generation with voice cloning and virtual avatars. BRIEF DESCRIPTION OF DRAWINGS
[0002] The accompanying drawings, which are incorporated in and form a part of the specification, illustrate one or more examples and, together with the description, explain certain principles and implementations of certain examples.
[0003] Figure 1 An example system that provides chat and video conferencing functionality to various client devices is shown;
[0004] Figure 2 An example system that provides chat and video conferencing functionality to various client devices is shown;
[0005] Figure 3 An example system that can establish virtual communication sessions is shown;
[0006] Figure 4 An example of an operating environment for generating clips with voice cloning and virtual avatars is shown;
[0007] Figure 5 An example diagram of a clip generation system for generating digital content with voice cloning and virtual avatars is shown;
[0008] Figure 6 An example of a video generator for generating clips with voice cloning and virtual avatars is shown; Figure 5
[0009] An example GUI for consent authorization requests for accessing personal data is shown; Figure 7
[0010] An example process for generating clips with voice cloning and virtual avatars is shown; Figure 8
[0011] An example process for generating clips with voice cloning and virtual avatars by participants during a video conference is shown; Figure 9
[0012] An example computing device suitable for use in example systems or methods for clip generation with voice cloning and virtual avatars is shown. Figure 10 DETAILED DESCRIPTION
[0013] Examples of generating environmental cutouts with voice cloning and virtual avatars are described herein. Those of ordinary skill in the art will realize and appreciate that the following description is illustrative only and is not intended to limit in any way. Reference will now be made in detail to implementations of the example embodiments as illustrated in the accompanying drawings. The same reference indicators will be used throughout the drawings and the following description to refer to the same or like items.
[0014] For clarity, not all of the routine features of the examples described herein are shown and described. It will of course be appreciated that in the development of any such actual implementation, numerous implementation-specific decisions must be made in order to achieve the developer's specific goals, such as compliance with application- and business-related constraints, and that these specific goals will vary from one implementation to another and from one developer to another. Without intent to limit the scope of what can be claimed, the text below and the accompanying drawings are illustrative only.
[0015] Video conferencing and video content have become common ways for people to interact and access information. People can be invited to participate in a video conference, join using a personal computer or phone, be able to see and hear the other participants, and essentially have a conversation as if in a face-to-face group meeting or event. In some cases, a participant can not want to show their real appearance or can not want to speak using their own voice. Accordingly, the participant can use artificial intelligence (AI) to generate a virtual character (e.g., a virtual avatar) that simulates their appearance and voice. The AI-generated avatar reflects the user’s head movements, facial expressions, and voice without showing the user’s real-time view or requiring the user to speak.
[0016] AI-generated avatars bring significant advantages to digital content creation, including increased engagement and cost-effectiveness. Compared to traditional video production, AI-generated avatars provide a consistent, scalable way of video production that ensures brand image while saving time and resources. AI-generated avatars also promote inclusivity and customization, making content resonate with different audiences’ needs. Additionally, AI-generated avatars protect privacy and enhance communication by effectively conveying emotions and expressions, making them an ideal choice for various applications such as marketing, training, and customer service.
[0017] One way to enable users to easily adopt AI-generated virtual avatars is to provide one or more template avatars. A template avatar is a pre-defined avatar that enables users to quickly start using AI-generated avatars, such as by using a text-to-avatar function to create an avatar narration or apply an avatar in a video conference. Template avatars allow users to quickly generate video content without the need for custom avatar creation. For appearance, a template avatar can reflect a participant’s head movements and facial expressions without showing the user’s real face. For voice, a text-to-speech (TTS) model is used to generate a voice that is very similar to the user’s voice.
[0018] Another approach is to allow users to create personalized avatars. Personalized avatars are user-generated clone avatars created by capturing video (including audio) of the user and generating a virtual avatar that resembles the user’s appearance and voice. Personalized avatars enable a more personalized and realistic user representation for digital content creation, video conference participation, and other environments.
[0019] One example for illustrating a method of generating digital content includes a clip generation system that generates a voice clone and a virtual avatar from a template video provided by a user. The template video includes audio and video data and is relatively short in length (e.g., about thirty seconds long). To create the template video, the user records their own speech while looking at a camera. The user can speak freely what they want to say, or the user can read a pre-defined script provided by the clip generation system. As part of recording the template video, and related to data privacy and security, the user also speaks a mandatory consent portion in the template video. The mandatory consent portion includes the user speaking their name, reciting a consent statement, and reciting a randomly generated password (e.g., a random number or string of characters). The user must recite the mandatory consent portion as a security protocol for future access to their avatar to ensure that the person accessing the avatar is indeed the user themselves, and not a malicious actor. The randomly generated password, along with the audio and video features, can be used to verify that the user has recorded a live video, and not a malicious actor provided a generated video to impersonate the user, which can help prevent spoofing or deepfakes. It can also verify that the user has consented to creating an avatar from their likeness and voice. After providing the template video, the clip generation system performs pre-processing on the template video to extract audio and video features from the template video. The extracted audio feature data and video feature data are saved in a database for the clip generation system to process in the future when the user wishes to use their virtual avatar.
[0020] To generate a virtual avatar for creating a voice clone for digital content, whether a template avatar or a personalized avatar, the clip generation system needs to include two models. The first model is a TTS model. Generally, a TTS model is used to generate human-like speech from a given text input. In this case, when a user wishes to use their virtual avatar to generate a clip, the user can provide a text input to the user interface of the clip generation system. The text input can be provided from a document (e.g., by uploading a slide deck, a word processing document, a script, etc.), a chat window on the user interface within the clip generation system that the user can type or speak into, or from a machine learning (ML) model that can be used to generate text data from a given input (e.g., keywords, general guidelines) such as a large language model (LLM). Additionally, an additional ML model can be integrated with the user interface on which the user uploads or provides the text input, such that the additional ML model can provide suggestions and improvements for the text input provided by the user.
[0021] Once the TTS model receives the text input, the TTS model of the clip generation system implements voice cloning by encoding the text input based on audio features extracted from the template video. The TTS model includes a neural network for synthesizing speech that closely matches the pitch, spectral features (e.g., acoustic features), and prosody of the audio of the template video. The TTS model employs a zero-shot or few-shot system that implements voice cloning by encoding speech from a short reference audio from the template video with little to no additional training or fine-tuning. The TTS model of the clip generation system then generates a cloned audio portion of digital content based on the text input.
[0022] A second model that is part of the clip generation system is a video generation model. The video generation model receives as input cloned audio data generated by the TTS model, the template video, and a reference image formed from a set of individual video frames captured from the template video. The reference image received by the video generation model is extracted from a single frame of the template video and provides the video generation model with identity features of the user with different facial expressions and head poses to help the clip generation system render a realistic video of the user.
[0023] As described above, the template video is provided as input to the video generation model for generating the video clip portion. However, the clip generation system will mask the mouth region of the user in the template video. The mouth region of the user in the template video is masked because the mouth region of each video frame of the template video will be replaced in a future processing step with a warped mouth region that corresponds to the speech audio generated by the TTS model. In other words, the video generation model uses the reference image to extract and warp the mouth region of the user from the audio input and subsequently blends the warped mouth region into each frame of the masked template video as part of the output presented by the virtual avatar reading the input text. In cases where the length of the input audio is longer than the template video (e.g., there are more audio frames than video frames), the template video is restarted once all of the video frames have been processed. Thus, the warped mouth region is continuously blended into each frame of the template video until the entire duration of the input audio has been processed, with the template video being reused as necessary for portions of the content.
[0024] More specifically, feature vectors are extracted for the entire length of the input audio. For a predefined time window (e.g., one to three seconds in length) of the input audio, a feature vector is extracted using ML techniques, e.g., using an autoencoder. Similarly, a feature vector is extracted from the reference image using ML techniques. Through an iterative process (e.g., iterating through each predefined time window using a sliding window), the video generation model receives the feature vector corresponding to the audio in the consecutive time window, the feature vector of the reference image, and a single frame of the template video with the mouth region masked. The video generation model warps the mouth region of the user based on the feature vector of the reference image and the feature vector of the input audio corresponding to the time window. The warped mouth region of the reference image is then merged from the template video into the masked region of the corresponding video frame. This process is iteratively performed until the entire input audio is processed, repeating the template video portion as needed to generate enough video frames for the clip. The output is a generated digital content (e.g., a video) that displays a virtual avatar with the appearance and voice of the user speaking the input text. Post-processing operations can be performed on the output digital content to smooth any artifacts that arise as part of the rendering process. In particular, a loss function is utilized to maintain the time smoothness of the virtual avatar in reading the input text in each video frame.
[0025] As described herein, certain embodiments improve digital content creation by more effectively generating virtual avatars using voice cloning that closely resembles the appearance and sound of the user. These virtual avatars with voice cloning allow users to use virtual avatars for various applications, such as video conferencing, video creation for marketing purposes, presentations, etc., allowing the appearance and sound of the virtual avatar to match the user’s likeness, voice, facial expressions, and other characteristics. The embodiments described herein tailor each virtual avatar to the user’s characteristics, making the virtual avatar more realistic than a template virtual avatar, significantly improving the speed of generating virtual avatars for users. This enables more efficient generation of digital content because users are able to quickly generate digital content with their voice and appearance, thereby reducing the time spent and computational and network resources (e.g., recording videos individually for each case) for digital content creation.
[0026] Additionally, the implementations described herein provide improvements in data privacy and security. By requiring a user to speak a randomly generated password, the clip generation system can require future users to re-enter the password when requesting to use a virtual avatar with voice cloning as a security protocol for accessing the virtual avatar using voice cloning. As detailed throughout this disclosure, other characteristics of the user, such as biometric characteristics, can also be used to verify the identity of the requester of the virtual avatar. For example, a camera or sensor of the user device can capture one or more biometric characteristics (e.g., voice characteristics, facial characteristics, etc.) and compare them to the audio and video characteristics extracted from the template video to determine whether there is a match. With the template video and additional security protocols such as passwords and biometric characteristics, the ability of a counterpart party to access the virtual avatar using voice cloning to impersonate or deepfake the user in the template video is inhibited.
[0027] This example is used to introduce the reader to the general subject matter discussed herein and is not intended to limit the disclosure to this example. The following sections describe various other non-limiting examples and examples of clip generation with voice cloning and virtual avatars.
[0028] Reference is now made to Figure 1 , Figure 1 An example system 100 that provides video conferencing functionality to various client devices is shown. The system 100 includes a chat and video conferencing provider 110 that is connected to a plurality of communication networks 120, 130 through which various client devices 140-180 can participate in video conferences hosted by the chat and video conferencing provider 110. For example, the chat and video conferencing provider 110 can be located within a private network to provide video conferencing services to devices within the private network, it can also be connected to a public network, such as the Internet, so that anyone can access it. Some examples can even provide a hybrid model in which the chat and video conferencing provider 110 can provide components to enable private organizations to host private internal video conferences or connect their systems to the chat and video conferencing provider 110 over a public network.
[0029] The system also selectively incorporates one or more authentication and authorization providers, such as an authentication and authorization provider 115 that can provide authentication and authorization services to users of the client devices 140-160. The authentication and authorization provider 115 can authenticate users of the chat and video conferencing provider 110 and manage user authorization for various services provided by the chat and video conferencing provider 110. In this example, the authentication and authorization provider 115 is operated by a different entity than the chat and video conferencing provider 110, but in some examples, they can be the same entity.
[0030] Chat and video conferencing provider 110 allows customers to create video conferences (or "meetings") and invite others to participate in those meetings, as well as perform other related functions such as recording meetings, generating transcripts of meeting audio, generating summaries and translations of meeting audio, managing user functionality in meetings, enabling texting during meetings, creating and managing breakout rooms from virtual meetings, etc. Figure 2 As described below, a more detailed description of the architecture and functionality of chat and video conferencing provider 110 is provided. It should be understood that the term "meeting" includes the meaning of the term "webinar" as used herein.
[0031] Meetings of this example chat and video conferencing provider 110 are held in a virtual room to which participants connect. In this case, the room is a construct provided by the server that provides a common point of reception of various video and audio data, which is then multiplexed and provided to different participants. While "room" is a conceptual label in this disclosure, any suitable functionality that enables multiple people to participate in a common video conference can be used.
[0032] To create a meeting with chat and video conferencing provider 110, a user can contact chat and video conferencing provider 110 using a client device 140-180 and select an option to create a new meeting. This option can be provided in a web page accessed by client devices 140-160 or by a client application executed by client devices 140-160. For telephone devices, an audio menu can be provided to the user that the user can navigate by pressing number buttons on their telephone device. To create a meeting, chat and video conferencing provider 110 can prompt the user for certain information, such as the date, time, and duration of the meeting, the number of participants, the type of encryption to be used, whether the meeting is private or open to the public, etc. After receiving the various meeting settings, chat and video conferencing provider can create a meeting record and generate a meeting identifier, and in some examples, a corresponding meeting passcode or password (or other authentication information), all of which meeting information is provided to the meeting host.
[0033] Upon receiving the meeting information, the user can distribute the meeting information to one or more users to invite them to the meeting. To start the meeting at the scheduled time (or immediately if the meeting is set to start immediately), the host provides the meeting identifier and, where applicable, the corresponding authentication information (e.g., passcode or password). The video conferencing system then starts the meeting and allows the users to join the meeting. Depending on the options set for the meeting, the users can enter immediately upon providing the appropriate meeting identifier (and authentication information, as applicable), even if the host has not arrived, or can be presented with information indicating that the meeting has not started, or can require special permission from the host for one or more users.
[0034] During the meeting, participants can use their client devices 140-180 to capture audio or video information and stream that information to the chat and video conferencing provider 110. They can also receive audio or video information from the chat and video conferencing provider 110 that is displayed by the respective client devices 140 to engage various users in the meeting.
[0035] At the conclusion of the meeting, the host can select an option to terminate the meeting, or the meeting can automatically terminate at a scheduled end time or after a predetermined duration. When the meeting terminates, the participants disconnect from the meeting and they will no longer receive audio or video streams of the meeting (and will stop transmitting audio or video streams). The chat and video conferencing provider 110 can also invalidate the meeting information (e.g., the meeting identifier or password).
[0036] To provide such functionality, one or more client devices 140-180 can communicate with the chat and video conferencing provider 110 using one or more communication networks, such as the network 120 or the public switched telephone network (“PSTN”) 130. The client devices 140-180 can be any suitable computing or communication device having audio or video functionality. For example, the client devices 140-160 can be conventional computing devices, such as desktop or laptop computers having a processor and computer readable media, that connect to the chat and video conferencing provider 110 using the Internet or other suitable computer network. Suitable networks include the Internet, any local area network (“LAN”), metropolitan area network (“MAN”), wide area network (“WAN”), cellular network (e.g., 3G, 4G, 4G LTE, 5G, etc.), or any combination of these networks. Alternatively, other types of computing devices, such as tablets, smartphones, and dedicated video conferencing equipment, can also be used. These devices can all provide audio and video functionality and can allow one or more users to participate in a video conference hosted by the chat and video conferencing provider 110.
[0037] In addition to the computing devices discussed above, the client devices 140-180 can also include one or more telephonic devices, such as a cellular telephone (e.g., cellular telephone 170), an Internet Protocol (“IP”) telephone (e.g., telephone 180), or a conventional telephone. Such telephonic devices can allow a user to place conventional telephone calls to other telephonic devices using the PSTN, including the chat and video conferencing provider 110. It should be understood that certain computing devices can also provide telephonic functionality and can operate as telephonic devices. For example, smartphones typically provide cellular telephone functionality and can therefore operate as telephonic devices. Similarly, IP telephones can also provide telephonic functionality and can operate as telephonic devices. Figure 1The phone device operates in the example system 100 shown in the middle. Additionally, a conventional computing device can execute software to enable phone functionality, which can allow a user to place and receive phone calls using a headset and microphone, among other things. Such software can communicate with a PSTN gateway to route calls from the computer network to the PSTN. Thus, a phone device includes any device that can place a conventional phone call, and is not limited to a dedicated phone device like a conventional phone.
[0038] Referring again to the client devices 140-160, these devices 140-160 contact the chat and video conferencing provider 110 using the network 120 and can provide information to the chat and video conferencing provider 110 to access functionality provided by the chat and video conferencing provider 110, such as accessing creating a new conference or joining an existing conference. To do so, the client devices 140-160 can provide user authentication information, a conference identifier, a conference passcode or password, among other things. In examples that employ an authentication and authorization provider 115, a client device, such as the client devices 140-160, can operate in conjunction with the authentication and authorization provider 115 to provide authentication and authorization information or other user information to the chat and video conferencing provider 110.
[0039] The authentication and authorization provider 115 can be any entity trusted by the chat and video conferencing provider 110 that can help authenticate users of the chat and video conferencing provider 110 and authorize users to access services provided by the chat and video conferencing provider 110. For example, the trusted entity can be a server operated by an enterprise or other organization with which a user has created an account containing authentication and authorization information, such as an employer or a trusted third party. A user can log into the authentication and authorization provider 115, for example, by providing a username and password, to access their account information at the authentication and authorization provider 115. The account information contains information established and maintained at the authentication and authorization provider 115 that can be used to authenticate and facilitate authorization for a particular user, regardless of which client device they can be using. An example of account information can be an email account established by a user at the authentication and authorization provider 115 and secured by a password or additional security features, such as single sign-on, hardware tokens, two-factor authentication, among others. However, such account information can be different from functionality such as email. For example, a healthcare provider can establish accounts for its patients. While the relevant account information can have associated email accounts, the account information is different from those email accounts.
[0040] Accordingly, a user's account information relates to a secure, verified set of information that can be used to authenticate and provide authorized services for a particular user and is accessible only by that user. Upon proper authentication, the associated user can then verify identity at other computing devices or services, such as chat and video conferencing provider 110. Chat and video conferencing provider 110 can need the prior explicit consent of the user before allowing access to the user's account information for authentication and authorization purposes by authentication and authorization provider 115.
[0041] Once the user is authenticated, authentication and authorization provider 115 can provide chat and video conferencing provider 110 with information about the services that the authorized user has access to. For example, authentication and authorization provider 115 can store information about user roles associated with the user. The user roles can include a set of services provided by chat and video conferencing provider 110 for which users assigned to those user roles have access. Alternatively, more or less granular methods of user authorization can be used.
[0042] When a user accesses chat and video conferencing provider 110 using a client device, chat and video conferencing provider 110 communicates with authentication and authorization provider 115 using information provided by the user to verify the user's account information. For example, the user can provide a username or cryptographic signature associated with authentication and authorization provider 115. Authentication and authorization provider 115 then confirms the information provided by the user or denies the request. Based on this response, chat and video conferencing provider 110 provides or denies access to its services accordingly.
[0043] For telephony devices, such as client devices 170-180, a user can call chat and video conferencing provider 110 to access video conferencing services. After answering the call, the user can provide information about the video conference, such as a conference identifier ("ID"), password or passcode, etc., to allow the telephony device to join the conference and participate using the audio devices of the telephony device (e.g., microphone and speaker), even if the telephony device does not provide video functionality.
[0044] Because phone devices are generally less functional than conventional computing devices, they can not be able to provide certain information to the chat and video conferencing provider 110. For example, a phone device can not be able to provide authentication information to the chat and video conferencing provider 110 to authenticate the phone device or the user. As a result, the chat and video conferencing provider 110 can provide more limited functionality to such phone devices. For example, a user can be allowed to join a conference after being provided with conference information, such as a conference identifier and password, but only as an anonymous participant in the conference. This can limit their ability to interact with the conference in some examples, such as limiting their ability to speak, hear, or view certain content shared during the conference, or access other conference functionality, such as joining breakout rooms or text chatting with other participants in the conference.
[0045] It should be appreciated that even in cases where a user is able to authenticate and employ a client device that is able to authenticate the user to the chat and video conferencing provider 110, the user can choose to participate in a conference anonymously and refuse to provide account information to the chat and video conferencing provider 110. The chat and video conferencing provider 110 can determine whether to allow such anonymous users to use the services provided by the chat and video conferencing provider 110. As with anonymous users using phone devices, discussed above, anonymous users for whatever reason, can be limited and in some cases can be blocked from accessing certain conferences or other services, or can be blocked from accessing the chat and video conferencing provider 110 altogether.
[0046] Referring again to the chat and video conferencing provider 110, in some examples it can allow client devices 140-160 to encrypt their respective video and audio streams to help improve their privacy in a conference. Encryption can be provided between the client devices 140-160 and the chat and video conferencing provider 110, or can be provided in an end-to-end configuration, where multimedia streams (e.g., audio or video streams) transmitted by a client device 140-160 are not decrypted until received by another client device 140-160 participating in a conference. Encryption can also be provided only for portions of a communication, for example, unencrypted communications that cross a national boundary can be encrypted.
[0047] Client-to-server encryption can be used to protect communications between the client devices 140-160 and the chat and video conferencing provider 110 while allowing the chat and video conferencing provider 110 to access the decrypted multimedia streams to perform certain processing, such as recording the conference for the participants or generating a conference record for the participants. End-to-end encryption can be used to make the conference completely private to the participants without having to worry that the chat and video conferencing provider 110 will gain access to important content of the conference. Any suitable encryption method can be employed, including key pair encryption of the streams. For example, to provide end-to-end encryption, the conference host's client device can obtain the public key of each other client device participating in the conference and securely exchange a set of keys to encrypt and decrypt the multimedia content transmitted during the conference. Thus, the client devices 140-160 can securely communicate with each other during the conference. Moreover, in some examples, certain types of encryption can be limited by the type of device participating in the conference. For example, a telephone device can lack the functionality to encrypt and decrypt the multimedia streams. Thus, while it can be desirable in many cases to encrypt the multimedia streams, it can not be necessary as it can prevent some users from participating in the conference.
[0048] By using Figure 1 the example system shown, users can use their respective client devices 140-180 to create and participate in conferences via the chat and video conferencing provider 110. Moreover, such a system enables users to use a variety of different client devices 140-180, from traditional standard-based video conferencing hardware to specialized video conferencing equipment, laptops or desktop computers, to handheld devices, to traditional telephone devices, and so on.
[0049] Referring now to Figure 2 , Figure 2 The example system 200 shown, the chat and video conferencing provider 210 provides video conferencing functionality to various client devices 220-250. The client devices 220-250 include two conventional computing devices 220-230, specialized equipment for a video conference room 240, and a telephone device 250. Each of the client devices 220-250 communicates with the chat and video conferencing provider 210 over a communications network, such as the Internet for the client devices 220-240 or the PSTN for the client device 250, substantially as described above Figure 1 with reference to FIG. 1. The chat and video conferencing provider 210 also communicates with one or more authentication and authorization providers 215, which can authenticate various users of the chat and video conferencing provider 210, substantially as described above Figure 1 with reference to FIG. 1.
[0050] In this example, chat and video conferencing provider 210 utilizes multiple different servers (or groups of servers) to provide various video conferencing functionality examples, enabling various client devices to create and participate in video conferences. Chat and video conferencing provider 210 uses one or more real-time media servers 212, one or more network service servers 214, one or more video room gateways 216, one or more messaging and online status gateways 217, and one or more telephone gateways 218. Each of these servers 212-218 is connected to one or more communication networks, enabling them to collectively provide client devices 220-250 with access to and participation rights in one or more video conferences.
[0051] Real-time media server 212 provides information to meeting participants (e.g., Figure 2 The client devices 220-250 shown provide multiplexed multimedia streams. While video and audio streams typically originate from the respective client devices, they are transmitted from client devices 220-250 via one or more networks to chat and video conferencing provider 210, where they are received by real-time media server 212. Real-time media server 212 determines which protocol is optimal based on factors such as proxy settings and the presence of firewalls. For example, client devices can choose audio and video over UDP, TCP, TLS, or HTTPS, and select content screen sharing over UDP.
[0052] Real-time media server 212 then multiplexes various video and audio streams based on the target client devices and delivers the multiplexed streams to each client device. For example, real-time media server 212 receives audio and video streams from client devices 220-240, while receiving only the audio stream from client device 250. Real-time media server 212 then multiplexes the streams received from devices 230-250 and provides the multiplexed streams to client device 220. Real-time media server 212 is adaptive in how it provides these streams, for example, reacting to real-time network and client changes. For example, real-time media server 212 can monitor parameters such as client bandwidth, CPU usage, memory, and network I / O, as well as network parameters such as packet loss, latency, and jitter, to determine how to modify the way the streams are provided.
[0053] The client devices 220 receive the streams, perform any decryption, decoding, and demultiplexing of the received streams, and then output audio and video using the client devices' video and audio devices. In this example, the live media server does not multiplex its own video and audio signals when sending streams to the client devices 220. Instead, each client device 220-250 receives multimedia streams only from other client devices 220-250. For telephone devices that lack video capabilities, such as client device 250, the live media server 212 transmits only multiplexed audio streams. The client devices 220 can receive multiple streams for a particular communication, allowing the client devices 220 to switch between streams to provide higher quality of service.
[0054] In addition to multiplexing multimedia streams, in some examples the live media server 212 can also decrypt incoming multimedia streams. As noted above, multimedia streams can be encrypted between the client devices 220-250 and the chat and video conferencing provider 210. In some such examples, the live media server 212 can decrypt incoming multimedia streams, appropriately multiplex the multimedia streams for the various clients, and encrypt the multiplexed streams for transmission.
[0055] As noted above Figure 1 The chat and video conferencing provider 210 can provide certain functionality with respect to unencrypted multimedia streams upon request of a user, as described above. For example, a conference host can request that a conference be recorded or that a transcript of the audio stream be prepared, which can then be performed by the live media server 212 using the decrypted multimedia streams, or the recording or transcription functionality can be offloaded to a dedicated server (or servers), such as a cloud recording server, to record the audio and video streams. In some examples, the chat and video conferencing provider 210 can allow a conference participant to notify it of inappropriate behavior or content in the conference. Such a notification can trigger the live media server 212 to record portions of the conference for review by the chat and video conferencing provider 210. Other functionality can also be implemented to take action based on decrypted multimedia streams at the chat and video conferencing provider, such as monitoring video or audio quality, adjusting or changing media encoding mechanisms, etc.
[0056] It will be appreciated that multiple real-time media servers 212 can participate in the data communications for a single conference, and that multimedia streams can be routed through multiple different real-time media servers 212. Moreover, the various real-time media servers 212 can not be located in the same location, but can be located in multiple different geographic locations, which can enable high quality communications between clients dispersed over a wide geographic area, e.g., located in different countries or different continents. Moreover, in some examples, one or more of these servers can be co-located at a client's premises, e.g., an enterprise or other organization. For example, different geographic regions can each have one or more real-time media servers 212 to enable client devices in the same geographic region to have high quality connections with the chat and video conferencing provider 210, e.g., to send and receive multimedia streams, rather than connecting to real-time media servers located in different countries or different continents. The local real-time media servers 212 can then communicate with physically distant servers using high speed network infrastructure, e.g., the Internet backbone, that would otherwise not be directly available to the client devices 220-250. Thus, routing of multimedia streams can be distributed throughout the video conferencing system and handled by different real-time media servers 212.
[0057] With respect to the network services servers 214, these servers 214 provide management functions to enable client devices to create or participate in conferences, send conference invitations, create or manage user accounts or subscriptions, and other related functions. Moreover, these servers can be configured to perform different functions or operate at different levels, e.g., to manage portions of the chat and video conferencing provider with supervisory server sets for particular regions or locations. When a client device 220-250 accesses the chat and video conferencing provider 210, it typically communicates with one or more network services servers 214 in order to access its account or participate in a conference.
[0058] In this example, when a client device 220-250 first contacts the chat and video conferencing provider 210, the client device is routed to the web service server 214. The client device can then be provided with access credentials, such as a username and password or single sign-on credentials, for the user to gain authenticated access to the chat and video conferencing provider 210. This process can include the web service server 214 contacting the authentication and authorization provider 215 to verify the provided credentials. Once the user’s credentials are accepted, and the user has agreed, by interacting with the web service server 214, the web service server 214 can perform administrative functions, such as updating user account information, if the user has account information stored in the chat and video conferencing provider 210, or scheduling a new meeting. The authentication and authorization provider 215 can be used to determine which administrative functions a particular user can access based on assigned roles, permissions, groups, etc.
[0059] In some examples, users can access the chat and video conferencing provider 210 anonymously. When communicating anonymously, the client device 220-250 can communicate with one or more web service servers 214, but only provide information to create or join a meeting, depending on which feature functions the chat and video conferencing provider allows anonymous users to use. For example, an anonymous user can access the chat and video conferencing provider using a client device 220 and provide a meeting ID and password. The web service server 214 can use the meeting ID to identify an upcoming or ongoing meeting and verify that the password for the meeting ID is correct. After verification, the web service server 214 can then communicate information to the client device 220 to enable the client device 220 to join the meeting and communicate with the appropriate live media server 212.
[0060] In the case where a user wishes to schedule a meeting, the user (anonymous or authenticated) can select an option to schedule a new meeting and can then select various meeting options, such as the date and time of the meeting, the duration of the meeting, the type of encryption to use, one or more users to invite, privacy controls (e.g., no anonymous users allowed, block screen sharing, manual authorization to enter the meeting, etc.), meeting recording options, etc. The web service server 214 can then create and store a meeting record for the scheduled meeting. When the scheduled meeting time arrives (or within a threshold period of time in advance), the web service server 214 can accept requests to join the meeting from various users.
[0061] To process the request to join the conference, the network service server 214 can receive conference information (e.g., a conference ID and password) from one or more client devices 220-250. The network service server 214 finds the conference record corresponding to the provided conference ID and then confirms whether the scheduled start time of the conference has arrived, whether the conference host has started the conference, and whether the password matches the password in the conference record. If the request is made by the host, the network service server 214 activates the conference and connects the host to the live media server 212 so that the host can begin sending and receiving multimedia streams.
[0062] Once the host starts the conference, subsequent users requesting access are allowed to join the conference if the conference record is found and the password matches the password provided by the requesting client device 220-250. In some examples, additional access controls can also be used. However, if the network service server 214 determines that the requesting client device 220-250 is allowed to join the conference, the network service server 214 identifies the live media server 212 to handle the multimedia streams for the requesting client device 220-250 and provides the client device 220-250 with information to connect to the identified live media server 212. When additional client devices 220-250 request access through the network service server 214, they can be added to the conference.
[0063] After joining the conference, the client devices will send and receive multimedia streams via the live media server 212, but they can also communicate with the network service server 214 as needed during the conference. For example, if the conference host leaves the conference, the network service server 214 can designate another user as the new conference host and assign that user host management privileges. Hosts can have management privileges that allow them to manage the conference, such as enabling or disabling screen sharing, muting or removing users from the conference, assigning or moving users to main conference rooms or breakout rooms (if any), recording the conference, etc. Such functionality can be managed by the network service server 214.
[0064] For example, if the host wishes to remove a user from the conference, they can select the user for removal and issue a command through the user interface on their client device. The command can be sent to the network service server 214, which can then disconnect the selected user from the corresponding live media server 212. If the host wishes to remove one or more participants from the conference, such a command can also be handled by the network service server 214, which can terminate the authorization for one or more participants to join the conference.
[0065] In addition to creating and managing ongoing meetings, the web services server 214 can also be responsible for closing and dissolving meetings after they have ended. For example, a meeting host can issue a command to the web services server 214 to end an ongoing meeting. The web services server 214 can then remove any remaining participants from the meeting, communicate with one or more live media servers 212 to stop streaming audio and video for the meeting, and deactivate (e.g., by removing the corresponding password for the meeting from the meeting record) or delete the meeting record corresponding to the meeting. Thus, if a user later attempts to access the meeting, the web services server 214 can deny the request.
[0066] Depending on the functionality provided by the chat and video conferencing provider, the web services server 214 can provide additional functionality, such as by providing private meeting functionality for organizations, special types of meetings (e.g., webinars), etc. Such functionality can be provided in accordance with the various examples of video conferencing providers according to the present specification.
[0067] Referring now to the video room gateway servers 216, these servers 216 provide an interface between dedicated video conferencing hardware, such as can be used in a dedicated video conference room. Such video conferencing hardware can include one or more cameras and microphones, as well as a computing device for receiving video and audio streams from each camera and microphone and connecting with the chat and video conferencing provider 210. For example, the video conferencing hardware can be provided by the chat and video conferencing provider to one or more subscribers, who can provide access credentials to the video conferencing hardware for connecting to the chat and video conferencing provider 210.
[0068] The video room gateway servers 216 provide dedicated authentication and communication with dedicated video conferencing hardware, which can not be available to other client devices 220-230, 250. For example, the video conferencing hardware can be registered with the chat and video conferencing provider upon first installation, and the video room gateway can use such registration along with information provided to the video room gateway server when the dedicated video conferencing hardware connects to the video room gateway server 216 (e.g., device ID information, subscriber information, hardware capabilities, hardware version information, etc.) to authenticate the video conferencing hardware. Upon receiving such information and authenticating the dedicated video conferencing hardware, the video room gateway server 216 can interact with the web services server 214 and live media servers 212 to allow the video conferencing hardware to create or join meetings hosted by the chat and video conferencing provider 210.
[0069] Referring now to the telephony gateway servers 218, these servers 218 support and facilitate the participation of telephony devices in conferences hosted by the chat and video conferencing provider 210. Because telephony devices use PSTN communications and do not use computer networking protocols such as TCP / IP, the telephony gateway servers 218 act as an interface that translates between the PSTN and the network systems used by the chat and video conferencing provider 210.
[0070] For example, if a user connects to a conference using a telephony device, they can dial a telephone number that corresponds to one of the telephony gateway servers 218 of the chat and video conferencing provider. The telephony gateway server 218 will answer the call and generate audio information that asks the user to provide information such as a conference ID and password. The user can input such information using buttons on the telephony device, for example, by sending a dual-tone multi-frequency ("DTMF") audio stream to the telephony gateway server 218. The telephony gateway server 218 determines the numbers or letters entered by the user and provides the conference ID and password information to the web services server 214 along with a request to join or start a conference, as generally described above. Once the telephony client device 250 is accepted into the conference, the telephony gateway server replaces the telephony device in the conference.
[0071] After joining the conference, the telephony gateway server 218 receives audio streams from the telephony device and provides them to the corresponding real-time media server 212, and receives audio streams from the real-time media server 212, decodes them, and provides the decoded audio to the telephony device. Thus, the telephony gateway server 218 essentially operates as a client device, while the telephony device primarily operates as an input / output device for the corresponding telephony gateway server 218, such as a microphone and speaker, thereby enabling the user of the telephony device to participate in a conference without using a computing device or video.
[0072] It should be understood that the components of the chat and video conferencing provider 210 discussed above are merely examples and example architectures for such devices. Some video conferencing providers can provide more or fewer functions than described above, and can not separate the functions into the different types of servers described above. Rather, any suitable server and network architecture can be used according to different examples.
[0073] Referring now to Figure 3 , Figure 3 An example system 300 in which virtual communication sessions can be established is shown. In this example system 300, a communication platform 310 and a plurality of client devices 340A-340N (which can be referred to individually as a client device 340 or collectively as the client devices 340 herein) are connected via a network 320. The communication platform 310 can be a chat and video conferencing provider 110 or a web services server 214 of a chat and video conferencing provider 110 as described above. Figure 1 Figure 2 The chat and video conferencing provider 210. The network 320 may be the Internet or any suitable communication network or combination of communication networks, including LAN (e.g., within a company's private LAN), WAN, MAN, cellular network (e.g., 3G, 4G, 4G LTE, 5G, etc.) or any combination thereof.
[0074] Client device 340 can be any suitable computing or communication device. Figure 1 Client devices (e.g., 140, 150, 160, or 170) or Figure 2 The client devices (e.g., 220, 230, or 250) are connected to the communication platform 310 via the Internet or other suitable computer networks. For example, client device 340 may be a desktop computer, laptop computer, tablet computer, or smartphone with a processor and computer-readable media connected to the communication platform 310. Client devices 340 are equipped with communication software that enables them to connect to the communication platform 310 for chat, video conferencing, email, and any other suitable communication. For example, during a video conferencing session, the user of an associated client device (e.g., client device 340A) can interact with other users associated with other client devices (e.g., client devices 340B-340N) via the communication platform 310 through audio and video.
[0075] Now for reference Figure 4 , Figure 4 An example of an operating environment 400 is shown for use with voice cloning and virtual avatar generation clips using specific aspects described herein. Operating environment 400 includes a chat and video conferencing provider 410, configured to host and provide various video conferencing functionalities, for example, similar to... Figure 1 The chat and video conferencing provider described is 110 and Figure 2 The chat and video conferencing provider 210 is described. The chat and video conferencing provider 410 is configured to receive data from client computing devices 440A- 440C (which may be individually referred to herein as client computing device 440 or collectively as client computing device 440) receives a video conference stream and transmits it to the client computing device. The video conference stream may include video signals from the participants, audio signals captured at the corresponding client computing device associated with each participant, and other signals or streams related to the participants. Client computing device 440 may be any of the aforementioned... Figure 1 , Figure 2 and Figure 3 The client devices discussed are 140-180, 220-250, and 340A-340N.
[0076] One or more client computing devices, such as client device 440A, can install the clip generation system 460. The clip generation system 460 is configured to generate a virtual avatar 462 with a voice clone for a user based on input text 452 and a template video 454 received from the communication application 450. For example, the clip generation system 460 can extract audio features from audio data of the template video 454 and extract appearance attributes of the user from video data of the template video 454 to generate a virtual avatar 462 with a voice clone that mimics the appearance and voice of the user. The virtual avatar 462 with the voice clone generated by the clip generation system 460 can be used to generate digital content 420 associated with the input text 452. The digital content 420 can be streamed to other participants or users of the chat and video conferencing provider 410 via the chat and video conferencing provider 410. Alternatively, the digital content 420 can be generated and stored in a local (or remote) data store, such as data store 430, for future use by the user of the client device 440A.
[0077] The virtual avatar 462 with the voice clone can be generated by the clip generation system 460 based on the template video 454 and, as described above, the virtual avatar 462 with the voice clone can be stored in the data store 430 for subsequent use by the user. Thus, the clip generation system 460 only needs to generate the virtual avatar 462 with the voice clone in one instance. The user can reuse the virtual avatar 462 with the voice clone to generate digital content 420 based on any input text desired by the user, such as the input text 452. In this way, the present example saves computing resources because the clip generation system 460 can pre-generate the virtual avatar 462 with the voice clone for unlimited future use. In some examples, the user can update the virtual avatar 462 with the voice clone by providing a new template video 454, in which case the clip generation system 460 can generate a new virtual avatar 462 with the voice clone.
[0078] Figure 5 and Figure 6The clip generation system 460 is described in greater detail. Generally, the clip generation system 460 can include one or more ML models, such as an autoencoder or a generative adversarial network (GAN), for generating clips (e.g., digital content 420) using virtual avatars 462 with voice cloning. The one or more ML models of the clip generation system 460 can include one or more encoder models and one or more decoder models, such as one or more image encoder-decoder pairs, one or more audio encoder-decoder pairs, and the like. The encoders and decoders use, execute, or otherwise incorporate trained ML models, including transformer models or convolutional neural networks (CNNs), recurrent neural networks (RNNs), variational autoencoders, transformer models, autoencoders, GANs, and the like.
[0079] When a user wishes to use their virtual avatar 462 with voice cloning, the user can provide input text 452 to the communication application 450. The communication application 450 can retrieve a pre-generated virtual avatar with voice cloning 462 from the data store 430 (or generate a new virtual avatar with voice cloning 462 by communicating with the clip generation system 460) to generate digital content 420. More specifically, when a client device 440 (e.g., client device 440A, etc.) joins a video conference via the communication application 450, the communication application 450 can retrieve a virtual avatar with voice cloning 462 generated for the user associated with the client device from the data store 430 and present digital content 420 for the user according to the input text 452.
[0080] The template video 454 includes template audio data and template video data and is recorded by a user associated with a client device (e.g., client device 440A, etc.). The template video data includes consecutive frames that together form the template video data. In some examples, the template video 454 is relatively short in length (e.g., thirty seconds in length). In other examples, the template video 454 can be longer in length (e.g., two minutes in length). The template audio data includes audio data associated with the user’s voice. When recording the template video 454, the user associated with the client device looks at the camera while recording their own speech. In some examples, the user can speak freely whatever they want to say. In other examples, the user can read a predefined script provided by the communication platform 310. In other examples, the user can combine free speech and a predefined script to record the template video 454.
[0081] During the recording of the template video 454, the user can be required to read a piece of mandatory consent content. The mandatory consent content can include the user's name, the user's consent to use video and audio to generate a synthetic video and audio of the user as a virtual avatar, and to read a randomly generated password. In one example, the randomly generated password can be generated by the communication application 450 through a random number generator. In other examples, the randomly generated password can include numbers and characters such that the randomly generated password is a random number and string of characters. The randomly generated password can be unique to the user associated with the client device such that each user of various client devices has their own unique password to access their own unique virtual avatar 462 with voice cloning. Further, the template video 454 is a single continuous video recording of the user speaking while looking at the camera. Requiring the template video 454 to be a single continuous video recording can enhance security measures. In particular, the continuous recording ensures that the statements used to capture the audio semantics of the user and the mandatory consent content are spoken by the same user in the same video recording. As such, a malicious actor attempting to create the template video in a fraudulent manner would not be able to edit the template video to include consent content read by someone else.
[0082] Including the randomly generated password as part of the template video 454 can also enhance security and privacy. For example, if the user associated with the client device wishes to subsequently access the virtual avatar 462 with voice cloning, the user can be required to re-enter the randomly generated password. Upon entry, the communication application 450 can compare the randomly generated password recited in the template video 454 with the randomly generated password provided by the user. If the communication application 450 determines that the provided password matches, the communication application 450 can allow the user to continue generating digital content 420 using the virtual avatar 462 with voice cloning.
[0083] Other security and privacy features can be integrated into the computing environment 400. Figure 4 Although the foregoing examples have been described in the context of a single user, the examples are not limited to a single user. For example, the computing environment 400 can be configured to allow multiple users to access the virtual avatar 462 with voice cloning. In such an example, the communication application 450 can be configured to allow each user to create their own template video 454. The communication application 450 can then be configured to allow each user to access their own virtual avatar 462 with voice cloning. In such an example, the communication application 450 can be configured to allow each user to access their own virtual avatar 462 with voice cloning only if the user provides the correct randomly generated password. In such an example, the communication application 450 can be configured to allow each user to access their own virtual avatar 462 with voice cloning only if the user provides the correct randomly generated password. Figure 4The additional audio and video sensors can be used to verify the identity of a user associated with a client device seeking access to the virtual avatar with voice cloning 462. More specifically, after the virtual avatar with voice cloning 462 is pre-generated, a user associated with a client device can wish to use the virtual avatar with voice cloning 462 to generate digital content 420. Before the communication application 450 allows the user to access the virtual avatar with voice cloning 462, the communication application 450 can use one or more additional sensors integrated into or connected to the client device to extract one or more biometrics associated with the user. The communication application 450 can compare the extracted one or more biometrics to the biometrics of the template video 454 to determine whether the user attempting to access the virtual avatar with voice cloning 462 is the same user as in the template video 454. In the event of a biometric match, the communication application 450 can authorize the user to access the virtual avatar with voice cloning 462 to generate digital content 420.
[0084] With continued reference to Figure 4 When a user associated with a client device attempts to use the virtual avatar with voice cloning 462 to generate digital content 420, the user can provide input text 452 via the communication application 450. The input text 452 can be associated with a voice portion of the digital content 420 that the user wishes to generate. In other words, the virtual avatar with voice cloning 462 can speak the input text 452 as part of the digital content 420. As Figure 5-9 In detail, the input text 452 will be provided to a TTS model of the clip generation system 460 to generate cloned audio data that mimics the audio data extracted from the template video 454. Thus, the input text 452 is converted to synthetic speech that is similar in acoustic and related features to the user’s voice.
[0085] To generate cloned audio associated with input text 452, clip generation system 460 can receive a message from communication application 450. The message can include a payload with input text 452. A user associated with the client device can provide input text 452, for example, by typing input text 452 via a user interface of communication application 450 (e.g., using a keyboard or touch screen), by sending an SMS text message, by sending a chat message, by sending an email, etc. The user can provide input text 452 during a meeting hosted by chat and video conferencing provider 410 without having to join the meeting, such that participants of the meeting (e.g., other users associated with client devices such as client device 440B and client device 440C) can view digital content 420 as the user’s virtual avatar speaks. In some implementations, the user can provide input text 452 via a wearable electronic device or a virtual reality (VR) device.
[0086] In some implementations, the user can provide input text 452 by uploading a file (e.g., a slide deck, a word processing file, etc.). In this case, communication application 450 can scan the uploaded file and extract text data from the uploaded file. The extracted text data can then be provided as input text 452 to clip generation system 460.
[0087] In some implementations, the LLM can be integrated with (or otherwise accessed by) the communication application 450 to assist a user in generating text, resulting in input text 452. For example, a user can generate input text 452 via prompting the LLM with keywords, prompts, and guidelines. In one example, a user can be required to provide a speech or presentation as part of a meeting hosted by the chat and video conferencing provider 410. However, the user can not be able to join the meeting in real-time. Thus, the user can utilize a virtual avatar 462 with speech cloning to speak on their behalf. The user can provide guidelines, keywords, and prompts to the LLM to generate input text 452. In some examples, the user can provide a copy of a set of presentation slides to the LLM to generate input text 452. The input text 452 generated by the LLM can then be provided to the clip generation system 460 to generate digital content 420 to be delivered on the meeting. In this way, the user does not need to attend the live meeting, but other participants of the meeting (e.g., users associated with client devices such as client devices 440B-440C) can view the presentation with a speech cloned virtual avatar 462 that mimics the user’s voice and appearance. Additionally, the user can utilize the LLM to refine, edit, or otherwise receive suggestions related to the input text 452 provided by the user, rather than using the LLM to generate the entire input text 452. In this way, the LLM can provide suggestions to the user related to the input text 452 (e.g., recommend different word choices, options to make the input text 452 more concise, etc.).
[0088] In other implementations, a user associated with a client device seeking to generate digital content 420 using a virtual avatar 462 with speech cloning can tag input text 452 with emotional metadata (e.g., happy, sad, serious, bored, etc.). As detailed below, the emotional metadata can be utilized by the clip generation system 460 to generate a virtual avatar 462 with speech cloning that is based on the tagged emotional metadata. For example, a user can tag input text 452 with emotional metadata of “excited,” and a TTS model of the clip generation system 460 can generate cloned audio data that is louder in volume and heavier in tone to match the tagged emotional metadata. Figure 5
[0089] Reference is now made to Figure 5 , Figure 5 An example diagram 500 for a clip generation system 460 for generating digital content 420 with speech cloning and virtual avatars is shown. Figure 5 The example diagram 500 shown in FIG. 5 involves operations performed by various modules and models of the clip generation system 460 for generating digital content 420 with speech cloning and virtual avatars (e.g., virtual avatars 462) using input text 452. The example diagram 500 is described in the context of a user associated with a client device (e.g., client device 440A) seeking to generate digital content 420 using a virtual avatar 462 with speech cloning. The user can provide input text 452 to the clip generation system 460, which can include a natural language processing module 510, a speech cloning module 520, a virtual avatar generation module 530, and a TTS model 540. Figure 4 The described virtual avatar 462 with voice cloning generates digital content 420. More specifically, the clip generation system 460 can receive the template video 454 and the input text 452 as inputs, while using the pre-processing module 510, the TTS model 520, the video generation model 530, and the post-processing module 550, the clip generation system 460 can generate the digital content 420. The digital content 420 can be generated through offline processing and stored in a data store (not shown) for later access by a user of the clip generation system 460. Additionally, the digital content 420 can also be generated through online processing, such as in real-time or live. For example, the digital content 420 can be generated for a user participating in a video conference and wishing to speak in the video conference without having to show up or actually speak. In this case, the user can use the digital content 420, including its virtual avatar 462 with voice cloning, to speak on their behalf, such as by typing text in a text input field in the GUI of the communication application 450.
[0090] The pre-processing module 510 can receive the template video 454 as input and perform pre-processing operations on the template video 454. The pre-processing operations performed on the template video 454 extract relevant information in the template video 454 that can be used to accurately and efficiently generate the digital content 420 using the virtual avatar 462 with voice cloning. First, the pre-processing module can deconstruct the template video 454 into individual and consecutive video frames that together form the template video 454. For example, as previously discussed, a user of the clip generation system 460 can record themselves into a camera to generate the template video 454. The camera used to record the template video 454 can have a frame rate (e.g., thirty frames per second (fps)) for recording the user. Thus, for a thirty-second template video recorded by the user using a camera with a frame rate of thirty frames per second, the pre-processing module 510 can deconstruct the template video 454 into a consecutive set of video frames (e.g., 900 frames). The template video 454 can also have other example frame rates and lengths, such as a 60 fps and a two-minute template video length. Those of ordinary skill in the art will find many variations, modifications, and alternatives.
[0091] After the template video 454 is disassembled into each individual and consecutive video frame, the pre-processing module 510 can mask the mouth region of the relevant user in the template video 454 to generate a set of masked video frames 512. The mouth region of the user in the template video 454 can be masked using a masking algorithm configured to detect the edges or specific colors of each template video frame. For example, the masking algorithm can be configured to detect the edges of the user’s face in consecutive video frames of the template video 454 based on the color variations of the pixels of the background versus the mouth region of the user in the consecutive video frames of the template video 454. The masked video frames 512 can be stored in a data storage device for use in subsequent processing steps by the video generation model 530.
[0092] Additionally, the operation of generating the masked video frames 512 performed by the pre-processing module 510 can be performed once and the masked video frames 512 can be stored in a data storage device for future use. In other words, each time a user seeks to generate digital content 420 using the virtual avatar 462 with voice cloning, the clipping generation system 460 does not have to perform the processing steps of masking the mouth region of consecutive frames of the template video 454 again. Instead, the clipping generation system 460 can retrieve the masked version of the template video 454 from memory at runtime.
[0093] The pre-processing module 510 can also extract a set of reference images 514 using consecutive video frames of the template video 454. Selecting the reference images 514 is an important step in generating digital content 420 using the virtual avatar 462 with voice cloning that is highly similar to the user’s voice and appearance because the reference images 514 are used as a reference in subsequent processing steps by the video generation model 530 to capture various nuances and natural habits associated with the user’s voice and appearance. The reference images 514 can be selected by the pre-processing module 510 based on one or more features of the user in the template video 454, such as the user’s head pose, the user’s mouth shape, the user’s viewing direction, and the user’s eye features. Selecting the reference images to capture various perspectives can enable a more accurate virtual avatar with voice cloning in the digital content 420 because the warped video frames 536 can be closer to the semantics and natural body movements of the user when speaking.
[0094] Additionally, the pre-processing module 510 can extract template audio data 516 from the template video 454. The template audio data 516 can refer to the sound components that accompany the visual content of the template video 454. In other words, the template audio data 516 is synchronized with the successive frames of the template video 454 to ensure that the sound matches the user’s visual actions (e.g., the user’s speech and vocalizations conform to the user’s lip movements, mouth shapes, head poses, etc.). The template audio data 516 can include all the auditory elements from the template video 454, such as speech, music, sound effects, and other background noise. The template audio data 516 can be stored as a separate audio file, such as a WAV file, MP3 file, etc., for use by the TTS model 520. As described below, the template audio data 516 will be used by the TTS model 520 to generate cloned audio data 522 with synthesized acoustic features 524 and synthesized prosodic features 526 that mimic the template audio data 516 and are based on the input text 452.
[0095] With continued reference to Figure 5 , the masked video frames 512, the reference images 514, and the template audio data 516 are provided as inputs to the TTS model 520 and the video generation model 530 of the clip generation system 460. The template audio data 516 and the input text 452 are provided as inputs to the TTS model 520. As described above Figure 4 , the input text 452 can be provided via user typing into a user interface window of the communication application 450, via an LLM, via file upload and scanning, via capturing user speech, etc.
[0096] The TTS model 520 of the clip generation system 460 can process the input text 452 and the template audio data 516 to generate cloned audio data 522. In particular, the TTS model 520 can process the template audio data 516 using one or more ML models with one or more audio encoder-decoder pairs, e.g., to extract audio features / parameters associated with one or more acoustic features and one or more prosodic features of the user, which the TTS model 520 can then use to generate cloned audio data 522 with synthesized acoustic features 524 and synthesized prosodic features 526 that mimic the acoustic features and prosodic features of the template audio data 516. For example, the TTS model 520 can include a trained ML model, such as an audio encoder that can generate an embedded representation of the acoustic features and prosodic features of the template audio data 516. The embedded representation can be one or more high-dimensional vectors that include quantized components that characterize the input speech, thereby efficiently encoding the extracted feature set. The audio encoder of the TTS model 520 can be any suitable type and trained ML model, including but not limited to a CNN, an RNN, a transformer, an autoencoder, a variational autoencoder, a GAN, etc.
[0097] The acoustic features of the template audio data 516 can refer to the physical properties of the template audio data 516, such as pitch, energy, duration, rhythm, stress, and intonation. The pitch can be associated with the fundamental frequency of the template audio data 516, which affects the intonation of the template audio data 516. The rhythm can define the rate or speed of the template audio data 516. The tone (e.g., inflection or intonation) can control the emphasis on certain words. For example, text containing all capital letters, italicized, bolded, or underlined words, or certain words in color or highlighted, or emojis or GIFs in the input text 452 can cause the intonation of the output cloned audio data 522 to change such that the cloned audio data 522 emphasizes certain words when outputted (e.g., mimics the emphasis spoken). The energy can control the energy level of the template audio data 516, such as the volume of the speech and the dynamic range of the speech. The duration can refer to the length of time a particular phoneme or segment of the template audio data 516 lasts. The encoder of the TTS model 520 can be trained to encode these acoustic features to fit into the vector space of other downstream ML components (e.g., audio decoder) of the TTS model 520.
[0098] The prosodic features of the template audio data 516 can refer to the patterns and variations in the pitch, loudness, and timing of the template audio data 516. The prosodic features refer to features related to long stretches of speech (e.g., phrases or sentences). The prosodic features can include elements of rhythm, tone, and stress patterns that convey the user’s meaning, emotion, and intent. Thus, the prosodic features of the template audio data 516 refer to the high-level patterns of the template audio data 516. For example, the prosodic features of intonation can refer to the changes in pitch over a phrase or an entire sentence to indicate a question, statement, emphasis, or emotion. The stress can refer to the emphasis on certain syllables or words that can change the meaning of a sentence. The rhythm and timing can refer to the recognition of patterns in the timing and duration of syllables, words, and pauses by the user when speaking. Similar to the acoustic features described above, the encoder of the TTS model 520 can be trained to encode these prosodic features to fit into the vector space of other downstream ML components (e.g., audio decoder) of the TTS model 520.
[0099] After extracting the acoustic and prosodic features of the template audio data 516, the TTS model 520 can use these extracted features (e.g., embedded representations of the template audio data 516) along with the input text 452 to generate cloned audio data 522. More specifically, the TTS model 520 can use one or more additional ML models, such as an audio decoder, to generate the cloned audio data 522 based on the embedded representations of the template audio data 516 and the input text 452. The decoder can be a trained ML model to convert the embedded representations into transformed embedded representations that are sound predictions of the cloned audio data 522 under different sets of features. Similar to the audio encoder, the audio decoder of the TTS model 520 can be any suitable type and trained ML model, including but not limited to CNNs, RNNs, transformers, autoencoders, variational autoencoders, GANs, etc.
[0100] As previously mentioned, the TTS model 520 can include one or more pre-trained ML models to generate the cloned audio data 522. In some examples, the TTS model 520 can be trained directly from audio data without any training based on text data. For example, one or more ML models of the TTS model 520, such as the aforementioned encoder-decoder pair, can be trained using a training dataset that includes data samples corresponding to recorded speech samples of one or more users or pre-recorded speech samples selected by the user (e.g., audio segments of selected speech, which can not be the user’s own speech but speech selected by the user).
[0101] Specifically, the audio encoder of the TTS model 520 can be trained to encode audio data into continuous latent representations that predict acoustic and prosodic features in the audio data. In some examples, a Bidirectional Encoder Representations from Transformers (BERT)-based encoder can be employed to generate the representations. The encoder can further apply clustering operations (e.g., k-means clustering) to the latent representations to enable the encoder to capture long-range temporal relationships between the latent representations and generate labels that can be used by downstream ML models of the TTS model 520 (e.g., a generative language model of the TTS model 520, an audio decoder of the TTS model 520, etc.) for training purposes. In some examples, the audio encoder can alternate between generating predictions of acoustic and prosodic features and clustering, thereby providing progressive improvements of the audio encoder as the audio encoder can learn from previous predictions to better cluster newly learned representations. Thus, the clustering operations included in the training phase of the audio encoder of the TTS model 520 improve the ability of the TTS model 520 to more accurately learn the structure of input audio during the inference phase when generating cloned audio data 522.
[0102] Additionally, the training (or retraining) of the various ML components of the TTS model 520 can be periodic, such as updating the ML model at discrete time intervals (e.g., once a week or once a month), or otherwise updating the ML model. The training data set can be from multiple recorded speech samples (e.g., shorter audio segments), or a particular one of the recorded speech samples (e.g., a longer audio segment). The recorded speech samples used to train the one or more ML models of the TTS model 520 can be obtained by the TTS model in different ways. In one example, the recorded speech samples can be obtained by offline recording, such as by the TTS model requesting a user to speak certain words and capturing audio data corresponding to the spoken words. In this case, the training data set can omit certain data samples that are determined to be outliers, such as noise or recorded speech samples of other users.
[0103] With continued reference to Figure 5 The cloned audio data 522, the reference image 514, and the masked video frames 512 are provided to a video generation model 530. The video generator 530 includes a warping model 532 and a inpainting model 534 to generate warped video frames 536. Similar to the TTS model 520, the video generation model 530 can include one or more ML models to implement one or more aspects of the disclosure. Further details regarding the warping model 532 and the inpainting model 534, and the various ML components therein, are described in further detail below Figure 6 In further detail.
[0104] The video generation model 530 generates a set of warped video frames 536 based on the provided inputs. In particular, each video frame of the generated warped video frames 536 includes a warped mouth region of the user that is synchronized with the cloned audio data 522, while maintaining the features and head pose of the user to be consistent with the consecutive video frames of the template video 454. More specifically, the video generation model 530 performs spatial warping on the feature maps of the reference image 514 and the masked video frames 512 according to the embedded features extracted from the cloned audio data 522 to inpaint the mouth region. The cloned audio data 522 is segmented into a set of segments covering the entire duration of the cloned audio data 522. The video generation model 530 generates a warped video frame for each segment, where the mouth region is warped according to the cloned audio data 522 of the corresponding segment. Thus, the output of the video generation model 530 is a set of warped video frames 536, where each warped video frame includes a warped mouth region of the user that is synchronized with the cloned audio data 522 of the corresponding segment.
[0105] After generating warped video frames 536 using the video generation model 530, the warped video frames 536 are provided to a post-processing module 550. The post-processing module 550 performs operations to stitch and blend the warped mouth regions of the warped video frames 536 into the masked mouth regions of the masked video frames 512. To perform these operations, a sliding window is used. As previously described, the masked video frames 512 are associated with consecutive video frames of the template video 454 in which the mouth region of the user has been masked. Each warped video frame of the warped video frames 536 is blended into a corresponding consecutive video frame of the template video 454 using the sliding window. For example, a first warped mouth region corresponding to a first segment of the cloned audio data 522 from zero to one second is mapped to a first masked video frame, a second warped mouth region corresponding to a second segment of the cloned audio data 522 from one to two seconds is mapped to a second masked video frame, and so on.
[0106] As previously described, the length of the cloned audio data 522 and the corresponding set of segments can be longer than the number of masked video frames 512. Thus, when the number of warped video frames 536 exceeds the number of masked video frames 512, the post-processing module 550 can restart the masked video frames 512 of the template video 454 one or more times to process the entire cloned audio data 522. For example, in a case where there are five thousand warped video frames and only two thousand five hundred masked video frames 512, the masked video frames 512 of the template video 454 will be restarted once (e.g., two full cycles) to process the entire cloned audio data 522 length.
[0107] The post-processing module 550 then applies a smoothing algorithm to the masked video frames 512 to stitch in the warped mouth regions to blend the warped mouth regions into the masked video frames 512 to render a voice cloning avatar with a natural and realistic appearance. Specifically, blending the warped mouth regions from the warped video frames 536 into each masked video frame 512 causes the transition between the inpainted region (e.g., the masked region) and the existing image (e.g., the unmasked portion of the masked video frame 512) to be natural and visually consistent. A variety of mechanisms can be employed to blend the inpainted pixels of the warped mouth regions into the masked video frames 512, such as interpolation, diffusion, via a neural network, and the like. After the post-processing module 550 performs the post-processing operations, the digital content 420 is fully rendered based on the input text 452 using the voice cloning avatar with an appearance and voice that resembles the user.
[0108] Reference is now made to Figure 6 , Figure 6 is a clip of a voice cloning and avatar generation Figure 5 An example graph 600 of the video generation model 530. As with the example graph 400, the graph 600 is a directed acyclic graph (DAG) that represents the architecture of the video generation model 530. The graph 600 includes a plurality of nodes 602, each of which represents a layer of the video generation model 530. The graph 600 includes an input node 604, an output node 606, and a plurality of intermediate nodes 608 between the input node 604 and the output node 606. The input node 604 represents the input to the video generation model 530, which is the warped video frames 536. The output node 606 represents the output of the video generation model 530, which is the post-processed video frames 610. The intermediate nodes 608 represent the layers of the video generation model 530. Figure 5As described, the video generation model 530 includes a warping model 532 and an image inpainting model 534. The warping model 532 receives the masked video frame 512, the reference image 514, and the cloned audio data 522 as inputs. The warping model 532 includes one or more ML components, such as a feature encoder, an audio encoder, and an alignment encoder, which can be used to generate warping features 616. The warping features 616 can represent an embedded representation of one or more features extracted from the various inputs to the warping model 532.
[0109] More specifically, the warping model 532 includes one or more feature encoders, such as feature encoder 602 and feature encoder 606. The feature encoders 602 and 606 can be any suitable type of trained ML model, including but not limited to a convolutional neural network (“CNN”), a recurrent neural network (“RNN”), a transformer, an autoencoder, a variational autoencoder, a generative adversarial network (“GAN”), and the like. The feature encoder 602 receives the masked video frame 512 as input, and the feature encoder 606 receives the reference image 514 as input. Both the feature encoders 602 and 606 are trained ML models that can generate an embedded representation of the masked video frame 512 and an embedded representation of the reference image 514. The embedded representations can be one or more high-dimensional vectors that include quantified components of the user’s appearance in the masked video frame 512 and the reference image 514, effectively encoding the extracted feature sets. For example, the feature encoders 602 and 606 can be trained to perform feature encoding, such as features of the masked video frame 512 and features of the reference image 514, including edges and contours, textures, colors and gradients, geometry, patterns, spatial relationships, symmetry and asymmetry, and motion and changes over time. Different layers of the neural networks used in the feature encoders 602 and 606 can detect these features, with lower layers focusing on more basic features, such as edges, and upper layers capturing more abstract and complex features.
[0110] The warping model 532 also includes an audio encoder 604 to perform feature extraction on the cloned audio data 522. The audio encoder 604 can be any suitable type of trained ML model, including but not limited to a CNN, an RNN, a transformer, an autoencoder, a variational autoencoder, a GAN, and the like. For example, the audio encoder 604 can include one or more trained ML models to generate an embedded representation of the cloned audio data 522. Similar to the discussion above regarding the feature encoders, the embedded representation of the cloned audio data 522 can be one or more high-dimensional vectors that include quantified components of the sounds in the cloned audio data 522, effectively encoding the extracted feature sets.
[0111] For example, the audio encoder 604 can receive the cloned audio data 522 as pre-processed time-domain waveforms. The audio encoder 604, which can be trained, can perform a dimensionality reduction function to convert the high-dimensional raw cloned audio data 522 into a lower-dimensional compact embedded representation, such as a vector. In this process, the audio encoder 604 can extract features such as pitch, harmonics, etc. from the cloned audio data 522. The embedded representation is lower-dimensional relative to the high-dimensional cloned audio data 522, which can be a waveform with high degrees of freedom.
[0112] The embedded representation of the reference image 514 and the embedded representation of the masked video frame 512 are then provided to a concatenation operator 608, and then to an alignment encoder 610. The concatenation operator 608 merges the multiple vector representations in the embedded representations into a single larger embedded representation. Thus, the generated single embedded representation includes the embedded representation of the reference image 514 and the embedded representation of the masked video frame 512. The alignment encoder 610 can include one or more trained ML model configurations to map one or more features in the embedded representation of the masked video frame 512 to one or more features of the embedded representation of the reference image 514. The alignment encoder 610 ensures that the various feature representations are properly aligned (e.g., different head poses, angles, etc. of the reference image 514 are mapped to the masked video frame 512). The output of the alignment encoder 610 is provided to a concatenation operator 612, where the aligned embedded representation is again concatenated with the embedded representation of the cloned audio data 522.
[0113] To generate the warped features 616, the output of the concatenation operator 612 and the embedded representation of the reference image 514 are provided to a warping operator 614. The warping operator 614 performs operations to warp the embedded representation of the reference image 514 according to the output of the concatenation operator 612. For example, the warping operator 614 can perform feature channel specific warping using affine transformations (e.g., de-warping, translation, rotation, and / or scaling) on the embedded representation of the reference image 514 and the output of the concatenation operator 612 to achieve spatial warping. Other techniques can be implemented by the warping operator 614, such as dense optical flow techniques, bilinear interpolation, etc. to achieve spatial warping. Those of ordinary skill in the art will recognize many variations, modifications, and alternatives.
[0114] The morphing features 616 are provided to the image inpainting model 534. The image inpainting model 534 uses the morphing features 616 and performs operations to generate warped mouth regions from the reference images 514 that are synchronized with the cloned audio data 522. More specifically, based on the morphing features 616, the image inpainting model 534 performs image inpainting on the mouth regions of the reference images 514, thereby generating a set of warped video frames 536. The image inpainting model 534 includes one or more ML components, including one or more feature decoders, such as the feature decoder 620, and the image inpainting model 534 includes one or more stitching operators, such as the stitching operator 618. Similar to the description above regarding the morphing model 532, the stitching operator 618 merges multiple vector representations in an embedded representation into a single larger embedded representation. As with the method used by the image inpainting model 534, the stitching operator 618 combines the embedded representation of the masked video frames 512 with the morphing features 616. After being stitched together, its output is provided to the feature decoder 620.
[0115] The feature decoder 620 can be a ML model trained to convert the embedded representation of the morphing features 616 into a transformed embedded representation, i.e., to predict the user’s mouth region topography from a set of acoustic features (e.g., from the cloned audio data 522) and a set of appearance features (e.g., from the reference images 514). For example, the feature decoder 620, as trained, can output the warped video frames 536, where each video frame includes a warped mouth region representing the prediction of the user’s mouth region topography by the video generation model 530 from a given audio signal. The feature decoder 620 can be any suitable type of trained ML model, including but not limited to CNNs, RNNs, transformers, autoencoders, variational autoencoders, GANs, etc. For example, the feature decoder 620 can include convolutional layers that inpaint the pixels of the user’s mouth region in the reference images by the prediction of the mouth region topography from the cloned audio data 522.
[0116] As with the description above regarding the morphing model 532, Figure 5The warped mouth region of each warped video frame 536 is then blended into the masked region of the masked video frame 512 of the template video 454 using post-processing techniques as described. In some examples, the post-processing techniques can include implementing a refiner to enhance details, correct errors, improve resolution, remove or reduce any artifacts that can arise during the blending process. The refiner can be any suitable type of trained machine learning model, including but not limited to a CNN, an autoencoder, or a GAN. In addition to the blended warped mouth region, the post-processing techniques implementing the refiner do not alter the regions of the masked video frame 512 outside of the masked mouth region, thereby enabling the final warped video frame 536 to maintain texture consistency with the original template video frame of the template video 454 while the mouth region is synchronized with the cloned audio data 522. The output rendering is a highly realistic avatar that mimics the user’s voice and appearance from the given input text 452.
[0117] Reference is now made to Figure 7 , Figure 7 is an example of a graphical user interface (GUI) 700 displaying a consent authorization request to access personal data. In accordance with the current disclosure, in some examples, a user can opt-in to use one or more optional AI features provided by a communication platform, such as the chat and video conferencing provider 110 or the chat and video conferencing provider 210. Using these optional AI features can involve providing the user’s personal information to the AI models underlying the AI features. The personal information can include the user’s contacts, calendar, communication records, video or audio streams, recordings of video or audio streams, transcribed text of audio or video conferences, or any other personal information available to the virtual meeting provider. In addition, the audio or video feeds can include the user’s speech, which contains the user’s speaking patterns, pause cadence, word choice, vocal tone, and pitch; the user’s appearance and likeness, which contains facial movements, eye movements, arm or hand movements, and body movements, all of which can be used to provide the optional AI features or to train the underlying AI models.
[0118] Before any such information is captured and used, whether for providing the optional AI features or for providing training data for the underlying AI models, the user can opt-in or opt-out of the access and use of some or all of the user’s personal information. In general, the applicant’s goal is to invest in AI-driven innovation to enhance user experience and productivity while prioritizing trust, security, and privacy. Without the user’s explicit informed consent, the user’s personal information will not be used for any AI features nor used as training data for any AI models. In addition, these optional AI features are off by default - account owners and administrators control whether these AI features are enabled for their accounts, and if enabled, individual users can decide whether to consent to the use of their personal information.
[0119] AsFigure 7 As can be seen, the user has engaged in a video conference and has elected to use the available optional AI functionality. In response, the GUI displays a consent authorization window 710 for the user to interact with. The consent authorization window informs the user that their request can involve the optional AI functionality accessing a variety of different types of personal information of the user. The user then decides whether to grant the optional AI functionality permission, whether to grant general permission or to grant limited permission only. For example, the user can select the option to allow the AI functionality to use personal information to provide the AI functionality but not for use in training the underlying AI model. In addition, the user has the option to select which types of information to share and for what purpose, e.g., to provide the AI functionality or to allow for use in training the underlying AI model.
[0120] Figure 8 is an example process 800 for clip generation with voice cloning and avatars. The discussion of the example process 800 will revolve around the clip generation system shown in Figure 4-6 ; however, any suitable clip generation system with voice cloning and avatars can be used. In addition, while the clip generation system 460 is described with respect to the client device 440 having the communication application 450, any suitable service provider or other computing device can be used in place of the communication device to perform the methods according to the present disclosure.
[0121] At block 802, the clip generation system 460 accesses a template video 454 of the user. The template video 454 includes template audio data 516 and template video data. The template video data includes consecutive frames and the template audio data 516 includes the user’s voice. The clip generation system 460 accesses the data store 430 to retrieve the template video 454. Alternatively, the template video 454 can be retrieved from a third party database. As discussed with respect to Figure 4 , the template video 454 can be relatively short in length (e.g., thirty seconds long) and can include audio and video data of the user speaking while looking at the camera. The template video 454 can be a continuously recorded video, can include a mandatory consent portion including content where the user states their name, a mandatory consent, and a randomly generated password in the recorded video.
[0122] At block 804, the clip generation system 460 receives input text 452. Similar to the template video 454, the clip generation system 460 can receive the input text 452 from the communication application 450. As discussed with respect to Figure 4The input text 452 can be generated using a variety of different mechanisms. For example, the user can receive the input text 452 by typing the input text 452 via a user interface of the communication application 450 (e.g., using a keyboard or touch screen), sending an SMS text message, sending a chat message, sending an email, etc. The input text 452 can also be provided via an uploaded file (e.g., a slide deck, a word processing file, etc.), or in some implementations, the LLM can be integrated with the communication application 450 (or the communication application 450 can otherwise have access to the LLM) to assist the user in generating text, in which case the input text 452 is generated.
[0123] In block 806, the clip generation system 460 generates cloned audio data 522 based on the template audio data 516 and the input text 452 using the TTS model 520. As described with respect to Figure 5 The cloned audio data 522 generated by the TTS model 520 contains synthesized speech features that mimic the user’s speech. For example, the cloned audio data 522 can include synthesized acoustic features 524, including cadence, intonation, energy, etc., and synthesized prosodic features 526 that mimic the high-level energy patterns, pronunciation, etc. of the user’s speech in the template video 454. Additionally, the length of the cloned audio data 522 corresponds to the input text 452 and is based on the synthesized speech features. For example, the TTS model 520 can determine that the user’s speech is relatively slow-paced, and thus, the length of the cloned audio data 522 can be longer to mimic the user’s speech pace.
[0124] In block 808, the clip generation system 460 extracts one or more reference images 514 from consecutive frames of the template video 454. The reference images 514 are provided as input to the video generator model 530 of the clip generation system 460. Using one or more feature encoders of the video generator model, such as the feature encoder 606, the reference images 514 are converted into an embedded representation. After processing by the video generator model 530, the user’s mouth region in the reference images 514 is spatially warped to synchronize the user’s mouth shapes with the cloned audio data 522. Additionally, selecting appropriate reference images 514 can be used to ensure that the reference images 514 have a sufficient variety of poses of the user to capture the user’s body appearance and natural tendencies when speaking. Thus, the reference images 514 can be selected based on the user’s head pose, the user’s mouth shapes, the user’s gaze direction, and the user’s eye features, etc.
[0125] In block 810, the clip generation system 460 provides the cloned audio data 522, the reference images 514, and the consecutive frames of the template video 454 to the video generation model 530. As described with respect to Figure 4-6As described, the successive frames of the template video 454 received by the video generation model 530 have been modified by the pre-processing module 510 such that, for each successive frame, the user’s mouth region in the template video 454 is masked, resulting in a masked video frame 512. The user’s masked mouth region in each masked video frame 512 is replaced with a warped mouth region of a reference image 514 that will be synchronized to the cloned audio data 522 in future processing steps.
[0126] In block 812, the clip generation system 460 segments the length of the cloned audio data 522 into a set of segments associated with predefined time periods using a sliding window. As mentioned previously, the masked video frames 512 are associated with individual video frames of the template video 454. The number of masked video frames 512 is finite and depends on the length of the template video 454, which in some examples is short (e.g., thirty seconds in length). Additionally, the input text 452 used to generate the cloned audio data 522 varies in length. For example, if a user wants to use a virtual avatar with voice cloning 462 to create digital content 420 to deliver a long-form speech or presentation, the length of the cloned audio data 522 can be relatively long compared to the length of the template video 454.
[0127] In block 814, for each segment of the cloned audio data 522, the clip generation system 460 generates one or more warped video frames 536 using the corresponding frame of the template video 454 (e.g., the masked video frame 512). Each warped video frame 536 contains a warped mouth region, where the corresponding warped mouth region is generated using one or more reference images 514 and based on the cloned audio data 522 of the corresponding segment. In other words, for a segment of the cloned audio data 522, the mouth region in the one or more reference images 514 is warped to be synchronized with the cloned audio data 522 from that particular segment. Thus, for the entire length of the cloned audio data 522, one warped video frame (e.g., the warped video frame 536) is generated for each segment of the cloned audio data 522.
[0128] In block 816, the clip generation system 460 blends the warped video frames 536, each with a warped mouth region, into the corresponding frame of the template video 454 to generate an output video (e.g., the digital content 420). For each warped video frame, the warped mouth region is grafted into the masked mouth region of the masked video frame. As with the template video 454, the output video is a video of the user speaking the text of the input text 452, but with the voice of the cloned audio data 522. Figure 4As described above regarding block 812, the number of segments associated with cloned audio data 522 may exceed the number of consecutive frames in template video 454. Using a sliding window, each distorted video frame is blended into the corresponding frame of the template video. If there are not enough template video frames to process every segment of the entire length of cloned audio data 522, the clip generation system 460 restarts template video 454 one or more times to ensure that each distorted video frame 536 is processed. The output digital content 420 contains a virtual avatar with voice clone 462, simulating the user's appearance and voice. Post-processing operations can be performed on distorted video frames 536 to ensure a natural and visually consistent transition between the repaired area (e.g., the masked mouth area in masked video frame 512 of consecutive video frames in template video 454) and the existing image (e.g., the unmasked portion of masked video frame 512).
[0129] Example flow 800 illustrates a clip generation method with voice cloning and virtual avatars. However, not every step in example flow 800 is necessary; sometimes additional steps may need to be added, and sometimes the order of steps may need to be changed. For example, clip generation system 460 may be implemented on a communication platform (not shown) communicatively coupled to client device 440. Alternatively, any other suitable service provider or other computing device may be used instead of client device 440 to perform the method according to this disclosure.
[0130] Figure 9 This is an example of a video editing workflow with voice cloning and virtual avatars for participants during video conferences. Figure 4 In the relevant description, the client device 440 is configured to implement this by executing appropriate program code (e.g., communication application 450). Figure 9 The operations described herein. Software or program code may be stored on a non-transient storage medium (e.g., on a memory device). Figure 9 The processes described herein and described below are intended to be illustrative, not restrictive. Although Figure 9 The process blocks that occur in a specific sequence or order are depicted, but this is not intended to impose limitations. In some alternative embodiments, blocks may be executed in a different order, or some blocks may be executed in parallel. For illustrative purposes, process 900 is described with reference to some examples depicted in the figures. However, other embodiments are also possible.
[0131] In block 902, process 900 involves participants joining a video conference. For example, a client device 440 associated with a participant can initiate a meeting on a chat and video conferencing provider 410 to provide information about... Figure 4The video conferencing features discussed include a user interface that presents participants' videos. Participants may not wish to speak or reveal their appearance during a video conference and may therefore want to use their virtual avatars 462 with voice clones to generate digital content 420 to represent them in the video conference.
[0132] To this end, the communication application 450 can present a user interface that allows participants to select a virtual avatar with a voice clone 462 that has been previously generated and stored in the data repository 430. When a user selects the option to use the virtual avatar 462 with a voice clone, in block 904, the communication application 450 retrieves the virtual avatar with the voice clone 462 from the data repository 430.
[0133] As mentioned above Figure 4 As described, the communication application 450 can implement additional security protocols to prevent participants from maliciously accessing another user's virtual avatar with voice clone 462. Therefore, in block 906, the communication application 450 can verify the identity of the requester of the virtual avatar with voice clone 462. For example, the communication application 450 can require the participant requesting access to the virtual avatar with voice clone 462 (e.g., the requester) to enter or repeat a randomly generated password that the user stated in template video 454 when the virtual avatar with voice clone 462 was initially created. If the password repeated / entered by the requester matches the password stated from template video 454, the communication application 450 grants the requester access to the virtual avatar with voice clone 462. If the password repeated / entered by the requester does not match the password stated from template video 454, the communication application 450 can deny the requester access to the virtual avatar with voice clone 462.
[0134] Alternatively, the communication application 450 may grant access to the virtual avatar with voice clone 462 based on a comparison of one or more biometric features of the requester with one or more biometric features of the user in template video 454. For example, the communication application 450 may capture one or more biometric features of the requester via a camera of client device 440. Biometric features may include facial recognition, iris recognition, retinal recognition, ear shape, voice recognition, etc. These biometric features of the requester may be compared with user biometric features extracted from template video 454. If the requester's biometric features match the user biometric features in template video 454, the communication application 450 may grant the requester access to the virtual avatar with voice clone 462. If the requester's biometric features do not match the user biometric features in template video 454, the communication application 450 may refuse to grant the requester access to the virtual avatar with voice clone 462.
[0135] After the communication application 450 authorizes access to the virtual persona with the voice clone 462, the process 900 proceeds to block 908 to receive input text 452. As discussed above with respect to Figure 4 and Figure 5 The input text 452 can be provided to the communication application 450 using a variety of different mechanisms. For example, a participant in a video conference can enter a text string within a chat window of the communication application 450 indicating that the participant would like a virtual persona with the voice clone 462 to speak on behalf of the participant. Alternatively, the participant can upload a document containing the input text 452, or the participant can prompt an LLM to generate the input text 452.
[0136] After receiving the input text 452, the process 900 proceeds to block 910 to generate digital content 420 using the virtual persona with the voice clone 462 and based on the input text 452. For example, using the clip generation system 460 of the communication application 450, digital content 420 can be generated of the virtual persona with the voice clone 462. The digital content 420 generated by the clip generation system 460 can then be displayed to other participants of the video conference. Thus, and in some examples, the digital content 420 can replace the video of the participant, with the virtual persona with the voice clone 462 speaking on behalf of the participant, and its appearance and voice matching that of the participant. Thus, it can appear that the participant is speaking live in the video conference, when in fact, the virtual persona with the voice clone 462 is speaking on their behalf.
[0137] At block 912, the process 900 involves determining whether the client device 440 should exit the video conference, e.g., the participant decides to leave the video conference, or the host of the video conference has ended the meeting. If not, the process 900 continues to generate digital content 420 using the virtual persona with the voice clone 462, as described by the operations in blocks 908-910 discussed above. If it is determined that the client device 440 should exit the meeting, the process 900 involves disconnecting from the video conference at block 914.
[0138] Referring now to Figure 10 , Figure 10 is an example computing device 1000 suitable for use in an example system or method for clip generation with a voice clone and virtual persona. The example computing device 1000 includes a processor 1010 that communicates with a memory 1020 and other components using one or more communication buses 1002. The processor 1010 is configured to execute processor-executable instructions stored in the memory 1020 to implement one or more methods for clip generation with a voice clone and virtual persona according to different examples, such as discussed above with respect to Figure 8 andFigure 9 The example processes 800 and 900, in part or in whole. In some embodiments, a computing device can include software 1060 for performing one or more of the methods described herein, such as one or more steps of the processes 800 and 900. In this example, the computing device 1000 also includes one or more user input devices 1050, such as a keyboard, mouse, touchscreen, microphone, etc., to accept user input, such as input text 452. The computing device 1000 also includes a display 1040 to provide visual output to a user, such as digital content 420.
[0139] The computing device 1000 also includes a communication interface 1030. In some examples, the communication interface 1030 can enable communication using one or more networks, including a local area network (“LAN”); a wide area network (“WAN”), such as the Internet; a metropolitan area network (“MAN”); a point-to-point or peer-to-peer connection; etc. Communication with other devices can be achieved using any suitable networking protocol. For example, one suitable networking protocol can include the Internet Protocol (“IP”), the Transmission Control Protocol (“TCP”), the User Datagram Protocol (“UDP”), or a combination thereof, such as TCP / IP or UDP / IP.
[0140] While some of the method and system examples herein are described in terms of software executed on various machines, the methods and systems can also be implemented in a specially configured hardware, such as a field-programmable gate array (FPGA) that is specifically configured to perform various methods in accordance with the present disclosure. For example, examples can be implemented in digital electronic circuitry, or in computer hardware, firmware, software, or in combinations of them. In one example, a device can include one or more processors. A processor includes a computer-readable medium, such as a random-access memory (RAM) coupled to the processor. The processor executes computer-executable program instructions, such as one or more computer programs, stored in the memory. Such processors can include a microprocessor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), field-programmable gate arrays (FPGAs), and state machines. Such processors can also include programmable electronic devices such as PLCs, programmable interrupt controllers (PICs), programmable logic devices (PLDs), programmable read-only memories (PROMs), electronically programmable read-only memories (EPROMs or EEPROMs), or other similar devices.
[0141] Such a processor can include or can be in communication with a medium, for example, one or more non-transitory computer-readable mediums, which can store processor-executable instructions that, when executed by the processor, can cause the processor to perform and / or assist the methods described in the present disclosure. Examples of non-transitory computer-readable mediums can include, but are not limited to, electronically, optically, magnetically, or otherwise stored data that can be accessible by a processor (e.g., in a web server). Other examples of non-transitory computer-readable mediums include, but are not limited to, floppy disks, CD-ROMs, DVDs, memory chips, hard drives, RAM, ROM, ASICs, configured processors, all optical media, all magnetic media, or any other medium that can be used to store desired program code means in the form of computer- processor executable instructions. The processor and described processes can be in one or more structures, and can be spread across one or more structures. The processor can include code for executing the methods (or portions thereof) described in the present disclosure.
[0142] The foregoing description of some examples is merely intended to be illustrative and descriptive, and is not intended to be exhaustive or to limit the disclosure to the precise form disclosed. Various modifications, variations, and alternatives can be apparent to those skilled in the art without departing from the spirit and scope of the present disclosure.
[0143] Reference throughout this document to an example or embodiment means that a particular feature, structure, operation, or other characteristic described in connection with the example is included in at least one embodiment of the disclosure. The disclosure is not limited to the particular examples or embodiments described. The phrase “in one example,” “in an example,” “in one embodiment,” or “in an embodiment” as used herein does not necessarily refer to the same example or embodiment, although it can. Any particular feature, structure, operation, or other characteristic described in this specification in connection with an example or embodiment can be combined with other features, structures, operations, or other characteristics described in connection with any of the other examples or embodiments.
[0144] The use of the word “or” in this document covers the inclusive and exclusive OR conditions. In other words, A or B or C includes any or all of the alternative combinations of A, B, and C, as applicable, individually, A only, B only, C only, A and B only, A and C only, B and C only, and A and B and C.
Claims
1. A method comprising: accessing a template video of a user, the template video comprising template audio data and template video data, wherein the template video data comprises consecutive frames and the template audio data comprises user speech; receiving, by a text-to-speech (TTS) model, a text input; generating, using the TTS model, cloned audio data based on the template audio data and the text input, the cloned audio data comprising synthesized speech features that mimic the user speech, wherein the cloned audio data has a length; extracting one or more reference images from the consecutive frames; providing, to a video generation model, the cloned audio data, the one or more reference images, and the consecutive frames of the template video, wherein for each consecutive frame of the template video, a user mouth region in the template video is masked; segmenting, using a sliding window, the length of the cloned audio data into a set of segments associated with predefined time periods; for each segment in the set of segments, generating one or more warped video frames using a corresponding frame from the template video, each warped video frame having a warped mouth region, wherein the corresponding warped mouth region is generated using the one or more reference images and based on the cloned audio data of the corresponding segment; and for each warped video frame, blending the warped mouth region into the corresponding frame of the template video to generate an output video.
2. The method of claim 1, wherein the template video is shorter than the length of the cloned audio data, and generating the warped video frames further comprises: restarting the template video one or more times to generate the warped video frames corresponding to the length of the cloned audio data.
3. The method of claim 1, further comprising: labeling the text input with metadata corresponding to a user emotion; and generating the warped video frames based on the metadata.
4. The method of claim 1, wherein the one or more reference images are extracted based on one or more of a head pose of the user, a mouth shape of the user, a gaze direction of the user, and eye features of the user.
5. The method of claim 1, wherein the template video has a length of at least thirty seconds and is a single continuous video recording.
6. The method of claim 1, wherein the template audio data of the template video is associated with a predefined script read by the user, wherein the predefined script contains a mandatory consent portion that contains a name of the user, a consent of the user, and a recitation of a randomly generated password.
7. The method of claim 6, further comprising: verifying access to the output video by: receiving an input password; comparing the input password to the randomly generated password in the template video to determine a match; and determining to grant access to the output video when there is a match.
8. The method of claim 1, further comprising: verifying access to the output video by comparing one or more biometric features of a requester of the output video to one or more biometric features of the user in the template video; and granting access to the output video when the one or more biometric features of the requester of the output video match the one or more biometric features of the user in the template video.
9. A system comprising: a communication interface; a non-transitory computer-readable medium; and one or more processors communicatively coupled with the communication interface and the non-transitory computer-readable medium, the one or more processors configured to execute processor-executable instructions stored in the non-transitory computer-readable medium to: access a template video of a user, the template video comprising template audio data and template video data, wherein the template video data comprises consecutive frames and the template audio data comprises a user voice; receive a text input by a text-to-speech (TTS) model; generate cloned audio data using the TTS model based on the template audio data and the text input, the cloned audio data comprising synthesized speech features that mimic the user voice, wherein the cloned audio data has a length; extract one or more reference images from the consecutive frames; provide the cloned audio data, the one or more reference images, and the consecutive frames of the template video to a video generation model, wherein for each consecutive frame of the template video, a user mouth region in the template video is masked; segment the length of the cloned audio data into a set of segments associated with predefined time periods using a sliding window; for each segment in the set of segments, generate one or more warped video frames using a respective frame from the template video, each warped video frame having a warped mouth region, wherein the respective warped mouth region is generated using the one or more reference images and based on the cloned audio data of the respective segment; and for each warped video frame, blend the warped mouth region into the respective frame of the template video to generate an output video.
10. The system of claim 9, wherein the template video is shorter than the length of the cloned audio data, and wherein the one or more processors are configured to execute further processor-executable instructions stored in the non-transitory computer-readable medium to: restart the template video one or more times to generate warped video frames corresponding to the length of the cloned audio data.
11. The system of claim 9, wherein the one or more processors are configured to execute further processor-executable instructions stored in the non-transitory computer-readable medium to: label the text input with metadata corresponding to a user emotion; and generate the warped video frames based on the metadata.
12. The system of claim 9, wherein the one or more reference images are extracted based on one or more of a head pose of the user, a mouth shape of the user, a gaze direction of the user, and eye features of the user.
13. The system of claim 9, wherein the template video has a length of at least thirty seconds and is a single continuous video recording.
14. The system of claim 9, wherein the template audio data of the template video is associated with a predefined script read by the user, wherein the predefined script contains a mandatory consent portion that contains a name of the user, a consent of the user, and a recitation of a randomly generated password.
15. The system of claim 14, wherein the one or more processors are configured to execute further processor-executable instructions stored in the non-transitory computer-readable medium to: verify access rights to the output video by: receiving an input password; and comparing the input password to the randomly generated password in the template video to determine a match; and granting access to the output video upon determining the match.
16. The system of claim 9, wherein the one or more processors are configured to execute further processor-executable instructions stored in the non-transitory computer-readable medium to: verify access to the output video by comparing one or more biometrics of a requestor of the output video to one or more biometrics of the user in the template video; and grant access to the output video when the one or more biometrics of the requestor of the output video match the one or more biometrics of the user in the template video.
17. A non-transitory computer-readable medium comprising processor-executable instructions configured to cause one or more processors to: access a template video of a user, the template video comprising template audio data and template video data, wherein the template video data comprises consecutive frames and the template audio data comprises speech of the user; receive text input from a text-to-speech (TTS) model; generate cloned audio data using the TTS model based on the template audio data and the text input, the cloned audio data comprising synthesized speech features that mimic the speech of the user, wherein the cloned audio data has a length; extract one or more reference images from the consecutive frames; provide the cloned audio data, the one or more reference images, and the consecutive frames of the template video to a video generation model, wherein for each consecutive frame of the template video, a mouth region of the user in the template video is masked; segment the length of the cloned audio data into a set of segments associated with predefined time periods using a sliding window; for each segment in the set of segments, generate one or more warped video frames using a corresponding frame from the template video, each warped video frame having a warped mouth region, wherein the corresponding warped mouth region is generated using the one or more reference images and based on the cloned audio data for the corresponding segment; and for each warped video frame, blend the warped mouth region into the corresponding frame of the template video to generate an output video.
18. The non-transitory computer-readable medium of claim 17, wherein the template video is shorter than the length of the cloned audio data, further comprising processor-executable instructions configured to cause one or more processors to: restart the template video one or more times to generate warped video frames corresponding to the length of the cloned audio data.
19. The non-transitory computer-readable medium of claim 17, wherein the template audio data of the template video is associated with a predefined script read by the user, wherein the predefined script contains a mandatory consent portion that contains a name of the user, a consent of the user, and a recitation of a randomly generated password, further comprising processor-executable instructions configured to cause one or more processors to: verify access to the output video by: receiving an input password; and comparing the input password to the randomly generated password in the template video to determine a match; and granting access to the output video upon a match being determined.
20. The non-transitory computer-readable medium of claim 17, further comprising processor-executable instructions configured to cause one or more processors to: verify access to the output video by comparing one or more biometric features of a requestor of the output video to one or more biometric features of a user in a template video; and grant access to the output video when the one or more biometric features of the requestor of the output video match the one or more biometric features of the user in the template video.