Model generation, acquisition methods, video generation methods, equipment, and media
The method enhances user experience by enabling video dialogue using virtual character interaction models, addressing the limitations of auditory-only voice assistants.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- ZTE CORP
- Filing Date
- 2024-01-18
- Publication Date
- 2026-04-10
AI Technical Summary
Current voice assistants lack the capability to enhance user experience beyond auditory dimension voice mode, limiting human-computer interaction.
A method for obtaining and generating a virtual character interaction model using facial expressions and audio data, enabling video dialogue through a model management server, model training server, and terminal devices, which includes training and distributing virtual character dialogue models.
Expands human-computer interaction to video dialogue, improving user experience by providing lively facial expressions and reducing hardware requirements on terminal devices.
Smart Images

Figure 2026510933000001_ABST
Abstract
Description
Technical Field
[0001] (Cross - reference to related applications) This application claims the priority of Chinese Patent Application No. 202310286640.2 filed on March 15, 2023, and incorporates the content of the Chinese patent application herein by reference in its entirety.
[0002] This application relates to the field of information technology. Specifically, it relates to a method for obtaining a virtual character interaction model, a method for generating a virtual character interaction model, a method for generating a video, an electronic device, and a computer - readable medium.
Background Art
[0003] With the development of artificial intelligence technology, many smart terminals are equipped with voice assistants. Through voice assistants, the interaction between humans and computers can be realized. However, the current voice assistants only stay in the auditory - dimension voice mode, are single, and cannot further improve the user experience.
Summary of the Invention
Problems to be Solved by the Invention
[0004] Embodiments of this application provide a method for obtaining a virtual character interaction model, a method for generating a virtual character interaction model, a method for obtaining a virtual character interaction model, a method for generating a video, an electronic device, and a computer - readable medium.
Means for Solving the Problems
[0005] As a first aspect of the present invention, a method for acquiring a virtual character dialogue model is provided, which is used in a model management server and includes the steps of: transmitting at least one model training data to the model training server, wherein the model training data is associated with the facial expression of a predetermined character and the audio data of a predetermined character, and a model training sample can be obtained based on the model training data, and the training sample includes the correspondence between the facial expression of the predetermined character and the audio data; receiving a virtual character dialogue model corresponding to various model training data; and, in response to a download request, transmitting a virtual character dialogue model corresponding to the download request to a terminal corresponding to the download request.
[0006] A second aspect of the present invention provides a method for generating a virtual character dialogue model, comprising the steps of: receiving at least one model training data used in a model training server, wherein the model training data is associated with the facial expressions of a predetermined character and the audio data of a predetermined character, and a model training sample can be obtained based on the model training data, the training sample includes a correspondence between the facial expressions of a predetermined character and the audio data; inputting each of the training samples into an initial model and training the initial model to obtain a virtual character dialogue model corresponding to each of the training samples; and transmitting each of the virtual character dialogue models to a model management server.
[0007] A third aspect of the present invention provides a method for acquiring a virtual character dialogue model, which is used in a terminal and includes the steps of: sending a download request to a model management server; and receiving a virtual character dialogue model, wherein the virtual character dialogue model is a model acquired by a model training server training an initial model with training samples, and the training samples include a correspondence between the facial expressions of a predetermined character and audio data.
[0008] A fourth aspect of the present invention provides a video generation method used in a terminal, comprising the steps of: determining a virtual character dialogue model, wherein the virtual character dialogue model is a model obtained by a model training server training an initial model with training samples, and the training samples include a correspondence between the facial expressions of a predetermined character and audio data; and inputting drive information to the determined virtual character dialogue model and obtaining an output video for the drive information, wherein the character in the output video is the predetermined character.
[0009] A fifth aspect of the present invention provides an electronic device comprising one or more processors and a memory that stores one or more programs and, when the one or more programs are executed by the one or more processors, implements at least one of the following methods: the acquisition method described in the first aspect, the generation method described in the second aspect, the acquisition method described in the third aspect, and the video generation method described in the fourth aspect.
[0010] As a sixth aspect of the present invention, a computer-readable medium is provided that stores a computer program and, when the program is executed by a processor, implements at least one of the acquisition method described in the first aspect, the generation method described in the second aspect, the acquisition method described in the third aspect, and the video generation method described in the fourth aspect.
[0011] The acquisition method described above may be performed by a model management server. The model management server sends training samples to a model training server, which uses the training samples to train an initial model and acquires a virtual character dialogue model. After the model management server acquires the virtual character dialogue model, it can store it locally. After receiving a download request sent by a terminal, the corresponding virtual character dialogue model is sent to the corresponding terminal based on the download request. [Brief explanation of the drawing]
[0012] [Figure 1] This is a flowchart of an embodiment of the acquisition method provided by the first aspect of the present application. [Figure 2] This is a schematic diagram illustrating the application scenarios of the method provided in this application. [Figure 3] This is a flowchart of another embodiment of the acquisition method provided by the first aspect of the present application. [Figure 4] This is a flowchart of yet another embodiment of the acquisition method provided by the first aspect of the present application. [Figure 5] This is a flowchart of yet another embodiment of the acquisition method provided by the first aspect of the present application. [Figure 6] This is a flowchart of yet another embodiment of the acquisition method provided by the first aspect of the present application. [Figure 7] This is a flowchart of an embodiment of the method for generating a virtual character dialogue model provided by the second aspect of the present application. [Figure 8] This is a flowchart of an embodiment of the method for obtaining a virtual character dialogue model provided by the third aspect of the present application. [Figure 9] This is a flowchart of another embodiment of the method for obtaining a virtual character dialogue model provided by the third aspect of the present application. [Figure 10] This is a flowchart of an embodiment of the video generation method provided by the fourth aspect of the present application. [Figure 11]It is a module schematic diagram of an embodiment of an electronic device provided by the present application. [Figure 12] It is a schematic diagram of a computer-readable medium provided by the present application. [Figure 13] It is a module schematic diagram of an embodiment of a model management server provided by the present application. [Figure 14] It is a module schematic diagram of an embodiment of a model training server provided by the present application. [Figure 15] It is a schematic diagram of an embodiment of a terminal provided by the present application. [Figure 16] It is a flowchart of a model training method provided by Embodiment 1 of the present application. [Figure 17] It is a flowchart for completing the interaction between the terminal of the model management server provided by Embodiment 2 of the present application and the push of the virtual character interaction model. [Figure 18] It is a flowchart for the selection and download of a model performed by the terminal provided by Embodiment 3 of the present application. [Figure 19] It is a flowchart for training a model provided by Embodiment 4 of the present application. [Figure 20] It is a flowchart for performing model application installation via the terminal provided by Embodiment 5 of the present application. [Figure 21] It is a schematic diagram of the usage method of the virtual character interaction model in the "Z Voice Assistant" scene provided by Embodiment 6 of the present application. [Figure 22] It is a schematic diagram of the usage method of the virtual character interaction model in the "Specific Incoming Call Application" scene provided by Embodiment 7 of the present application. [Figure 23] It is a schematic diagram of the usage method of the virtual character interaction model in the "Specific Voice Call" scene provided by Embodiment 8 of the present application.
Embodiments for Carrying Out the Invention
[0013] To enable those skilled in the art to better understand the technical proposal of this application, the following description, accompanied by drawings, will detail the method for acquiring a virtual character dialogue model, the method for generating a virtual character dialogue model, the method for generating a video, and the electronic equipment and computer-readable media provided by this application.
[0014] The following text provides a more detailed description of exemplary embodiments with reference to the drawings, but these exemplary embodiments may be embodied in different forms and should not be construed as being limited to the embodiments described herein. Rather, these embodiments are provided to make the present application more detailed and complete, and to enable those skilled in the art to fully understand the scope of the present application.
[0015] Unless contradictory, each embodiment and each feature in the embodiment can be combined with one another.
[0016] As used herein, the terms "and / or" include any and all combinations of one or more related enumerated items.
[0017] The terms used herein are for the sole purpose of describing specific embodiments and are not intended to limit the application. The singular terms “one” and “the said” as used herein are also intended to include the plural unless the context otherwise makes clear. Furthermore, where the terms “including” and / or “consisting of” are used herein, they refer to the presence of such features, wholes, steps, operations, elements and / or assemblies, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, assemblies and / or groups thereof.
[0018] Unless otherwise specified, all terms used herein (including technical and scientific terms) have the same meaning as those generally understood by those skilled in the art. Furthermore, these terms, as defined in general dictionaries, should be interpreted as having meanings consistent with their meanings in the context of the relevant art and the present application, and should not be interpreted as having idealized or overly formal meanings unless explicitly limited herein.
[0019] As a first aspect of this invention, a method for acquiring a virtual character dialogue model is provided. As shown in Figure 1, the acquisition method includes the following steps S110 to S130.
[0020] In step S110, at least one model training data is sent to the model training server, and the model training data is associated with the facial expression of a predetermined character and the audio data of a predetermined character.
[0021] In step S120, a virtual character dialogue model corresponding to various model training data is received.
[0022] In step S130, in response to a download request, a virtual character dialogue model corresponding to the download request is sent to the terminal corresponding to the download request.
[0023] This invention provides the system shown in Figure 2, which comprises a model management server, a model training server, and a terminal.
[0024] The method for obtaining a virtual character dialogue model provided by the first aspect of this application may be performed by a model management server. The model management server sends model training data to a model training server, which uses the model training data to train an initial model and obtain a virtual character dialogue model. After the current model management server receives the virtual character dialogue model, it can store it locally. After receiving a download request sent by a terminal, the corresponding virtual character dialogue model is further sent to the corresponding terminal based on the download request.
[0025] As described above, in this application, the model training data is associated with the facial expressions (including the shape of the mouth) of a predetermined character and the audio data of the predetermined character. In an optional embodiment, the model training data may include training samples, which include the correspondence between the facial expressions of a predetermined character and the audio data. The facial expressions are, for example, facial expressions corresponding to phrases expressing happiness, facial expressions corresponding to phrases expressing surprise, and facial expressions corresponding to phrases expressing doubt.
[0026] In another optional embodiment, the model training data may be video data of a predetermined character. In such a case, the model training server may pre-process the model training data in order to obtain the training samples.
[0027] Since the data used to train the virtual character dialogue model is associated with the facial expressions and audio data of a given character, the video output by the virtual character dialogue model will also correspond to the facial expressions and audio data. After the terminal receives the drive audio and inputs it to the virtual character dialogue model, the system can obtain video information with lively facial expressions of the character, improving the user experience.
[0028] After the terminal receives the virtual character dialogue model, it can input drive information (for example, one of voice information, text information, or video information) into the virtual character dialogue model and obtain a corresponding output video. The character in the output video is the predetermined character. Therefore, the method provided by this application expands human-computer interaction beyond voice dialogue to video dialogue, thereby improving the user experience.
[0029] By using a model management server and a model training server to allow terminals to acquire virtual character dialogue models, the hardware requirements on terminal devices can be reduced, which is advantageous for the widespread adoption of virtual character dialogue models.
[0030] This application does not particularly limit the type of model training data transmitted to the model training server. Characters in training samples of the same type are the same, and characters in different training samples are different.
[0031] When model training data for one character is sent to the model training server, the model training server can obtain a virtual character interaction model for that character. When model training data for multiple characters is sent to the model training server, the model training server can obtain multiple virtual character interaction models corresponding to each of the multiple characters.
[0032] After storing multiple virtual character interaction models in the current model management server, terminal users can select the corresponding virtual character according to their preference, thereby further improving the user experience.
[0033] This application does not specifically limit the definition of "data for model training." For example, data for model training may be video material containing facial expressions of a given character and corresponding audio data. In other words, data for model training may include the final video file of a given character. In this application, the "final video file" should provide a sufficient amount of data to satisfy the parameter requirements of the training model. For example, the length, clarity, and amount of audio in the final video file should satisfy the requirements of model training.
[0034] Optionally, the "final video file" is a video file obtained by processing the initial video file and that satisfies the parameter requirements of the training model. The model management server sends the final video file containing the specified characters to the model training server, which processes the received final video file and obtains training samples.
[0035] To reduce the computational load on the model training server, the model management server may optionally process the final video file corresponding to each character directly and obtain training samples for each character. In other words, in another optional embodiment, the data for model training includes training samples. As described above, the training samples include the correspondence between the facial expressions of a given character obtained from the final video file of that character and the audio data. Accordingly, as shown in Figure 3, before sending the data for at least one model training to the model training server (i.e., before step S110), the acquisition method may further include step S106, which processes the final video file of the given character and obtains the training samples.
[0036] In an optional embodiment, as shown in Figure 4, the acquisition method may further include step S140 of placing identification parameter information in each of the received virtual character dialogue models.
[0037] Step S140 allows us to distinguish between each of the virtual character dialogue models corresponding to different characters.
[0038] In this application, the virtual character dialogue model transmitted in step S130 may be a virtual character dialogue model in which identification parameter information is placed.
[0039] This application does not particularly limit the identification parameter information, and it is sufficient if it can distinguish each virtual character dialogue model corresponding to a different character. Optionally, the identification parameter information of the virtual character dialogue model includes at least one of the following parameter information: a model identifier, a model nickname, model introduction information, model attribution information, information characterizing whether or not to make it public, cost information, application scene information, and information characterizing whether or not to push it to the user.
[0040] Specifically, it is as follows:
[0041] Model Identifier (ID): Uniquely identifies the trained virtual character dialogue model. In this example, the format is defined as M-XXXXXX.
[0042] Model Nickname: This is the model's name, and can also be the name of the virtual character. It facilitates communication, interaction, and memory.
[0043] Model Introduction: This section provides an introduction and explanation of the virtual character dialogue model. It is intended to facilitate administrators' understanding of the model information before deployment or before users download it. The introduction may include animations, videos, images, text, or a combination of these.
[0044] Model Attribution Information: Uniquely identifies the copyright holder of the model. Only the copyright holder of the model, in addition to the background administrator, can modify or delete the model on the server. In this embodiment, the format is defined as O-XXXXXX.
[0045] Information identifying whether to make something public: This defines whether the model is publicly available to all users. The options are: make it public / make it partially public / do not make it public. Selecting "make it public" means the model is public to all users. Selecting "make it partially public" requires further specification of which users it will be made public to. Selecting "do not make it public" means the model will only be used by its owners, which generally refers to a model trained by an individual user.
[0046] Cost Information: You can set whether or not model downloads are paid, and how much the cost is. If it's free, the cost will be set to 0.
[0047] Applicable Scene Information: You can configure which specific application or scene the model will be used for. If targeting a particular application, you will need to specify the application ID information. If targeting a specific scene, you will need to specify the number information, such as triggering for a particular incoming phone number.
[0048] Information characterizing whether or not to push to the user: The server can be configured to either spontaneously push this model to the user or not. If "yes," the server will spontaneously send a message to the user and push this model. If "no," the server will not spontaneously push this model to the user.
[0049] Table 1 illustrates four virtual character dialogue models in which identification parameter information is placed.
[0050] [Table 1]
[0051] Regarding the first model in Table 1, the ID is M-000001, the nickname is Xiaojia (Little Jia), and she is a famous female live streamer on a certain Douyin (a short video social software) who has already obtained the right of portrait. The model production company or the copyright owner is Company Z, and the unique attribution ID assigned by the system is O-000001. This model indicates that in addition to the background administrator, only Company Z has the right to modify or delete this model. The selection of the model is public to all, and users who have registered in this system but have not downloaded and installed it can all download it. The model fee is set to 0, indicating that users can download and use it for free. Since the model scene is defined only for the Z voice assistant, it can only be applied to the Z voice assistant after the model is downloaded. The setting of whether to send a push is "Yes", and the system will automatically push this model to users who meet the conditions.
[0052] [[ID=z3]] Regarding the second model in Table 1, the ID is M-000002, the nickname is Xiaomei (Little Mei), and she is the female protagonist of a certain TV drama who has already obtained the right of portrait. The model production company or the copyright owner is Company Z, and the unique attribution ID assigned by the system is O-000001. This model indicates that in addition to the background administrator, only Company Z has the right to modify or delete this model. The selection of the model is public to all, and users who have registered in this system but have not downloaded and installed it can all download it. The model fee is set to 5, indicating that a payment of 5 CNY is required for downloading this model. Since the model scene is defined as general, the user can define the usage scene after the model is downloaded. The setting of whether to send a push is "No", and the system will not automatically push this model to users.
[0053] JPEG2026510933000003.jpg59170
[0054] Regarding the fourth model in Table 1, the ID is M-000004, the nickname is Xiaozhu (Little Pig), the model introduction is a personal customization, the model producer or copyright owner is User A, and the unique attribution ID assigned by the system is O-000003. This model indicates that, in addition to the background administrator, only User A has the right to modify or delete this model. Since the model selection is "not public", only User A can view it. The model fee is set to 0, indicating that users can download and use it for free. Since the model scene is defined as general-purpose, the user can define the usage scene after downloading the model. The setting of whether to send a push is "yes", and the system will automatically push this model to User A. Here, only User A meets the push conditions.
[0055] As described above, in this application, after receiving the virtual character interaction model sent by the model training server, the virtual character interaction model can be downloaded by the terminal.
[0056] This application does not particularly limit how to obtain the initial video file. It may be directly uploaded to the model management server by the user or the administrator, or it may be uploaded to the model management server by the user via the terminal. Accordingly, as shown in FIG. 5, before step S110, the acquisition method may include step S102 of receiving the final video file.
[0057] This application does not particularly limit the source of the "final video file".
[0058] As an optional embodiment, it is possible to receive the final video file sent by the terminal device. Such a solution corresponds to "character customization". That is, when a user desires to obtain a virtual character interaction model of a certain character, the terminal device can be used to send the final video file of the character to the model management server.
[0059] In this embodiment, the terminal is allowed to upload an initial video file, which enables customization of the virtual character dialogue model and further improves the user experience.
[0060] Of course, the present invention is not limited thereto, and in another arbitrary embodiment, the administrator of the model management server may directly upload a video file of a predetermined character to the model management server, and the model management server may process it to obtain the data for model training.
[0061] In the acquisition method provided by this application, the current model management server only allows downloads from terminals registered with the current model management server. Therefore, after receiving the initial video file sent by the terminal, authentication is performed on the terminal, and step S106 is executed only if the authentication is successful. In other words, in an optional embodiment, the acquisition method may further include performing authentication on the terminal between step S102 and step S106.
[0062] To allow terminals to easily know when to download virtual character dialogue models, optionally, after the step of receiving virtual character dialogue models corresponding to each training model, as shown in Figure 6, the acquisition method may further include step S122 of sending a notification message to each terminal that transmits the final video file of the predetermined character.
[0063] After receiving the corresponding notification message, the device can send a download request to the model management server.
[0064] In this application, the model management server can establish a mapping relationship between the model name and the terminal's identification information. When a terminal sends a download request, terminal identifier information is attached. After receiving the download request from within the terminal, the model management server analyzes the download request, identifies the terminal's identifier information, determines the corresponding virtual character dialogue model, and sends the virtual character dialogue model to the terminal.
[0065] A second aspect of the present invention provides a method for generating a virtual character dialogue model. As shown in Figure 7, the generation method includes the following steps S210 to S230.
[0066] In step S210, at least one model training data is received, and the model training data is associated with the facial expression of a predetermined character and the audio data of a predetermined character.
[0067] In step S220, each training sample based on various model training data is input to the initial model, and the initial model is trained to acquire each virtual character dialogue model corresponding to each of the training samples, the training samples include the correspondence between the facial expressions of a predetermined character obtained from the final video file of the predetermined character and audio data.
[0068] In step S230, each of the virtual character dialogue models is sent to the model management server.
[0069] This generation method may be performed by the model training server. The model management server sends data for model training to the model training server, and the model training server trains an initial model using training samples based on the model training data to obtain a virtual character dialogue model. After the model management server receives the virtual character dialogue model, it can store it locally. After receiving a download request sent by a terminal, the model management server further sends the corresponding virtual character dialogue model to the corresponding terminal based on the download request.
[0070] As described above, in this application, the training sample includes a correspondence between the facial expressions of a predetermined character and audio data. For example, facial expressions corresponding to phrases expressing happiness, facial expressions corresponding to phrases expressing surprise, and facial expressions corresponding to phrases expressing doubt. Therefore, after the terminal inputs the received driving audio into the virtual character dialogue model, it is possible to obtain video information with lively facial expressions of the character, thereby improving the user experience.
[0071] After the terminal receives the virtual character dialogue model, it can input drive information (e.g., voice information or text information) into the virtual character dialogue model and obtain a corresponding output video. The character in the output video is the predetermined character. Therefore, the method provided by this application expands human-computer interaction beyond voice dialogue to video dialogue, thereby improving the user experience.
[0072] By using a model management server and a model training server to allow terminals to acquire virtual character dialogue models, the hardware load during terminal operation can be reduced, while also reducing the hardware requirements on the terminals. This is advantageous for the widespread adoption of virtual character dialogue models.
[0073] This application does not limit how training samples are used to train the initial model. For example, a virtual character dialogue model can be obtained using Audio Driven Neural Radiance Fields (AD-NeRF) technology.
[0074] When using AD-NeRF, the final video file only requires 3 to 5 minutes of audio spoken by the character. By simply providing new audio material, the virtual character dialogue model can be driven to output a video that matches the character.
[0075] This application does not particularly limit the voice tone of the character output by the virtual character dialogue model. For example, the voice tone of the character output by the virtual character dialogue model may be similar to the voice tone in the final video file, other computer-generated sounds may be used as the voice of the virtual character, or other sample sounds may be used as the voice of the virtual character.
[0076] As described above, the data for model training may be the "final video file" or it may be a "training sample." When the data for model training is a training sample, the model training server does not need to perform the step of "pre-processing the final video file in order to obtain the training sample."
[0077] If the data for model training is the final video file of a predetermined character, the generation method may further include processing the final video file between step S210 and step S220 to obtain the corresponding training sample.
[0078] A third aspect of the present invention provides a method for acquiring a virtual character dialogue model. As shown in Figure 8, the acquisition method includes the following steps S310 to S320.
[0079] In step S310, a download request is sent to the model management server.
[0080] In step S320, a virtual character dialogue model is received.
[0081] Here, the virtual character dialogue model is a model obtained by training an initial model using training samples on a model training server, and the training samples include a correspondence between the facial expressions of a predetermined character and audio data.
[0082] The above acquisition method may be performed by a terminal (i.e., terminal device). The virtual character dialogue model is trained by the model training server, acquired, and sent to the model management server.
[0083] The model management server sends data for model training to the model training server, which trains the initial model using training samples based on the model training data to obtain a virtual character interaction model. After the model management server receives the virtual character interaction model, it can store it locally. The terminal sends a download request to the model management server, and after the model management server receives the download request sent by the terminal, it further sends the corresponding virtual character interaction model to the corresponding terminal based on the download request.
[0084] In this application, the training sample includes a correspondence between the facial expressions of a predetermined character and audio data. For example, facial expressions corresponding to phrases expressing happiness, facial expressions corresponding to phrases expressing surprise, and facial expressions corresponding to phrases expressing doubt. Therefore, after the terminal inputs the received driving audio into the virtual character dialogue model, it is possible to obtain lively video information of the character's facial expressions, improving the user experience.
[0085] After the terminal receives the virtual character dialogue model, it inputs drive information (e.g., voice information or text information) into the virtual character dialogue model and obtains a corresponding output video. The character in the output video is the predetermined character. Therefore, the method provided by this application expands human-computer interaction beyond voice dialogue to video dialogue, thereby improving the user experience.
[0086] By using a model management server and a model training server to allow terminals to acquire virtual character dialogue models, the hardware load during terminal operation can be reduced, while also reducing the hardware requirements on the terminals. This is advantageous for the widespread adoption of virtual character dialogue models.
[0087] In this application, prior to the step of sending a download request to the model management server, the acquisition method may further include step S300, as shown in Figure 9, of sending the final video file of a predetermined character to the management server.
[0088] Step S300 enables character customization, improving the user experience.
[0089] To further enhance the user experience, different virtual dialogue models can be used for different application scenarios. Therefore, in step S310, multiple download requests can be sent to the model management server, and different characters will correspond to different download requests. Accordingly, in step S320, multiple virtual character dialogue models can be received.
[0090] Naturally, this invention is not limited to this, and step S310 can be executed multiple times, with each download request sent at different times corresponding to a different character. Therefore, after multiple downloads, different virtual dialogue models for multiple characters are stored in the current terminal. The virtual character dialogue models for different characters can correspond to different application scenarios. The user can select the appropriate virtual character dialogue model according to the application scenario.
[0091] In an optional embodiment, the virtual character interaction model may have a scene parameter in which a model training server is located.
[0092] In another optional embodiment, the model training server may not perform scene placement for the virtual model, and the user may place scene parameters for the virtual character interaction model via a terminal.
[0093] In an optional embodiment, a correspondence can be established between the name of a virtual character, a model identifier (e.g., model ID, model nickname), and a model storage address.
[0094] Different virtual character interaction models can be activated in different scenes, further enhancing the user experience.
[0095] A fourth aspect of the present invention provides a video generation method. As shown in Figure 10, the video generation method includes the following steps S410 and S420.
[0096] In step S410, a virtual character dialogue model is determined, which is a model acquired by a model training server training an initial model using training samples, and the training samples include a correspondence between the facial expressions of a predetermined character and audio data.
[0097] In step S420, drive information is input to the confirmed virtual character dialogue model, an output video corresponding to the drive information is obtained, and the character in the output video is the predetermined character.
[0098] The above video generation method may be performed by a terminal. The virtual character dialogue model is a model obtained by training an initial model using training samples on a model training server, and the training samples include the correspondence between the facial expressions of a predetermined character obtained from the final video file of the predetermined character and audio data, and the character in the output video is the predetermined character.
[0099] By inputting drive information into a virtual character dialogue model and obtaining the final video file corresponding to the drive information, video dialogue with the terminal can be realized, improving the user experience.
[0100] In this application, the virtual character dialogue model is acquired by the acquisition method provided in the third aspect of this application. Therefore, both the initial video processing and model training are all performed locally on the terminal, thereby reducing the operational load on the terminal hardware and improving dialogue efficiency.
[0101] As described above, the terminal's local memory can store multiple virtual character dialogue models that match several different application scenes. In step S410, the confirmed virtual character dialogue model matches the target application scene.
[0102] In this application, different application scenarios correspond to different driving information. For example, in the application scenario of a "voice assistant," the corresponding driving information is a specific driving audio. For example, in the case of Siri, the driving information is the audio "hey siri." After receiving the driving audio, the voice assistant application inputs the driving audio into the corresponding virtual character dialogue model in order to obtain the final video file.
[0103] Furthermore, if the trigger information is, for example, "incoming call audio," the incoming call application can acquire the incoming call information and trigger the corresponding virtual character dialogue model. In this case, the audio output by the character in the final video file is the voice of the caller. In such application scenarios, video calls can be realized on the terminal side using only audio information transmitted through the channel, which not only reduces the occupation of the bandgap during the call but also improves the user experience.
[0104] A fifth aspect of the present invention is provided: an electronic device. As shown in Figure 11, the electronic device includes one or more processors 101 and one or more programs stored therein, and when the one or more programs are executed by the one or more processors 101, The acquisition method described in the first aspect of this application, The generation method described in the second aspect of this application, The acquisition method described in the third aspect of this application, The present invention comprises a video generation method described in the fourth aspect of the present invention, and a memory 102 that enables one or more processors to implement at least one of the methods described therein.
[0105] Optionally, the electronic device may further include one or more I / O interfaces 103, which are connected between the processor and memory and arranged to enable information interaction between the processor and memory.
[0106] Here, the processor 101 is a device having data processing capabilities, including but not limited to a central processing unit (CPU), the memory 102 is a device having data storage capabilities, including but not limited to random access memory (RAM, more specifically SDRAM, DDR, etc.), read-only memory (ROM), electrically erasable and programmable read-only memory (EEPROM), and flash memory (FLASH), the I / O interface (read / write interface) 103 is connected between the processor 101 and the memory 102 to enable information interaction between the processor 101 and the memory 102, including but not limited to a data bus.
[0107] In some embodiments, the processor 101, memory 102, and I / O interface 103 are interconnected via a bus 104 and further connected to other components of the computer equipment.
[0108] Accordingly, the electronic device may be the model management server, model training server, or terminal described above.
[0109] As a sixth aspect of the present invention, as shown in Figure 12, a computer program is stored and when the program is executed by a processor, The acquisition method described in the first aspect of this application, The generation method described in the second aspect of this application, The acquisition method described in the third aspect of this application, The present invention provides a computer-readable medium that implements at least one of the video generation methods described in the fourth aspect of this application.
[0110] A seventh aspect of the present application provides a system for a virtual character dialogue model. As shown in Figure 2, the system comprises a model management server, a model training server, and a terminal. The model management server is configured to perform the acquisition method provided in the first aspect of the present application, the model training server is configured to perform the generation method provided in the second aspect of the present application, and the terminal is configured to perform the acquisition method provided in the third aspect of the present application and / or the video generation method provided in the fourth aspect of the present application.
[0111] (Examples) This embodiment provides a model management server. As shown in Figure 13, the model management server may include a first management module, a pre-processing module, a storage module, a first authentication module, interface a, interface b, and interface c.
[0112] The first management module is the core module of the model management server, responsible for determining and processing service logic, and for scheduling and managing other modules.
[0113] Interface a is the background management interface for the model management server, providing services such as deployment and management of the model management server, transmission of background training materials, and service subscription management.
[0114] Interface b is the interface between the terminal and the model management server. It is used for communication between the terminal and the model management server, and is used for uploading model training materials and downloading trained virtual character models.
[0115] Interface c is a server-to-server interface that provides a service for transmitting training materials and models.
[0116] The pre-processing module is responsible for pre-processing training materials, generating trainable material from video files. The memory module stores information such as the initial training material, pre-processed training material, and the trained model.
[0117] The authentication module is responsible for performing bidirectional identity authentication with respect to other communication network elements.
[0118] The present invention further provides a model training server. As shown in Figure 14, the model training server may include a second management module, a training module, a second storage module, a second authentication module, interface a, and interface c.
[0119] The second management module is the core module for processing the service logic of the model training server, and is responsible for making decisions and processing the service logic, as well as scheduling and managing other modules.
[0120] The training module is configured to train an initial model based on the installed training samples and to acquire a virtual character interaction model. To ensure training functionality, the training module may optionally be a professional GPU cluster card or NPU card.
[0121] The second memory module is configured to store training samples and the trained virtual character interaction model.
[0122] The second authentication module is positioned to perform bidirectional identity authentication with respect to other communication network elements.
[0123] Interface a is a background management interface that provides model training and model deployment and management.
[0124] Interface c is a server-to-server interface that provides a service for transmitting training samples and models.
[0125] The present invention further provides a terminal. As shown in Figure 15, the terminal comprises a third management module, a third storage module, a drive module, a third authentication module, interface b, interface d, and interface e.
[0126] Interface b is the interface between the terminal and the model management server. It is configured for communication between the terminal and the model management server, and is used for uploading model training materials and downloading trained virtual character models.
[0127] Interface d is a user interface, configured to provide a user interaction interface, offering functions such as browsing and setting up virtual character dialogue models and loading local personalized training materials.
[0128] Interface e is an application interface, where other applications or services are configured to input model ID information and drive audio information.
[0129] The third management module is the core module for processing the terminal's service logic. It is responsible for making decisions and processing the service logic, and for scheduling and managing the other modules.
[0130] Based on the model information and driving audio information introduced by the third management module, the drive module is configured to synthesize virtual person videos in real time so that the terminal's display panel is shown to the user and interaction with the user can be completed. To ensure real-time conversion effects, the drive module is equipped with a dedicated hardware accelerator (e.g., high-performance GPU + NPU) and ensures that the latency is within 20ms.
[0131] The third memory module is configured to store the local initial video and the downloaded virtual character interaction model.
[0132] The third authentication module is positioned to perform bidirectional identity authentication with respect to other communication network elements.
[0133] (Example 1) As shown in Figure 16, this embodiment 1 provides a process in which a model management server interacts with a model training server to acquire a virtual character dialogue model.
[0134] The administrator logs in to the model training server (interface a) via a background management program, which can be a standalone client or a web browser.
[0135] The model management server management module receives the connection request transmitted by interface a and calls the first authentication module to perform bidirectional ID authentication. If authentication is successful, the process proceeds to the next step; if authentication is unsuccessful, the flow ends.
[0136] After authentication is complete, the administrator imports the training materials for the target character into the model management server (interface a) via the background interface, names this target character Jia Jia, and this character is a movie star whose likeness rights have already been acquired.
[0137] The first management module of the model management server receives training material transmitted through interface a, and calls the pre-processing module to perform pre-processing on the material.
[0138] After receiving notification that the pre-processing module has completed its processing, the first management module of the model management server initiates a connection with the model training server via interface c.
[0139] The model management server calls the first authentication module, and the model training server calls the second authentication module, completing bidirectional identity authentication.
[0140] The model management server transmits training materials to the model training server via interface c.
[0141] The second management module of the model training server calls the memory module to store the data and then calls the training module to start training.
[0142] After receiving notification that training is complete, the second management module of the model training server calls the second memory module to store the model.
[0143] The second management module of the model training server returns the trained model to the model management server via interface c.
[0144] The first management module of the model management server receives model information transmitted by interface c, and then calls the first storage module to store the model.
[0145] The primary management module of the model management server notifies the background administrator via interface a that model training is complete. The administrator can then confirm the completion of training or initiate training for a new model.
[0146] (Example 2) As shown in Figure 17, this embodiment 2 provides a flow of interaction between the model management server and the terminal.
[0147] The first management module of the model management server detects that model "M-000001" is ready, and if the option to push parameters is "yes", it initiates the push flow for a new virtual character interaction model.
[0148] The first management module selects users who meet the criteria, chooses users who are online and registered on this model management server but have not installed this virtual character interaction model, and sends them a push message.
[0149] Furthermore, the server can further configure the number of pushes and the push interval, such as pushing three times, each time every other day. In this embodiment, it is set to push only once. In other words, after the initial creation of the virtual character dialogue model, the model management server starts the model push flow only once.
[0150] The model management server and the terminal's virtual character client establish communication via interface b, and each calls its respective authentication module bidirectionally to complete bidirectional ID authentication.
[0151] After authentication is successful, the model management server sends the model information of the virtual character interaction model to the terminal.
[0152] After the terminal performs information analysis on the virtual character dialogue model, it displays the information to the user via the terminal's interface d, allowing the user to view the information of the virtual character dialogue model.
[0153] After the user views the virtual character dialogue model, if the information is satisfactory, a download request is sent to the model management server via the terminal after consent, and the download is initiated. If the information is not satisfactory, the process ends.
[0154] After the virtual character interaction model download is complete, the terminal's third management module verifies the integrity of the model information. Following successful verification, the model is installed and deployed according to the model deployment parameters (Table 1).
[0155] After installation is complete, the user is notified of the installation results, and the local model deployment information library is updated. The installation results are also fed back to the server.
[0156] The updated local model deployment library information is shown in Table 2. [Table 2]
[0157] (Example 3) As shown in Figure 18, this embodiment 3 illustrates the flow of selecting and downloading a virtual character interaction model to be executed by the terminal.
[0158] User A enters their account and password into their smart device and logs into the virtual persona client.
[0159] The model management server and the terminal's virtual character client establish communication via interface b, and each calls its respective authentication module bidirectionally to complete bidirectional ID authentication.
[0160] User A views downloadable model information via interface d. Here, M-000002 and M-000003 are available for download. Note that M-000001 has already been downloaded and therefore will not appear in the new model browser window where it has not yet been downloaded.
[0161] The user selects and downloads the desired model via their device. In this case, the user selected M-000002.
[0162] Since this model is paid, the interface transitions to the payment interface, where the user completes the payment. A payment of 5 CNY is required here. If payment is not required, this step can be skipped.
[0163] After payment is complete, the download flow will start.
[0164] After the virtual character interaction model download is complete, the terminal's third management module verifies the integrity of the model information. Following successful verification, the model is installed and deployed according to the model deployment parameters (Table 1).
[0165] In this example, since the model application scene is generic, the terminal further guides the user setting application scene, allowing the user to configure the settings according to the guide interface, or choose to configure them later. The flow details are the same as in Example 6, which will be described later. In this example, the user chose to configure the settings later.
[0166] After installation is complete, the user is notified of the installation results, and the local model deployment information library is updated. The installation results are also fed back to the server.
[0167] By repeating the above steps, the user downloaded M-000003 again. The updated local model deployment library information is shown in Table 3.
[0168] [Table 3]
[0169] (Example 4) As shown in Figure 19, Example 4 illustrates the flow in which the model training server performs model training.
[0170] The user logs in to the virtual persona client on their device.
[0171] The first management module of the model management server receives the connection request transmitted by interface b and calls the first authentication module to perform bidirectional ID authentication. If authentication is successful, the process proceeds to the next step; if authentication is unsuccessful, the flow ends.
[0172] After authentication is complete, the user imports training materials for the target character via interface d, in this case Xiao Zhu, which are training videos of the target character collected by the user (i.e., the final character video). The terminal's third management module calls interface b and uploads the training materials to the model management server.
[0173] When importing model assets, it is necessary to specify additional model attribute parameters. The parameters that can be set here mainly include the model nickname, model description, whether to make it public or not, cost, applicable scenes, and whether to push it or not. Other attribute parameters are automatically assigned by the server.
[0174] The first management module of the model management server receives the training material transmitted through interface b and calls the pre-processing module to perform pre-processing on the material.
[0175] After receiving notification that the pre-processing module has completed its processing, the first management module of the model management server initiates a connection with the model training server via interface c.
[0176] The model management server and the model training server each call their respective authentication modules to complete bidirectional identity authentication.
[0177] The model management server transmits training materials to the model training server via interface c.
[0178] The second management module of the model training server calls the memory module to store the data and then calls the training module to start training.
[0179] After receiving notification that training is complete, the second management module of the model training server calls the second memory module to store the model.
[0180] The management module of the model training server returns the trained model to the model management server via interface c.
[0181] After receiving model information transmitted by interface c, the management module of the model management server calls the first storage module to store the model.
[0182] The management module of the model management server notifies the terminal via interface b that the training of the virtual character interaction model is complete and it is available for download.
[0183] In this embodiment, no fees are charged to the user for model training and download. In actual service deployments, in any other form, model training or model download may be charged. The specific billing criteria may be set by the service provider. Further explanation is omitted in this embodiment.
[0184] After downloading and installing the model, the local model deployment library will be updated as shown in Table 4 below.
[0185] [Table 4]
[0186] (Example 5) Table 4 in the above embodiment shows that there are two virtual character models that cannot be used yet because no application scenes have been set. This embodiment will explain how to set the application scenes.
[0187] The flow is as shown in Figure 20.
[0188] Log in to the virtual character client.
[0189] Select "View Installed Model List," and you will find M-000001, M-000002, M-000003, and M-000004.
[0190] Select a model for which no application scene has been set; in this case, select M-000002. Note that the usage scenes for the two models M-000001 and M-000003 are determined by the model owner and cannot be modified by the local user.
[0191] The user sets the application scene, and in this case, it is set to a voice call scene with the contact "Teacher Zhao".
[0192] After confirmation, complete the setup.
[0193] The local model deployment library will be updated as shown in Table 5.
[0194] [Table 5]
[0195] (Example 6) This example demonstrates the application of a virtual character. By transforming the voice assistant into the image of a pre-defined virtual character and completing the interaction with the user, the user experience is significantly improved, making them feel as if they are interacting with a real person rather than a machine. The overall flowchart is shown in Figure 21.
[0196] The Z voice assistant is activated, and the model ID information, in this case M-000001, is transmitted to the virtual person client management module via interface e.
[0197] The virtual person client management module reads the model data and calls the model-driven module to load the model.
[0198] The Z voice assistant transmits voice information in real time to the virtual person client management module via interface e, and the management module then transparently transmits it to the drive module.
[0199] The virtual character client management module is voice-driven and calculates, generates, and displays virtual character videos to the user in real time. Using dedicated hardware acceleration, the overall rendering delay is controlled to 20ms during calculation.
[0200] Through these steps, the Z voice assistant is transformed into a real person named Xiao Jia, making the user feel as if they are interacting with Xiao Jia rather than a machine. This significantly improves the user experience of the Z voice assistant.
[0201] (Example 7) This embodiment introduces a second application scenario for the virtual character. For a specific incoming phone number, the mobile phone's ringtone or incoming video is converted into a pre-set virtual character singing video; that is, the virtual character sings the ringtone, giving the human ear a new incoming call experience. The overall flowchart is shown in Figure 22.
[0202] In this embodiment, an incoming call from a specific number is designated as 10695555, which is a service phone number belonging to Company G. The telephone application transmits the configured model ID information, namely M-000003, to the virtual person client management module via interface e.
[0203] The virtual person client management module reads the model data and calls the model-driven module to load the model.
[0204] The phone application transmits ringtone data in real time to the virtual person client management module via interface e, and the management module transparently transmits it to the drive module.
[0205] The virtual character client-driven module calculates, generates, and displays a virtual character singing video in real time based on a ringtone. Using dedicated hardware acceleration, the overall rendering delay is controlled to 20ms during calculation.
[0206] JPEG2026510933000008.jpg28170
[0207] (Example 8) This embodiment introduces a third application scenario for virtual characters. For a specific contact, in a voice call, the other party's voice is replaced with a pre-configured virtual character's spoken video, thereby replacing a conventional voice call with a new "video call" experience. However, there is no traffic consumption for the video call, and the overall flowchart is shown in Figure 23.
[0208] A voice call is established with a specific contact, in this embodiment, Mr. Zhao. The phone application transmits the configured model ID information, namely M-000002, to the virtual person client management module via interface e.
[0209] The virtual person client management module reads the model data and calls the model-driven module to load the model.
[0210] The telephone application transmits the other party's voice data in real time to the virtual person client management module via interface e, and the management module transparently transmits it to the drive module.
[0211] The virtual character client management module is voice-driven and calculates, generates, and displays virtual character speech videos in real time. Using dedicated hardware acceleration, the overall rendering delay is controlled to 20ms during calculation.
[0212] The above steps convert an audio call with Master Zhao into a "video call" with Xiaomei. Since Xiaomei is a character specially created by the user based on Master Zhao's persona, the call with Master Zhao provides the user with an entirely new experience. Furthermore, there is no need to worry about the traffic consumption associated with traditional video calls, making it a win-win situation.
[0213] Although the embodiments of the present invention only show three types of virtual character application scenes, actual application scenes are not limited to these. They may also include scenes such as conference calls or ringback tone scenes during voice calls, and can be implemented according to the plans and concepts provided by the present invention.
[0214] Those skilled in the art will understand that all or some of the steps, systems, and devices disclosed in the above description may be implemented as software, firmware, hardware, or a suitable combination thereof. In hardware embodiments, the distinctions between the functional modules / units mentioned above do not necessarily correspond to distinctions between physical components; for example, one physical component may have multiple functions, or one function or step may be performed by multiple physical components working together. Some or all physical modules may be implemented as software executed by a processor such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit such as a dedicated integrated circuit. Such software may be placed on a computer-readable medium, which may include computer storage media (or non-temporary media) and communication media (or temporary media). As is known to those skilled in the art, the technical term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technique for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital multifunction disk (DVD) or other optical disk memory, magnetic boxes, magnetic tapes, magnetic disk memory or other magnetic storage devices, or any other media that store desired information and are accessible by a computer. Furthermore, it is known to those skilled in the art that communication media generally include computer-readable instructions, data structures, program modules or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information transmission media.
[0215] While examples and specific terminology are disclosed in this specification, these are used and should be interpreted solely as general illustrative terms and are not intended to be limiting. It will be obvious to those skilled in the art that, in some examples, features, characteristics, and / or elements described in combination with a particular example may be used alone or in combination with features, characteristics, and / or elements described in combination with other examples, unless otherwise specified. Therefore, it will be understood that modifications can be made in various forms and details, without departing from the scope of the Application as described by the appended claims.
Claims
1. Used as a model management server, A step of sending at least one model training data to a model training server, wherein the model training data is associated with the facial expression of a predetermined character and the audio data of a predetermined character, and a model training sample can be obtained based on the model training data, and the training sample includes the correspondence between the facial expression of the predetermined character and the audio data. The steps include receiving a virtual character interaction model corresponding to various model training data, The step of responding to a download request by sending a virtual character interaction model corresponding to the download request to the terminal corresponding to the download request. How to obtain a virtual character interaction model.
2. The data for model training includes a final video file of a predetermined character, and the final video file satisfies the parameter requirements of the training model. The acquisition method described in claim 1.
3. The data for model training includes training samples, the training samples include the correspondence between the facial expressions of a predetermined character obtained from the final video file of the predetermined character and audio data, and before the step of sending at least one piece of model training data to the model training server, The step includes processing the final video file of the predetermined character and obtaining the training sample. The acquisition method described in claim 1.
4. Before the step of processing the final video file of the predetermined character, The step further includes receiving the aforementioned final video file. The acquisition method according to claim 2 or claim 3.
5. This further includes placing identification parameter information in each of the received virtual character interaction models. The acquisition method according to any one of claims 1 to 3.
6. The identification parameter information of the virtual character dialogue model is, It includes at least one of the following parameter information: model identifier, model nickname, model introduction information, model attribution information, information characterizing whether it is public or not, cost information, application scene information, and information characterizing whether it is pushed to users or not. The acquisition method described in claim 5.
7. Used as a model training server, A step of receiving at least one model training data, wherein the model training data is associated with the facial expression of a predetermined character and the audio data of a predetermined character, and a model training sample can be obtained based on the model training data, wherein the training sample includes the correspondence between the facial expression of the predetermined character and the audio data. The steps include inputting each of the aforementioned training samples into the initial model and training the initial model to obtain a virtual character dialogue model corresponding to each of the aforementioned training samples, The step includes sending each of the aforementioned virtual character interaction models to a model management server. A method for generating a virtual character dialogue model.
8. Used in terminals, The steps include sending a download request to the model management server, The process includes the step of receiving a virtual character dialogue model, wherein the virtual character dialogue model is a model obtained by a model training server training an initial model using training samples, and the training samples include a correspondence between the facial expressions of a predetermined character and audio data. How to obtain a virtual character interaction model.
9. Before the step of sending a download request to the model management server, The step further includes sending the final video file of a predetermined character to the model management server. The acquisition method described in claim 8.
10. Used in terminals, A step of determining a virtual character dialogue model, wherein the virtual character dialogue model is a model obtained by a model training server training an initial model with training samples, and the training samples include a correspondence between the facial expressions of a predetermined character and audio data. The steps include inputting drive information to the confirmed virtual character dialogue model and obtaining an output video corresponding to the drive information, wherein the character in the output video is the predetermined character. Video generation method.
11. In the step of determining the virtual character dialogue model, the determined virtual character dialogue model matches the target application scene. The video generation method according to claim 10.
12. One or more processors, When one or more programs are stored and the one or more programs are executed by the one or more processors, A method of acquisition according to any one of claims 1 to 6, The generation method described in claim 7, The acquisition method described in claim 8 or claim 9, A video generation method according to claim 10 or claim 11, and a memory that implements at least one of the methods by one or more processors, comprising: electronic equipment.
13. When a computer program is stored and executed by the processor, A method of acquisition according to any one of claims 1 to 6, The generation method described in claim 7, The acquisition method described in claim 8 or claim 9, A video generation method according to claim 10 or claim 11, and at least one of the methods described above. A computer-readable medium.