Training method, driving method, device, readable medium and equipment for driving model
By using 2D video data to train deep learning models, combined with lip and image recognition technology, the robustness problem of the existing 3D facial capture data training model in multi-language and multi-pronunciation scenarios is solved, and efficient and low-cost lip-type driving is achieved, adapting to multi-language and multi-pronunciation scenarios and meeting real-time response.
Patent Information
- Application Number
- CN202210709551.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-21
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2042-06-21
AI Technical Summary
The existing lip-driven model based on 3D facial capture data training is poorly robust in multilingual and multipronunciation scenarios, and cannot meet the needs of multilingual and multipronunciation users. The data acquisition cost is high and the training and deployment cost is high.
The model is trained using 2D video data, audio and image features are extracted through deep learning models, and iterative training is carried out in combination with lip recognition and image quality recognition to obtain the first lip-type driver model, and the second lip-type driver model is obtained through supervised end-to-end training to adapt to scenes of multilingual and multi-pronunciation people.
It improves the robustness and versatility of the model, reduces labor costs, improves the accuracy and efficiency of lip-driven, can adapt to the scene needs of multilingual and multi-pronunciation people, and meets the lip-driven needs of real-time response.
Smart Images

Figure CN115546575B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of lip-activated driving, and in particular to a training method, a driving method, an apparatus, a readable medium, and a device for a driving model. Background Art
[0002] In current game production or virtual object application scenarios in a virtual space, it is necessary to produce corresponding lip animations according to target audio, so as to lip-drive a preset virtual object according to the lip animations.
[0003] In the related art, 3D facial capture data can be collected by cameras set at different positions, and then a model is trained end-to-end based on the 3D facial capture data. The trained model outputs lip parameters corresponding to different audios so that the preset virtual objects can be lip-driven according to the lip parameters. However, the model trained based on such 3D facial capture data has poor robustness and cannot meet the requirements of multi-language and multi-speaker scenarios. Summary of the Invention
[0004] This summary is provided to briefly introduce concepts that will be described in detail in the detailed description below. This summary is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0005] In a first aspect, the present disclosure provides a method for training a driving model, the method comprising:
[0006] For each video frame in the preset video sample, inputting data of the video frame into a first preset lip-activated model to obtain a virtual image corresponding to the video frame;
[0007] Training the first preset lip-activated driving model according to the data of the virtual image and the video frame to obtain a first driving sub-model;
[0008] Obtaining first preset lip parameters, and inputting the first preset lip parameters into a preset image rendering model to obtain a first rendered image;
[0009] Training a second preset lip-shaped driving model according to the first rendered image and the first preset lip-shaped parameters to obtain a second driving sub-model;
[0010] A first lip-shaped driving model is determined based on the first driving sub-model and the second driving sub-model. The first lip-shaped driving model is used to determine target lip-shaped parameters corresponding to the target audio based on a preset virtual image and the target audio to be identified, so as to lip-drive the virtual object according to the target lip-shaped parameters. The preset virtual image is an image of the virtual object to be driven.
[0011] In a second aspect, the present disclosure provides a method for driving a virtual object, the method comprising:
[0012] Get the target audio to be recognized;
[0013] Inputting a preset virtual image and the target audio into a pre-trained first lip-activated model to obtain target lip parameters corresponding to the target audio, wherein the preset virtual image is an image of a virtual object to be driven;
[0014] performing lip-driven operation on the virtual object according to the target lip-shaped parameters;
[0015] The first lip-activated model is a first lip-activated model trained by the method provided in the first aspect of the present disclosure.
[0016] In a third aspect, the present disclosure provides a training device for a driving model, the device comprising:
[0017] A first training module is configured to input data of each video frame in a preset video sample into a first preset lip-activated model to obtain a virtual image corresponding to the video frame; and train the first preset lip-activated model based on the virtual image and the data of the video frame to obtain a first driver sub-model;
[0018] A second training module is configured to obtain first preset lip parameters and input the first preset lip parameters into a preset image rendering model to obtain a first rendered image; and train a second preset lip driving model based on the first rendered image and the first preset lip parameters to obtain a second driving sub-model.
[0019] A first determination module is used to determine a first lip-shaped driving model based on the first driving sub-model and the second driving sub-model. The first lip-shaped driving model is used to determine target lip-shaped parameters corresponding to the target audio based on a preset virtual image and the target audio to be identified, so as to lip-drive the virtual object according to the target lip-shaped parameters. The preset virtual image is an image of the virtual object to be driven.
[0020] In a fourth aspect, the present disclosure provides a device for driving a virtual object, the device comprising:
[0021] An acquisition module, used to acquire the target audio to be recognized;
[0022] a second determining module, configured to input a preset virtual image and the target audio into a pre-trained first lip-activated model to obtain target lip parameters corresponding to the target audio, wherein the preset virtual image is an image of the virtual object to be activated;
[0023] A lip-sync driving module, configured to lip-drive the virtual object according to the target lip-sync parameters;
[0024] The first lip-activated model is a first lip-activated model trained by the training device provided in the third aspect of the present disclosure.
[0025] In a fifth aspect, a computer-readable medium is provided, on which a computer program is stored, which, when executed by a processing device, implements the steps of the method described in the first aspect or the second aspect of the present disclosure.
[0026] According to a sixth aspect, an electronic device is provided, including:
[0027] a storage device having a computer program stored thereon;
[0028] A processing device is used to execute the computer program in the storage device to implement the steps of the method described in the first aspect or the second aspect of the present disclosure.
[0029] According to the above technical solution, for each video frame in the preset video sample, the data of the video frame is input into a first preset lip-shaped driving model to obtain a virtual image corresponding to the video frame; the first preset lip-shaped driving model is trained based on the virtual image and the data of the video frame to obtain a first driving sub-model; first preset lip-shaped parameters are obtained and input into a preset image rendering model to obtain a first rendered image; and a second preset lip-shaped driving model is trained based on the first rendered image and the first preset lip-shaped parameters to obtain a second driving sub-model; A first lip-shaped driving model is determined based on the first driving sub-model and the second driving sub-model. The first lip-shaped driving model is used to determine the target lip-shaped parameters corresponding to the target audio based on a preset virtual image and the target audio to be identified, so as to lip-drive the virtual object according to the target lip-shaped parameters. The preset virtual image is the image of the virtual object to be driven. In this way, the first lip-shaped driving model is obtained by model training using 2D video data. Since the 2D video data can be easily obtained from an open source video database and the data volume is large, model training based on large-scale 2D video data can learn richer emotional information and the model is more robust. Therefore, the model can adapt to multi-language and multi-speaker scenarios, which improves the versatility of the model. Lip-shaped driving based on the more universal first lip-shaped driving model can also improve the accuracy and efficiency of lip-shaped driving.
[0030] Other features and advantages of the present disclosure will be described in detail in the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale. In the drawings:
[0032] Figure 1 is a flowchart of a method for training a driving model according to an exemplary embodiment;
[0033] Figure 2 is a schematic diagram showing a process of training a first driver sub-model according to an exemplary embodiment;
[0034] Figure 3 is a schematic diagram showing a process of training a second driver sub-model according to an exemplary embodiment;
[0035] Figure 4is a flowchart of another method for training a driving model according to an exemplary embodiment;
[0036] Figure 5 is a flowchart of a method for driving a virtual object according to an exemplary embodiment;
[0037] Figure 6 is a flowchart of another method for driving a virtual object according to an exemplary embodiment;
[0038] Figure 7 is a block diagram of a training device for a driving model according to an exemplary embodiment;
[0039] Figure 8 is based on Figure 7 A block diagram of a training device for a driving model shown in the illustrated embodiment;
[0040] Figure 9 is a block diagram of a driving device for a virtual object according to an exemplary embodiment;
[0041] Figure 10 is based on Figure 9 A block diagram of a driving device for a virtual object shown in the illustrated embodiment;
[0042] Figure 11 The figure is a structural block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION
[0043] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0044] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0045] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Other terms are defined in the following description.
[0046] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0047] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".
[0048] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0049] All actions of acquiring signals, information or data in this disclosure are carried out in compliance with the relevant data protection laws and policies of the country where they are located and with the authorization given by the owner of the corresponding device.
[0050] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0051] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operation of the disclosed technical solution based on the prompt message.
[0052] As an optional but non-limiting implementation, in response to a user's active request, the prompt information may be sent to the user in the form of a pop-up window, for example, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also include a selection control for the user to select "Agree" or "Disagree" to provide personal information to the electronic device.
[0053] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0054] At the same time, it is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.
[0055] The present disclosure is mainly used in scenarios where preset virtual objects (such as preset game characters, preset virtual roles, etc.) are lip-driven according to the audio to be recognized. For example, in current game production, art students need to manually make corresponding lip animations frame by frame according to the pronunciation of each sentence, but this process requires a high time cost. For example, a 3-second audio generally takes an art student about 1 day. In addition to being used in the game field, scenarios such as live broadcast of virtual objects will also increasingly enter daily life, and in these scenarios, the preset virtual objects need to be lip-driven.
[0056] In order to improve the efficiency of lip-activated speech recognition in related technologies, art students can manually create lip animations corresponding to each phoneme, and then use the ASR (Automatic Speech Recognition) algorithm to identify the phoneme sequence of a sentence, and splice art resources according to the phoneme sequence. However, this also requires art students to spend time to create animations for each phoneme, and corresponding ASRs need to be trained for different languages. The cost of model training and deployment is high, and the splicing method itself has poor consistency, resulting in monotonous lip movements.
[0057] In another related technology, 3D facial capture data can be collected by cameras set at different orientations, and then a model is trained end-to-end based on the 3D facial capture data. Based on the trained model, lip parameters corresponding to different audio frequencies are output so that preset virtual objects can be lip-driven according to the lip parameters. However, this 3D facial capture data needs to be collected by pre-setting a large number of cameras at different orientations. The data collection cost is high, and the amount of data collected is small. The model trained based on a small amount of 3D facial capture data has poor robustness and cannot meet the needs of multi-language and multi-speaker scenarios. The accuracy of lip-driven speech based on this model needs to be improved.
[0058] To solve the above-mentioned problems, the present disclosure provides a training method, a driving method, an apparatus, a readable medium and a device for a driving model. The video data in a 2D preset video sample can be used for model training to obtain a first lip-shaped driving model. Since the 2D video data can be easily obtained from an open source video database and the data volume is large, model training based on large-scale 2D video data can learn richer emotional information and the model is more robust. Therefore, the model can adapt to multi-language and multi-speaker scenarios, which improves the versatility of the model. Lip-shaped driving based on the more universal first lip-shaped driving model can also improve the accuracy and efficiency of lip-shaped driving.
[0059] At the same time, the first lip-shaped driving model is obtained based on 2D video data training, and then the target lip-shaped driving parameters are determined based on the first lip-shaped driving model for lip-shaped driving. The entire process does not require human participation and does not rely on the ASR model. Therefore, it can reduce labor costs and model training and deployment costs.
[0060] The specific embodiments of the present disclosure are described in detail below with reference to the accompanying drawings.
[0061] Figure 1 FIG. 1 is a flow chart showing a method for training a driving model according to an exemplary embodiment. Figure 1 As shown, the method includes the following steps:
[0062] In step S101, for each video frame in a preset video sample, data of the video frame is input into a first preset lip-activated model to obtain a virtual image corresponding to the video frame.
[0063] Among them, the preset video sample can be obtained from an open source video database. For example, there are a large number of talking face videos in the real world. Video data including voice information can be intercepted as the preset video sample for training to obtain the first lip-shaped driving model. The data of the video frame includes the image frame and audio data of the video frame. The first preset lip-shaped driving model is a pre-set deep learning model.
[0064] For example, Figure 2 FIG. 1 is a schematic diagram showing a process of training to obtain a first driver sub-model according to an exemplary embodiment. Figure 2 As shown, the first preset lip-shaped driving model includes an audio feature extraction model, an image feature extraction model and a picture generation model. In this way, in the process of inputting the data of the video frame into the first preset lip-shaped driving model to obtain the virtual image corresponding to the video frame, the audio data of the video frame can be input into the audio feature extraction model to extract the audio features, and the image frame of the video frame can be input into the image feature extraction model to extract the image features, and then the audio features and the image features are input into the picture generation model to obtain the virtual image corresponding to the video frame.
[0065] In step S102, the first preset lip-activated driving model is trained based on the data of the virtual image and the video frame to obtain a first driving sub-model.
[0066] In this step, the virtual image and the audio data can be input into a preset lip reading recognition model to obtain a lip reading recognition result, and the lip reading loss value can be determined based on the lip reading recognition result; the virtual image and the image frame can be input into a preset picture quality recognition model to obtain an image quality recognition result corresponding to the video frame, and the image loss value can be determined based on the image quality recognition result; the first preset lip shape driving model can be iteratively trained based on the lip reading loss value and the image loss value to obtain the first driving sub-model.
[0067] Continue with Figure 2 For example, in the process of training the first preset lip-activated model according to the data of the virtual image and the video frame to obtain the first driving sub-model, Figure 2 As shown, the virtual image and the audio data corresponding to the video frame can be input into a preset lip reading recognition model to obtain a lip reading recognition result, and the lip reading loss value can be determined based on the lip reading recognition result; the virtual image and the image frame corresponding to the video frame can be input into a preset picture quality recognition model to obtain an image quality recognition result corresponding to the video frame, and the image loss value can be determined based on the image quality recognition result; thereafter, the first preset lip shape driving model can be iteratively trained based on the lip reading loss value and the image loss value to obtain the first driving sub-model.
[0068] The specific process of iteratively training the first preset lip-driven model based on the lip reading loss value and the image loss value can refer to the specific implementation process of training the model based on the loss value described in relevant literature, which will not be repeated here.
[0069] It should be noted that if Figure 2 The preset lip reading recognition model and preset picture quality recognition model shown are only used in the model training stage of the first driving sub-model, that is, the lip reading loss value is determined by the preset lip reading recognition model, and the image loss value is determined by the preset picture quality recognition model, so as to perform feedback training on the first preset lip shape driving model based on these two loss values. In the model application stage of the first driving sub-model, the preset lip reading recognition model and the preset picture quality recognition model are not needed in the process of outputting virtual images corresponding to different video frames through the first driving sub-model.
[0070] In step S103, first preset lip parameters are obtained, and the first preset lip parameters are input into a preset image rendering model to obtain a first rendered image.
[0071] Among them, the first preset lip shape parameters can be, for example, any lip shape parameters preset by game developers, the preset image rendering model can be a pytorch3d model, and the lip shape in the first rendered image is the lip shape corresponding to the first preset lip shape parameters.
[0072] In step S104, a second preset lip-shape driving model is trained according to the first rendered image and the first preset lip-shape parameters to obtain a second driving sub-model.
[0073] Figure 3 FIG. 1 is a schematic diagram showing a process of training a second driver sub-model according to an exemplary embodiment. Figure 3 As shown, in this step, the first rendered image can be input into the second preset lip-shaped driving model to obtain the first model output parameters; the first lip-shaped parameter loss value is determined based on the first preset lip-shaped parameters and the first model output parameters; the second preset lip-shaped driving model is iteratively trained based on the first lip-shaped parameter loss value to obtain the second driving sub-model, and the second preset lip-shaped driving model is also a preset deep learning model.
[0074] Specifically, the second preset lip-driven model can be used to extract dense motion features of the lip shape on the first rendered image, and then the corresponding lip shape parameters (ie, the first model output parameters) can be predicted based on the dense motion features.
[0075] Based on the above method, the first lip-shaped driven sub-model and the second lip-shaped driven sub-model can be pre-trained. The first lip-shaped driven sub-model is used to convert a real video including the target audio to be recognized into a virtual video made by a virtual object. The second lip-shaped driven sub-model extracts the dense motion features of the lip shape on each frame of the target virtual image in the virtual video, and then predicts the corresponding lip shape parameters based on the dense motion features, thereby outputting the target lip shape parameters corresponding to the target audio.
[0076] In step S105 , a first lip-sync driving model is determined according to the first driving sub-model and the second driving sub-model.
[0077] The first lip-activated model is used to determine target lip parameters corresponding to the target audio based on a preset virtual image and the target audio to be recognized, so as to lip-activate the virtual object based on the target lip parameters. The preset virtual image is the image of the virtual object to be driven.
[0078] By adopting the above method, the video data in the 2D preset video sample can be used for model training to obtain the first lip-shaped driven model. Since the 2D video data can be easily obtained from the open source video database and the data volume is large, model training based on large-scale 2D video data can learn richer emotional information and the model is more robust. Therefore, the model can adapt to multi-language and multi-speaker scenarios, which improves the versatility of the model. Lip-shaped driven model based on the more universal first lip-shaped driven model can also improve the accuracy and efficiency of lip-shaped driven.
[0079] Considering that the first lip-shaped drive model includes two models, a first drive sub-model and a second drive sub-model, the model structure is relatively complex, and the response time for determining the target lip-shaped parameters for lip-shaped drive based on the first lip-shaped drive model will be relatively long, which cannot meet the demand for faster response to lip-shaped drive in real-time scenarios. Therefore, in another possible implementation method of the present disclosure, the preset audio corresponding to each video frame in the preset video sample can be input into the first lip-shaped drive model, and the lip-shaped parameters output by the model are used as target training samples. Based on the target training samples, another preset deep learning model is supervised end-to-end trained to obtain a second lip-shaped drive model. In this way, the target audio to be recognized can be input into the second lip-shaped drive model, and the target lip-shaped drive parameters are directly output by the second lip-shaped drive model for lip-shaped drive. Since the model structure of the second lip-shaped drive model is relatively simple compared to the first lip-shaped drive model, the response time for determining the target lip-shaped parameters for lip-shaped drive based on the second lip-shaped drive model is also relatively fast, thereby more efficiently performing lip-shaped drive, meeting the real-time response demand for lip-shaped drive in real-time scenarios. Figure 4 Introducing the training method of the second lip-sync driven model:
[0080] Figure 4 is a flowchart of another driving model training method according to an exemplary embodiment. Figure 4 As shown, the method further includes the following steps:
[0081] In step S401, a target training sample is determined by the first lip-activated model.
[0082] The target training sample includes preset audio corresponding to different video frames and second preset lip parameters corresponding to the preset audio.
[0083] In this step, for each video frame in the preset video sample, the image frame and audio data of the video frame can be input into the first lip-shaped driving model to obtain the second model output parameters; for each video frame, the audio data in the video frame is used as the preset audio, and the second model output parameters are used as the second preset lip-shaped parameters corresponding to the preset audio.
[0084] In step S402, supervised training is performed on the third preset lip-activated model according to the preset audio and the second preset lip-activated parameters corresponding to the preset audio to obtain a second lip-activated model.
[0085] The third preset lip-shaped driving model may include a preset deep learning model, and the second lip-shaped driving model is used to determine the target lip-shaped parameters corresponding to the target audio.
[0086] In this step, the preset audio can be input into the third preset lip-shaped driving model to obtain the third model output parameters corresponding to the preset audio; the second lip-shaped parameter loss value is determined based on the second preset lip-shaped parameter and the third model output parameter; the third preset lip-shaped driving model is iteratively trained based on the second lip-shaped parameter loss value to obtain the second lip-shaped driving model.
[0087] use Figure 4 The method shown can be used to train the second lip-type drive model. The target audio to be recognized can then be input into the second lip-type drive model, and the target lip-type parameters corresponding to the target audio can be directly output through the second lip-type drive model, so that lip-type drive can be performed according to the target lip-type parameters. Since the second lip-type drive model has a relatively simpler model structure than the first lip-type drive model, the response time for lip-type drive based on the target lip-type parameters determined by the second lip-type drive model is also relatively fast, and thus lip-type drive can be performed more efficiently, meeting the real-time response requirements for lip-type drive in real-time scenarios. For example, after a 10-second audio segment is input into the second lip-type drive model, lip-type drive is usually performed in a response time of only 1 second, thereby meeting the real-time lip-type drive scenario requirements.
[0088] The following describes a specific method for lip-driven operation of a virtual object based on the first lip-driven model or the second lip-driven model obtained through the above training:
[0089] Figure 5 is a flow chart showing a method for driving a virtual object according to an exemplary embodiment. Figure 5 As shown, the method includes the following steps:
[0090] In step S501 , target audio to be recognized is obtained.
[0091] Among them, the target audio refers to the preset audio corresponding to the virtual object waiting for lip-syncing. Taking the offline production scene of the game as an example, the target audio refers to the preset game audio corresponding to different preset game characters. In the real-time game interaction scene, the preset game character is required to conduct real-time voice interaction with the user (referring to the user who is using the game software). At this time, the target audio can be the preset dialogue audio pre-configured for the preset game character. The virtual object can include a preset virtual role, such as a preset game character.
[0092] In step S502, the preset virtual image and the target audio are input into a pre-trained first lip-activated model to obtain target lip parameters corresponding to the target audio.
[0093] Among them, the preset virtual image (or avator image) is an image of the virtual object to be driven; for example, the preset virtual image includes a preset facial image corresponding to the virtual object, and the first lip shape driving model includes a first driving sub-model and a second driving sub-model. The first driving sub-model and the second driving sub-model are both preset deep learning models, and the output of the first driving sub-model is used as the input of the second driving sub-model. The first driving sub-model is used to convert a real-life video (the real-life video includes voice information) into an avator video (the avator video replaces the real object in the real-life video with the preset virtual object); the second driving sub-model is used to convert each picture frame in the avator video into a corresponding target lip shape driving parameter, and the target lip shape parameter includes parameters that respectively characterize lip opening, lip pursing, lip movement in all directions, lip smiling, pursing the lower lip, pursing the upper lip, and other lip shapes.
[0094] Therefore, in this step, the preset virtual image and the target audio can be input into the first driving sub-model to obtain the target virtual image corresponding to the target audio; and the target virtual image can be input into the second driving sub-model to obtain the target lip parameters corresponding to the target audio.
[0095] In step S503, the virtual object is lip-driven according to the target lip parameters.
[0096] In this step, the target lip parameters may be input into a preset image rendering model to obtain a second rendering image corresponding to the target lip parameters; and then the virtual object may be lip-driven according to the second rendering image.
[0097] Among them, the preset image rendering model can be a pytorch3d model, which can render the facial image corresponding to the target lip shape parameters to obtain the second rendered image, and the lip shape displayed in the second rendered image is the target lip shape corresponding to the target lip shape parameters.
[0098] The above method is used to train the model using 2D video data to obtain a first lip-sync-driven model. Since the 2D video data can be easily obtained from an open source video database and the data volume is large, model training based on large-scale 2D video data can learn richer emotional information and the model is more robust. Therefore, the model can adapt to multi-language and multi-speaker scenarios, which improves the versatility of the model. Lip-sync-driven based on the more universal first lip-sync-driven model can also significantly improve the accuracy and efficiency of lip-sync-driven.
[0099] Figure 6 is a flow chart showing a method for driving a virtual object according to an exemplary embodiment. Figure 6 As shown, the method further includes the following steps:
[0100] In step S504, the target audio to be recognized is input into the second lip-activated model to obtain the target lip parameters corresponding to the target audio.
[0101] Among them, the second lip-shaped driving model is a deep learning model pre-trained according to the target training sample, and the target training sample is a sample determined by the first lip-shaped driving model. The target training sample specifically includes the preset audio corresponding to different video frames in the preset video sample, and the second preset lip-shaped parameters corresponding to the preset audio. The second preset lip-shaped parameters are the lip-shaped parameters output by the model after the preset audio is input into the first lip-shaped driving model (i.e., the second model output parameters).
[0102] Using the above method, the target audio to be identified can be input into the second lip-type drive model, and the target lip-type drive parameters can be directly output through the second lip-type drive model for lip-type drive. Since the second lip-type drive model has a relatively simpler model structure than the first lip-type drive model, the response time for lip-type drive based on the target lip-type parameters determined by the second lip-type drive model is also relatively fast, thereby enabling more efficient lip-type drive, meeting the real-time response requirements for lip-type drive in real-time scenarios. For example, after a 10-second audio segment is input into the second lip-type drive model, lip-type drive is usually only required in a response time of 1 second, thereby meeting the real-time lip-type drive scenario requirements.
[0103] Figure 7FIG. 1 is a block diagram of a training device for a driving model according to an exemplary embodiment. Figure 7 As shown, the device includes:
[0104] The first training module 701 is configured to input data of each video frame in a preset video sample into a first preset lip-activated model to obtain a virtual image corresponding to the video frame; and train the first preset lip-activated model based on the virtual image and the data of the video frame to obtain a first driver sub-model.
[0105] The second training module 702 is configured to obtain first preset lip parameters and input the first preset lip parameters into a preset image rendering model to obtain a first rendered image; and train a second preset lip driving model based on the first rendered image and the first preset lip parameters to obtain a second driving sub-model.
[0106] The first determination module 703 is used to determine a first lip-shaped driving model based on the first driving sub-model and the second driving sub-model. The first lip-shaped driving model is used to determine target lip-shaped parameters corresponding to the target audio based on a preset virtual image and the target audio to be identified, so as to lip-drive the virtual object according to the target lip-shaped parameters. The preset virtual image is an image of the virtual object to be driven.
[0107] Optionally, the data of the video frame includes the image frame and audio data of the video frame. The first training module 701 is used to input the virtual image and the audio data into a preset lip reading recognition model to obtain a lip reading recognition result, and determine a lip reading loss value based on the lip reading recognition result; input the virtual image and the image frame into a preset picture quality recognition model to obtain an image quality recognition result corresponding to the video frame, and determine an image loss value based on the image quality recognition result; iteratively train the first preset lip shape driving model according to the lip reading loss value and the image loss value to obtain the first driving sub-model.
[0108] Optionally, the second training module 702 is used to input the first rendered image into the second preset lip-shaped driving model to obtain the first model output parameters; determine the first lip-shaped parameter loss value based on the first preset lip-shaped parameters and the first model output parameters; and iteratively train the second preset lip-shaped driving model based on the first lip-shaped parameter loss value to obtain the second driving sub-model.
[0109] Optionally, Figure 8 is based on Figure 7 The embodiment shown is a block diagram of a training device for a driving model, such as Figure 8As shown, the device also includes:
[0110] A training sample determination module 704 is configured to determine target training samples using the first lip-driven model, wherein the target training samples include preset audio corresponding to different video frames and second preset lip parameters corresponding to the preset audio;
[0111] The third training module 705 is used to perform supervised training on the third preset lip-shaped driving model based on the preset audio and the second preset lip-shaped parameters corresponding to the preset audio to obtain a second lip-shaped driving model, and the second lip-shaped driving model is used to determine the target lip-shaped parameters corresponding to the target audio based on the target audio.
[0112] Optionally, the training sample determination module 704 is used to input the image frame and audio data of each video frame in the preset video sample into the first lip-shaped drive model to obtain second model output parameters; for each video frame, use the audio data in the video frame as the preset audio, and use the second model output parameters as the second preset lip-shaped parameters corresponding to the preset audio.
[0113] Optionally, the third training module 705 is used to input the preset audio into the third preset lip-shaped drive model to obtain a third model output parameter corresponding to the preset audio; determine a second lip-shaped parameter loss value based on the second preset lip-shaped parameter and the third model output parameter; and iteratively train the third preset lip-shaped drive model based on the second lip-shaped parameter loss value to obtain the second lip-shaped drive model.
[0114] Figure 9 is a block diagram of a driving device for a virtual object according to an exemplary embodiment. Figure 9 As shown, the device includes:
[0115] An acquisition module 901 is used to acquire target audio to be recognized;
[0116] A second determining module 902 is configured to input a preset virtual image and the target audio into a pre-trained first lip-activated model to obtain target lip parameters corresponding to the target audio, wherein the preset virtual image is an image of the virtual object to be activated;
[0117] A lip-sync driving module 903 is configured to lip-drive the virtual object according to the target lip-sync parameters;
[0118] Wherein, the first lip-activated model is Figure 7 The first lip-sync driving model is obtained by training the training device provided in the illustrated embodiment.
[0119] Optionally, the first lip-shaped driving model includes a first driving sub-model and a second driving sub-model, and the second determination module 902 is used to input the preset virtual image and the target audio into the first driving sub-model to obtain the target virtual image corresponding to the target audio; and input the target virtual image into the second driving sub-model to obtain the target lip-shaped parameters corresponding to the target audio.
[0120] Optionally, Figure 10 is based on Figure 9 The embodiment shown is a block diagram of a driving device for a virtual object, such as Figure 10 As shown, the device also includes:
[0121] The third determining module 904 is used to input the target audio to be recognized into the second lip-type driving model to obtain the target lip-type parameters corresponding to the target audio. The second lip-type driving model is based on Figure 8 The model is pre-trained by the training device provided in the illustrated embodiment.
[0122] Optionally, the lip-shaping driving module 903 is configured to input the target lip-shaping parameters into a preset image rendering model to obtain a second rendered image corresponding to the target lip-shaping parameters; and perform lip-shaping driving on the virtual object according to the second rendered image.
[0123] By using the above-mentioned device, 2D video data can be used for model training to obtain a first lip-activated model. Since the 2D video data can be easily obtained from an open source video database and the data volume is large, model training based on large-scale 2D video data can learn richer emotional information and the model is more robust. Therefore, the model can adapt to multi-language and multi-speaker scenarios, improving the versatility of the model. Lip-activated based on the more universal first lip-activated model can also improve the accuracy and efficiency of lip-activated driving.
[0124] Reference below Figure 11 , which shows a schematic structural diagram of an electronic device 1100 suitable for implementing an embodiment of the present disclosure. The terminal device in the embodiment of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 11 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.
[0125] like Figure 11 As shown, the electronic device 1100 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 1101, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1102 or a program loaded from a storage device 1108 into a random access memory (RAM) 1103. Various programs and data required for the operation of the electronic device 1100 are also stored in the RAM 1103. The processing device 1101, the ROM 1102, and the RAM 1103 are connected to each other via a bus 1104. An input / output (I / O) interface 1105 is also connected to the bus 1104.
[0126] Typically, the following devices may be connected to the I / O interface 1105: an input device 1106 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 1107 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1108 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1109. The communication device 1109 may allow the electronic device 1100 to communicate with other devices wirelessly or by wire to exchange data. Figure 11 The electronic device 1100 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead.
[0127] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 1109, or installed from the storage device 1108, or installed from the ROM 1102. When the computer program is executed by the processing device 1101, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.
[0128] It should be noted that the computer-readable medium described above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0129] In some embodiments, the client can communicate using any currently known or later developed network protocol, such as HTTP (HyperText Transfer Protocol), and can interconnect with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or later developed network.
[0130] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0131] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device: for each video frame in a preset video sample, inputs the data of the video frame into a first preset lip-shaped driving model to obtain a virtual image corresponding to the video frame; trains the first preset lip-shaped driving model based on the virtual image and the data of the video frame to obtain a first driving sub-model; obtains first preset lip-shaped parameters, and inputs the first preset lip-shaped parameters into a preset image rendering model to obtain a first rendered image; trains a second preset lip-shaped driving model based on the first rendered image and the first preset lip-shaped parameters to obtain a second driving sub-model; determines a first lip-shaped driving model based on the first and second driving sub-models, wherein the first lip-shaped driving model is used to determine the target lip-shaped parameters corresponding to the target audio based on the preset virtual image and the target audio to be recognized, so as to lip-drive the virtual object according to the target lip-shaped parameters, and the preset virtual image is the image of the virtual object to be driven.
[0132] Alternatively, the computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device: obtains the target audio to be recognized; inputs a preset virtual image and the target audio into a pre-trained first lip-type drive model to obtain target lip-type parameters corresponding to the target audio, wherein the preset virtual image is an image of the virtual object to be driven; and lip-drives the virtual object according to the target lip-type parameters; wherein the first lip-type drive model is a first lip-type drive model trained by the method described in the first aspect of the present disclosure.
[0133] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0134] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or can be implemented using a combination of dedicated hardware and computer instructions.
[0135] The modules described in the embodiments of the present disclosure may be implemented in software or hardware. In some cases, the name of a module does not necessarily limit the module itself. For example, an acquisition module may also be described as a "module for acquiring target audio."
[0136] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0137] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0138] According to one or more embodiments of the present disclosure, Example 1 provides a method for training a driving model, including:
[0139] For each video frame in the preset video sample, inputting data of the video frame into a first preset lip-activated model to obtain a virtual image corresponding to the video frame;
[0140] Training the first preset lip-activated driving model according to the data of the virtual image and the video frame to obtain a first driving sub-model;
[0141] Obtaining first preset lip parameters, and inputting the first preset lip parameters into a preset image rendering model to obtain a first rendered image;
[0142] Training a second preset lip-shaped driving model according to the first rendered image and the first preset lip-shaped parameters to obtain a second driving sub-model;
[0143] A first lip-shaped driving model is determined based on the first driving sub-model and the second driving sub-model. The first lip-shaped driving model is used to determine target lip-shaped parameters corresponding to the target audio based on a preset virtual image and the target audio to be identified, so as to lip-drive the virtual object according to the target lip-shaped parameters. The preset virtual image is an image of the virtual object to be driven.
[0144] According to one or more embodiments of the present disclosure, Example 2 provides the method of Example 1, wherein the video frame data includes image frames and audio data of the video frame, and the training of the first preset lip-activated driving model based on the virtual image and the video frame data to obtain the first driving sub-model includes:
[0145] Inputting the virtual image and the audio data into a preset lip reading recognition model to obtain a lip reading recognition result, and determining a lip reading loss value according to the lip reading recognition result;
[0146] Inputting the virtual image and the image frame into a preset picture quality recognition model to obtain an image quality recognition result corresponding to the video frame, and determining an image loss value according to the image quality recognition result;
[0147] The first preset lip-activated driving model is iteratively trained according to the lip reading loss value and the image loss value to obtain the first driving sub-model.
[0148] According to one or more embodiments of the present disclosure, Example 3 provides the method of Example 1, wherein training a second preset lip-activated model based on the first rendered image and the first preset lip-activated parameters to obtain a second driving sub-model includes:
[0149] Inputting the first rendered image into the second preset lip-driven model to obtain first model output parameters;
[0150] Determining a first lip parameter loss value according to the first preset lip parameter and the first model output parameter;
[0151] The second preset lip-shaped driving model is iteratively trained according to the first lip-shaped parameter loss value to obtain the second driving sub-model.
[0152] According to one or more embodiments of the present disclosure, Example 4 provides the method of Example 1, further comprising:
[0153] Determining target training samples using the first lip-driven model, wherein the target training samples include preset audio corresponding to different video frames and second preset lip parameters corresponding to the preset audio;
[0154] A third preset lip-shape driving model is supervisedly trained based on the preset audio and the second preset lip-shape parameters corresponding to the preset audio to obtain a second lip-shape driving model, wherein the second lip-shape driving model is used to determine the target lip-shape parameters corresponding to the target audio based on the target audio.
[0155] According to one or more embodiments of the present disclosure, Example 5 provides the method of Example 4, wherein determining the target training sample by using the first lip-activated model includes:
[0156] For each video frame in the preset video sample, inputting the image frame and audio data of the video frame into the first lip-activated model to obtain second model output parameters;
[0157] For each video frame, the audio data in the video frame is used as the preset audio, and the second model output parameters are used as the second preset lip parameters corresponding to the preset audio.
[0158] According to one or more embodiments of the present disclosure, Example 6 provides the method of Example 4, wherein supervised training is performed on a third preset lip-activated model based on the preset audio and the second preset lip-activated model corresponding to the preset audio to obtain the second lip-activated model, including:
[0159] Inputting the preset audio into the third preset lip-driven model to obtain a third model output parameter corresponding to the preset audio;
[0160] Determining a second lip parameter loss value according to the second preset lip parameter and the third model output parameter;
[0161] The third preset lip-shaped drive model is iteratively trained according to the second lip-shaped parameter loss value to obtain the second lip-shaped drive model.
[0162] According to one or more embodiments of the present disclosure, Example 7 provides a method for driving a virtual object, the method comprising:
[0163] Get the target audio to be recognized;
[0164] Inputting a preset virtual image and the target audio into a pre-trained first lip-activated model to obtain target lip parameters corresponding to the target audio, wherein the preset virtual image is an image of a virtual object to be driven;
[0165] performing lip-driven operation on the virtual object according to the target lip-shaped parameters;
[0166] The first lip-activated model is a first lip-activated model trained by the method described in any one of Examples 1-3.
[0167] According to one or more embodiments of the present disclosure, Example 8 provides the method of Example 7, wherein the first lip-activated model includes a first driving sub-model and a second driving sub-model, and inputting the preset virtual image and the target audio into the pre-trained first lip-activated model to obtain target lip parameters corresponding to the target audio includes:
[0168] Inputting the preset virtual image and the target audio into the first driving sub-model to obtain a target virtual image corresponding to the target audio;
[0169] The target virtual image is input into the second driving sub-model to obtain the target lip shape parameters corresponding to the target audio.
[0170] According to one or more embodiments of the present disclosure, Example 9 provides the method of Example 7, further comprising:
[0171] The target audio to be recognized is input into a second lip-sync driving model to obtain the target lip-sync parameters corresponding to the target audio, where the second lip-sync driving model is a model pre-trained according to the method described in any one of Examples 4-6.
[0172] According to one or more embodiments of the present disclosure, Example 10 provides the method of Example 7, wherein lip-driven operation of the virtual object according to the target lip parameters includes:
[0173] Inputting the target lip shape parameters into a preset image rendering model to obtain a second rendered image corresponding to the target lip shape parameters;
[0174] The virtual object is lip-driven according to the second rendered image.
[0175] According to one or more embodiments of the present disclosure, Example 11 provides a training device for a driving model, the device comprising:
[0176] A first training module is configured to input data of each video frame in a preset video sample into a first preset lip-activated model to obtain a virtual image corresponding to the video frame; and train the first preset lip-activated model based on the virtual image and the data of the video frame to obtain a first driver sub-model;
[0177] A second training module is configured to obtain first preset lip parameters and input the first preset lip parameters into a preset image rendering model to obtain a first rendered image; and train a second preset lip driving model based on the first rendered image and the first preset lip parameters to obtain a second driving sub-model.
[0178] A first determination module is used to determine a first lip-shaped driving model based on the first driving sub-model and the second driving sub-model. The first lip-shaped driving model is used to determine target lip-shaped parameters corresponding to the target audio based on a preset virtual image and the target audio to be identified, so as to lip-drive the virtual object according to the target lip-shaped parameters. The preset virtual image is an image of the virtual object to be driven.
[0179] According to one or more embodiments of the present disclosure, Example 12 provides a driving device for a virtual object, including:
[0180] An acquisition module, used to acquire the target audio to be recognized;
[0181] a second determining module, configured to input a preset virtual image and the target audio into a pre-trained first lip-activated model to obtain target lip parameters corresponding to the target audio, wherein the preset virtual image is an image of the virtual object to be activated;
[0182] A lip-sync driving module, configured to lip-drive the virtual object according to the target lip-sync parameters;
[0183] The first lip-activated model is a first lip-activated model trained by the training device described in Example 11.
[0184] According to one or more embodiments of the present disclosure, Example 13 provides a computer-readable medium having a computer program stored thereon, which implements the steps of the method described in any one of Examples 1-6 or Examples 7-10 when executed by a processing device.
[0185] According to one or more embodiments of the present disclosure, Example 14 provides an electronic device, including:
[0186] a storage device having a computer program stored thereon;
[0187] A processing device is used to execute the computer program in the storage device to implement the steps of the method described in any one of Examples 1-6 or Examples 7-10.
[0188] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of the present disclosure is not limited to technical solutions formed by a specific combination of the above-mentioned technical features, but also encompasses other technical solutions formed by any combination of the above-mentioned technical features or their equivalents, without departing from the above-mentioned disclosure. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.
[0189] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.
[0190] Although the subject matter has been described using language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims. Regarding the apparatus in the above-described embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method and will not be elaborated upon here.
Claims
1. A training method for a driving model, characterized in that: The method comprises: For each video frame in the preset video sample, inputting data of the video frame into a first preset lip-activated model to obtain a virtual image corresponding to the video frame; Training the first preset lip-activated driving model according to the data of the virtual image and the video frame to obtain a first driving sub-model; Obtaining first preset lip parameters, and inputting the first preset lip parameters into a preset image rendering model to obtain a first rendered image; Training a second preset lip-shaped driving model according to the first rendered image and the first preset lip-shaped parameters to obtain a second driving sub-model; A first lip-shaped driving model is determined based on the first driving sub-model and the second driving sub-model. The first lip-shaped driving model is used to determine target lip-shaped parameters corresponding to the target audio based on a preset virtual image and the target audio to be identified, so as to lip-drive the virtual object according to the target lip-shaped parameters. The preset virtual image is an image of the virtual object to be driven.
2. The method according to claim 1, characterized in that The data of the video frame includes the image frame and audio data of the video frame, and the training of the first preset lip-activated driving model based on the virtual image and the video frame data to obtain the first driving sub-model includes: Inputting the virtual image and the audio data into a preset lip reading recognition model to obtain a lip reading recognition result, and determining a lip reading loss value according to the lip reading recognition result; Inputting the virtual image and the image frame into a preset picture quality recognition model to obtain an image quality recognition result corresponding to the video frame, and determining an image loss value according to the image quality recognition result; The first preset lip-activated driving model is iteratively trained according to the lip reading loss value and the image loss value to obtain the first driving sub-model.
3. The method according to claim 1, characterized in that The training of the second preset lip-shaped driving model according to the first rendered image and the first preset lip-shaped parameters to obtain a second driving sub-model includes: Inputting the first rendered image into the second preset lip-driven model to obtain first model output parameters; Determining a first lip parameter loss value according to the first preset lip parameter and the first model output parameter; The second preset lip-shaped driving model is iteratively trained according to the first lip-shaped parameter loss value to obtain the second driving sub-model.
4. The method according to claim 1, wherein The method further comprises: Determining target training samples using the first lip-driven model, wherein the target training samples include preset audio corresponding to different video frames and second preset lip parameters corresponding to the preset audio; A third preset lip-shape driving model is supervisedly trained based on the preset audio and the second preset lip-shape parameters corresponding to the preset audio to obtain a second lip-shape driving model, wherein the second lip-shape driving model is used to determine the target lip-shape parameters corresponding to the target audio based on the target audio.
5. The method according to claim 4, characterized in that Determining the target training sample by using the first lip-activated model includes: For each video frame in the preset video sample, inputting the image frame and audio data of the video frame into the first lip-activated model to obtain second model output parameters; For each video frame, the audio data in the video frame is used as the preset audio, and the second model output parameters are used as the second preset lip parameters corresponding to the preset audio.
6. The method according to claim 4, characterized in that The performing supervised training on the third preset lip-shaped driving model according to the preset audio and the second preset lip-shaped parameters corresponding to the preset audio to obtain the second lip-shaped driving model includes: Inputting the preset audio into the third preset lip-driven model to obtain a third model output parameter corresponding to the preset audio; Determining a second lip parameter loss value according to the second preset lip parameter and the third model output parameter; The third preset lip-shaped drive model is iteratively trained according to the second lip-shaped parameter loss value to obtain the second lip-shaped drive model.
7. A method for driving a virtual object, characterized in that: The method comprises: Get the target audio to be recognized; Inputting a preset virtual image and the target audio into a pre-trained first lip-activated model to obtain target lip parameters corresponding to the target audio, wherein the preset virtual image is an image of a virtual object to be driven; performing lip-driven operation on the virtual object according to the target lip-shaped parameters; Wherein, the first lip-activated model is a first lip-activated model trained by the method according to any one of claims 1 to 3.
8. The method according to claim 7, characterized in that The first lip-activated model includes a first driving sub-model and a second driving sub-model. Inputting the preset virtual image and the target audio into the pre-trained first lip-activated model to obtain target lip parameters corresponding to the target audio includes: Inputting the preset virtual image and the target audio into the first driving sub-model to obtain a target virtual image corresponding to the target audio; The target virtual image is input into the second driving sub-model to obtain the target lip shape parameters corresponding to the target audio.
9. The method according to claim 7, characterized in that The method further comprises: The target audio to be recognized is input into a second lip-sync driving model to obtain the target lip-sync parameters corresponding to the target audio, wherein the second lip-sync driving model is a model pre-trained according to the method according to any one of claims 4 to 6.
10. The method according to claim 7, characterized in that The lip-driven operation of the virtual object according to the target lip parameters includes: Inputting the target lip shape parameters into a preset image rendering model to obtain a second rendered image corresponding to the target lip shape parameters; The virtual object is lip-driven according to the second rendered image.
11. A training device for a driving model, characterized in that: The device comprises: A first training module is configured to input data of each video frame in a preset video sample into a first preset lip-activated model to obtain a virtual image corresponding to the video frame; and train the first preset lip-activated model based on the virtual image and the data of the video frame to obtain a first driver sub-model; A second training module is configured to obtain first preset lip parameters and input the first preset lip parameters into a preset image rendering model to obtain a first rendered image; and train a second preset lip driving model based on the first rendered image and the first preset lip parameters to obtain a second driving sub-model. A first determination module is used to determine a first lip-shaped driving model based on the first driving sub-model and the second driving sub-model. The first lip-shaped driving model is used to determine target lip-shaped parameters corresponding to the target audio based on a preset virtual image and the target audio to be identified, so as to lip-drive the virtual object according to the target lip-shaped parameters. The preset virtual image is an image of the virtual object to be driven.
12. A driving device for a virtual object, characterized in that: The device comprises: An acquisition module is used to acquire the target audio to be recognized; a second determining module, configured to input a preset virtual image and the target audio into a pre-trained first lip-activated model to obtain target lip parameters corresponding to the target audio, wherein the preset virtual image is an image of the virtual object to be activated; A lip-sync driving module, configured to lip-drive the virtual object according to the target lip-sync parameters; Wherein, the first lip-activated model is a first lip-activated model trained by the training device according to claim 11.
13. A computer-readable medium having a computer program stored thereon, characterized in that: When the program is executed by a processing device, the steps of the method according to any one of claims 1 to 6 or 7 to 10 are implemented.
14. An electronic device, characterized in that: include: a storage device having a computer program stored thereon; A processing device, configured to execute the computer program in the storage device to implement the steps of the method according to any one of claims 1 to 6 or 7 to 10.
Citation Information
Patent Citations
Face driving model training and face mouth shape animation generation method
CN112396182A
Face driving method and device for virtual image, equipment and medium
CN113223125A