Animation synthesis method, device, mobile terminal and electronic device

By obtaining speech recognition features and generating speech lip movement parameters in 3D lip movement synthesis, and generating lip movement animations with face models, the problem of low efficiency and large error in the existing technology is solved, and efficient tone-independent lip movement synthesis on the mobile terminal is realized.

CN112541956BActive Publication Date: 2025-07-22BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011226145.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-11-05
Publication Date
2025-07-22
Estimated Expiration
2040-11-05

AI Technical Summary

Technical Problem

The diversified input features and phoneme duration requirements of existing 3D lip synthesis technology lead to low synthesis efficiency and large errors, which cannot be applied to mobile devices.

Method used

By obtaining the speech recognition features in the sound file, using the lip movement parameters to obtain the model for processing, generating the speech lip movement parameters, and combining the face model to generate lip movement animations to avoid phoneme duration requirements and realize timbre-independent lip movement synthesis.

Benefits of technology

It improves the efficiency of lip animation synthesis, reduces synthesis errors, and reduces bandwidth requirements, and is suitable for mobile terminals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112541956B_ABST
    Figure CN112541956B_ABST
Patent Text Reader

Abstract

The present application discloses an animation synthesis method, apparatus, mobile terminal, and electronic device, which relate to the field of computer technologies, specifically to artificial intelligence technologies such as speech technology and deep learning. The specific implementation solution is as follows: obtain a sound file; obtain the speech recognition features in the sound file; process the speech recognition features by a model according to lip movement parameters to obtain speech lip movement parameters; and generate a lip movement animation according to the speech lip movement parameters and a face model. In the animation synthesis method of the embodiments of the present application, by using the speech recognition features as the input features for extracting the speech lip movement parameters, the phoneme duration is not required, and it is independent of the timbre, which can not only improve the synthesis efficiency but also reduce the synthesis error.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, specifically to artificial intelligence technologies such as speech technology and deep learning, and particularly to an animation synthesis method, device, mobile terminal, and electronic device. Background Art

[0002] Modal synthesis is an extension of TTS (TextToSpeech), and the goal is to synthesize an animated image that matches the synthesized voice. The most core and crucial part is that the lip movement and lip shape of the synthesized image need to match the speech.

[0003] However, existing 3D lip movement synthesis is basically based on the scheme of driving Blendshape (blended shape) deformers. The input features are diverse, some are text features, some are the Mel spectrogram of speech, and many existing schemes need to use phoneme duration and are not mobile-terminal oriented. Summary of the Invention

[0004] This application provides an animation synthesis method, device, mobile terminal, and electronic device.

[0005] According to one aspect of this application, an animation synthesis method is provided, including:

[0006] Obtain a sound file;

[0007] Obtain the speech recognition features in the sound file;

[0008] Process the speech recognition features by a model according to lip movement parameters to obtain speech lip movement parameters; and

[0009] Generate a lip movement animation according to the speech lip movement parameters and a face model.

[0010] According to another aspect of this application, an animation synthesis device is provided, including:

[0011] A first acquisition module, configured to obtain a sound file;

[0012] A second acquisition module, configured to obtain the speech recognition features in the sound file;

[0013] A third acquisition module, configured to process the speech recognition features by a model according to lip movement parameters to obtain speech lip movement parameters; and

[0014] A generation module, configured to generate a lip movement animation according to the speech lip movement parameters and a face model.

[0015] According to another aspect of this application, a mobile terminal is provided, including the animation synthesis device described in the above-mentioned embodiment of one aspect.

[0016] According to another aspect of the present application, there is provided an electronic device, including:

[0017] at least one processor; and

[0018] a memory communicatively connected to the at least one processor; wherein,

[0019] the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the animation synthesis method described in the above-mentioned embodiment of one aspect.

[0020] According to another aspect of the present application, there is provided a non-transitory computer-readable storage medium storing computer instructions, on which a computer program is stored, and the computer instructions are used to cause the computer to execute the animation synthesis method described in the above-mentioned embodiment of one aspect.

[0021] According to another aspect of the present application, there is provided a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the animation synthesis method described in the above-mentioned embodiment of one aspect.

[0022] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become easily understandable through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The drawings are used to better understand the solution and do not constitute a limitation to the present application. Among them:

[0024] Figure 1 is a schematic flowchart of an animation synthesis method provided by an embodiment of the present application;

[0025] Figure 2 is a schematic flowchart of another animation synthesis method provided by an embodiment of the present application;

[0026] Figure 3 is a schematic flowchart of yet another animation synthesis method provided by an embodiment of the present application;

[0027] Figure 4 is a schematic block diagram of an animation synthesis device provided by an embodiment of the present application;

[0028] Figure 5 is a schematic block diagram of another animation synthesis device provided by an embodiment of the present application;

[0029] Figure 6 is a schematic block diagram of a mobile terminal provided by an embodiment of the present application; and

[0030] Figure 7 A block diagram of an electronic device for an animation synthesis method according to an embodiment of the present application. Specific embodiments

[0031] The following describes exemplary embodiments of the present application with reference to the accompanying drawings. Various details of the embodiments of the present application are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present application. Similarly, descriptions of well-known functions and structures are omitted below for clarity and conciseness.

[0032] The following describes an animation synthesis method, apparatus, mobile terminal, electronic device, and storage medium according to embodiments of the present application with reference to the accompanying drawings.

[0033] Artificial intelligence is a discipline that studies the use of computers to simulate certain thinking processes and intelligent behaviors of humans (such as learning, reasoning, thinking, planning, etc.). It has technical fields at both the hardware level and the software level. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, as well as deep learning, big data processing technology, and knowledge graph technology.

[0034] Speech technology refers to the key technologies in the field of computers, namely automatic speech recognition technology and speech synthesis technology.

[0035] Deep learning is a new research direction in the field of machine learning. Deep learning is to learn the internal laws and representation levels of sample data, and the information obtained during these learning processes is very helpful for the interpretation of data such as text, images, and sounds. Its ultimate goal is to enable machines to have the ability to analyze and learn like humans, and be able to recognize data such as text, images, and sounds. Deep learning is a complex machine learning algorithm, and the effects achieved in speech and image recognition far exceed those of previous related technologies.

[0036] In the embodiments of the present application, in view of the problem in the related technology that lip movement synthesis is basically based on the scheme of driving Blendshape deformers, the input features are diverse, some are text features, some are the mel spectrogram of speech, and many existing schemes need to use phoneme duration and cannot be independent of timbre, resulting in low synthesis efficiency and large errors, an animation synthesis method is proposed.

[0037] The animation synthesis method according to the embodiments of the present application processes the speech recognition features obtained from a sound file by a model according to lip movement parameters to obtain speech lip movement parameters, and generates a lip movement animation according to the speech lip movement parameters and a face model, solving the problems in the above related technologies and reducing the bandwidth required for lip movement animation synthesis at the same time.

[0038] The animation synthesis method provided by the embodiments of the present application can be executed by an electronic device, which can be a PC (Personal Computer), a tablet computer, a handheld computer, etc., without any limitation here.

[0039] In the embodiments of the present application, a processing component, a storage component, and a driving component can be provided in the electronic device. Optionally, the driving component and the processing component can be integrally provided. The storage component can store an operating system, an application program, or other program modules. The processing component implements the animation synthesis method provided by the embodiments of the present application by executing the application program stored in the storage component.

[0040] Figure 1 It is a schematic flowchart of an animation synthesis method provided by the embodiments of the present application.

[0041] The animation synthesis method according to the embodiments of the present application can also be executed by the animation synthesis device provided by the embodiments of the present application. The device can be configured in a mobile terminal to implement processing the speech recognition features obtained from a sound file by a model according to lip movement parameters to obtain speech lip movement parameters, and generating a lip movement animation according to the speech lip movement parameters and a face model. It should be noted that the device described in this embodiment can also be configured in an electronic device.

[0042] As a possible situation, the animation synthesis method according to the embodiments of the present application can also be executed on the server side. The server can be a cloud server, and the animation synthesis method can be executed in the cloud.

[0043] In the embodiments of the present application, the animation synthesis method according to the embodiments of the present application can be applied to a search virtual assistant APP (Application), and the search virtual assistant APP can be installed on a mobile terminal. It should be noted that the application interface of the search virtual assistant APP described in this embodiment can include a face (person) model, and the user can interact with the face (person) model through voice. For example, when the user asks "What's the weather like today?", the face (person) model can give a voice answer, such as "The temperature today is 13 - 20 degrees, cloudy".

[0044] As Figure 1 shown, the animation synthesis method can include the following steps:

[0045] Step 101: Obtain a sound file. The sound file can be a file in WAV format.

[0046] To better describe this application, take the animation synthesis method of the embodiments of this application applied to the Search Virtual Assistant APP as an example. The mobile terminal can obtain the sound file through the Search Virtual Assistant APP. It should be noted that the sound file can be the sound file of the face model in the Search Virtual Assistant APP replying to the user's query.

[0047] Specifically, when the user needs to use the above-mentioned Search Virtual Assistant APP, first start the Search Virtual Assistant APP in the mobile terminal to enter the application interface. Then, the user can make a query through voice or text on this application interface. The Search Virtual Assistant APP receives the content of the query, searches the Internet based on the content of the query, and converts the obtained answer information into a sound file.

[0048] Step 102: Obtain the speech recognition features in the sound file. It should be noted that the speech recognition features described in this embodiment can be the near-field recognition features of speech, and the near-field recognition features can be independent of the voice timbre.

[0049] Step 103: Process the speech recognition features according to the lip movement parameter acquisition model to obtain the speech lip movement parameters. The lip movement parameter acquisition model can be a unidirectional recurrent neural network model.

[0050] It should be noted that the lip movement parameter acquisition model described in this embodiment can be pre-trained in advance and stored in the storage space of the mobile terminal for convenient retrieval and application. The storage space is not limited to the physical storage space, such as a hard disk. The above storage space can also be the storage space of a network hard disk (cloud storage space) connected to the mobile terminal.

[0051] Specifically, after the mobile terminal obtains the sound file through the Search Virtual Assistant APP, it can also obtain the speech recognition features in the sound file through the Search Virtual Assistant APP, and then input the speech recognition features into the lip movement parameter acquisition model, so that the lip movement parameter acquisition model processes the speech recognition features to extract the speech lip movement parameters from the speech recognition features.

[0052] Step 104: Generate a lip movement animation according to the speech lip movement parameters and the face model. The face model can be a 3D face model.

[0053] Specifically, after the mobile terminal obtains the voice lip movement parameters through the search virtual assistant APP, it can substitute the voice lip movement parameters into the face model through the search virtual assistant APP for calculation to generate a lip movement animation. Among them, the above calculation process can be based on a preset algorithm, and the preset algorithm can be calibrated according to the actual situation, and no specific limitation is made here.

[0054] In the embodiment of the present application, first, a sound file is obtained, and the speech recognition feature in the sound file is obtained. Then, according to the model for obtaining lip movement parameters, the speech recognition feature is processed to obtain the voice lip movement parameters. Finally, a lip movement animation is generated according to the voice lip movement parameters and the face model. Thus, by using the speech recognition feature as the input feature for extracting the voice lip movement parameters, the phoneme duration is not required, and it is independent of the timbre, which can not only improve the synthesis efficiency but also reduce the synthesis error.

[0055] To clearly illustrate the previous embodiment, in an embodiment of the present application, generating a lip movement animation according to the voice lip movement parameters and the face model may include converting the face model into a face mesh model, and substituting the voice lip movement parameters into the face mesh model for calculation to generate a lip movement animation.

[0056] It should be noted that assuming the face model is a 3D face model, the face mesh model described in this embodiment may include at least one three-dimensional coordinate system and multiple feature points.

[0057] Specifically, after the mobile terminal obtains the voice lip movement parameters through the search virtual assistant APP, it can first convert the face model into a face mesh model, substitute the voice lip movement parameters into the face mesh model for calculation to obtain the vectors of the feature points near the lips in the face mesh model, and generate a lip movement animation by executing the vectors of the feature points near the lips in the face mesh model. Thus, it is not necessary to use the rendering farm of the GPU (Graphics Processing Unit) cluster for drawing, which can greatly reduce the bandwidth required for lip movement animation synthesis.

[0058] To improve the accuracy of lip movement animation synthesis, in an embodiment of the present application, obtaining the speech recognition feature in the sound file may include inputting the sound file into a near-field recognition model, and extracting the speech information in the sound file through the near-field recognition model to obtain the speech recognition feature.

[0059] It should be noted that the near-field recognition model described in this embodiment can be pre-stored in the storage space of the mobile terminal for convenient retrieval and application. The storage space is not limited to the physical storage space, such as a hard disk. The above storage space can also be the storage space of a network hard disk (cloud storage space) connected to the mobile terminal.

[0060] Specifically, after the mobile terminal obtains the sound file by searching for the virtual assistant APP, it can also input the sound file into the near-field recognition model through the search virtual assistant APP. The near-field recognition model extracts the voice information in the sound file and performs relevant processing on the voice information to obtain voice recognition features. Thus, the voice recognition features obtained by the near-field recognition model are independent of the timbre, thereby improving the accuracy of lip movement animation synthesis.

[0061] To further improve the accuracy of obtaining lip movement parameters, in an embodiment of the present application, as Figure 2 shown, the lip movement parameter acquisition model can be trained in the following manner:

[0062] Step 201: Obtain sample face videos and corresponding voices. Among them, there can be multiple sample face videos and corresponding voices.

[0063] In the embodiment of the present application, there are multiple ways to obtain sample face videos and corresponding voices. Among them, the face videos and corresponding voices of multiple speakers when speaking can be collected through a video device, or videos on the network (for example, videos of news broadcast hosts when speaking) can be directly obtained and used as sample face videos and corresponding voices.

[0064] Step 202: Process the voice according to the near-field recognition model to obtain sample voice recognition features.

[0065] Step 203: Process the sample face video according to the face modeling model to obtain target lip movement parameters. Among them, the face modeling model can be a 3DMM face modeling model.

[0066] It should be noted that the face modeling model described in this embodiment can be pre-stored in the storage space of the mobile terminal for convenient retrieval and application.

[0067] Specifically, after obtaining the sample face video and the corresponding voice, the sample face video can be input into the face modeling model for processing to obtain the target lip movement parameters. It should be noted that the target lip movement parameters described in this embodiment can be a deformation coefficient of the face. The voice is input into the near-field recognition model for processing to obtain sample voice recognition features.

[0068] Step 204: Input the sample voice recognition features into the lip movement parameter acquisition model to generate predicted voice lip movement parameters.

[0069] Step 205: Generate a loss value according to the predicted voice lip movement parameters and the target lip movement parameters, and train the lip movement parameter acquisition model according to the loss value.

[0070] Specifically, after obtaining the sample speech recognition features, the sample speech recognition features can be input into the lip movement parameter acquisition model to generate predicted speech lip movement parameters, and a loss value can be generated based on the predicted speech lip movement parameters and the target lip movement parameters, and the lip movement parameter acquisition model can be trained according to the loss value, so as to optimize the lip movement parameter acquisition model and further improve the accuracy of obtaining lip movement parameters.

[0071] In an embodiment of the present application, the training and generation of the lip movement parameter acquisition model can be performed by a related server. The server can be a cloud server or the host of a computer. A communication connection is established between the server and the mobile terminal (or electronic device) that can execute the animation synthesis method provided by the embodiment of the application. The communication connection can be at least one of a wireless network connection and a wired network connection. The server can send the trained lip movement parameter acquisition model to the mobile terminal (or electronic device) for the mobile terminal (or electronic device) to call when needed, thus greatly reducing the computing pressure on the mobile terminal (or electronic device).

[0072] In order to make the synthesized lip movement animation more natural and accurate, in an embodiment of the present application, as Figure 3 shown, the animation synthesis method may further include the following steps:

[0073] Step 301, obtain the migration weight of the face model.

[0074] It should be noted that the migration weight described in this embodiment can be relative to the systems of different mobile terminals. For example, the systems of different brands of mobile terminals are designed by different art engineers, and there must be differences in everyone's designs. Therefore, when the same speech lip movement parameters are placed on the face models of different mobile terminal systems, the shapes of the lip movements after deformation are not the same.

[0075] In an embodiment of the present application, to solve the above problems, the face model is now placed in the system of a certain brand of mobile terminal, and then the face model is converted into a face mesh model, and the coordinates of multiple feature points in the face mesh model are obtained. Finally, the coordinates of the multiple feature points are compared and calculated with the standard coordinates of the preset multiple feature points, so as to obtain the difference value (for example, displacement vector) between the coordinates of the multiple feature points and the standard coordinates of the preset multiple feature points, and it can be used as the migration weight of the face model relative to the system of a certain brand of mobile terminal.

[0076] Step 302, obtain weighted lip movement parameters according to the migration weight and the speech lip movement parameters.

[0077] In an embodiment of the present application, the weighted lip movement parameters can be obtained through the following formula (1):

[0078] bs2 = bs1 * W(1)

[0079] Among them, bs2 is the weighted lip movement parameter, bs1 is the speech lip movement parameter, and W is the migration weight.

[0080] Step 303: Generate a lip movement animation according to the weighted lip movement parameter and the face model.

[0081] Specifically, after obtaining the migration weight of the face model, the migration weight and the speech lip movement parameter can be multiplied through the above formula (1) to obtain the weighted lip movement parameter. Then, the face model is converted into a face mesh model, and the weighted lip movement parameter is substituted into the face mesh model for calculation to obtain the vectors of the feature points near the lips in the face mesh model, and a lip movement animation is generated by executing the vectors of the feature points near the lips in the face mesh model, thereby avoiding the problem that the same speech lip movement parameter acts on the deformed lip shape differently due to different mobile terminal systems, making the synthesized lip movement animation more natural and accurate.

[0082] Figure 4 It is a schematic block diagram of an animation synthesis device provided by an embodiment of the present application.

[0083] The animation synthesis device of the embodiment of the present application can be configured in a mobile terminal to implement processing the speech recognition features obtained from a sound file according to a lip movement parameter acquisition model to obtain a speech lip movement parameter, and generating a lip movement animation according to the speech lip movement parameter and the face model. It should be noted that the device described in this embodiment can also be configured in an electronic device.

[0084] In the embodiment of the present application, the animation synthesis device of the embodiment of the present application can be applied to a search virtual assistant APP (Application), and the search virtual assistant APP can be installed on a mobile terminal. It should be noted that the application interface of the search virtual assistant APP described in this embodiment may include a face (person) model, and the user can interact with the face (person) model through voice. For example, when the user asks "What's the weather like today?", the face (person) model can give a voice answer, such as "The temperature today is 13 - 20 degrees, cloudy."

[0085] Such as Figure 4 As shown, the animation synthesis device 400 may include: a first acquisition module 410, a second acquisition module 420, a third acquisition module 430, and a generation module 440.

[0086] Among them, the first acquisition module 410 is used to acquire a sound file. Among them, the sound file can be a file in WAV format.

[0087] To better describe this application, the animation synthesis device of the embodiments of this application is taken as an example and applied to the search virtual assistant APP. Among them, the first acquisition module 410 can acquire a sound file through the search virtual assistant APP. It should be noted that the sound file can be the sound file of the face model in the search virtual assistant APP replying to the user's query.

[0088] Specifically, when the user needs to use the above-mentioned search virtual assistant APP, first start the search virtual assistant APP in the mobile terminal to enter the application interface. Then, the user can query through voice or text on this application interface. The search virtual assistant APP receives the content of the query, searches the Internet based on the content of the query, and converts the obtained answer information into a sound file. Then, the first acquisition module 410 acquires this sound file.

[0089] The second acquisition module 420 is used to acquire the speech recognition features in the sound file. It should be noted that the speech recognition features described in this embodiment can be the near-field recognition features of speech, and the near-field recognition features can be independent of the voice timbre.

[0090] The third acquisition module 430 is used to process the speech recognition features according to the lip movement parameters of the model to obtain the speech lip movement parameters.

[0091] It should be noted that the lip movement parameter acquisition model described in this embodiment can be pre-trained and stored in the storage space of the mobile terminal for convenient retrieval and application. The storage space is not limited to the entity-based storage space, such as a hard disk. The above storage space can also be the storage space of the network hard disk (cloud storage space) connected to the mobile terminal.

[0092] Specifically, after the first acquisition module 410 acquires the sound file through the search virtual assistant APP, the second acquisition module 420 can acquire the speech recognition features in the sound file. Then, the third acquisition module 430 inputs the speech recognition features into the lip movement parameter acquisition model, so as to process the speech recognition features through the lip movement parameter acquisition model to extract the speech lip movement parameters from the speech recognition features.

[0093] The generation module 440 is used to generate a lip movement animation according to the speech lip movement parameters and the face model. Among them, the face model can be a 3D face model.

[0094] Specifically, after the third acquisition module 430 acquires the speech lip movement parameters, the generation module 440 can substitute the speech lip movement parameters into the face model for calculation to generate a lip movement animation. Among them, the above calculation process can be based on a preset algorithm, and the preset algorithm can be calibrated according to the actual situation and is not limited here.

[0095] In an embodiment of the present application, a sound file is obtained through a first acquisition module, a speech recognition feature in the sound file is obtained through a second acquisition module, and a model processes the speech recognition feature according to lip movement parameters through a third acquisition module to obtain speech lip movement parameters; a lip movement animation is generated by a generation module according to the speech lip movement parameters and a face model. Thus, by using the speech recognition feature as the input feature for extracting the speech lip movement parameters, the phoneme duration is not required, and it is independent of the timbre, which can not only improve the synthesis efficiency but also reduce the synthesis error.

[0096] To clearly illustrate the previous embodiment, in an embodiment of the present application, as Figure 4 shown, the generation module 440 is specifically configured to convert the face model into a face mesh model, and substitute the speech lip movement parameters into the face mesh model for calculation to generate a lip movement animation.

[0097] It should be noted that assuming the face model is a 3D face model, the face mesh model described in this embodiment may include at least one three-dimensional coordinate system and a plurality of feature points.

[0098] Specifically, after the third acquisition module 430 obtains the speech lip movement parameters, the generation module 440 may first convert the face model into a face mesh model, substitute the speech lip movement parameters into the face mesh model for calculation to obtain the vectors of the feature points near the lips in the face mesh model, and generate a lip movement animation by executing the vectors of the feature points near the lips in the face mesh model. Thus, it is not necessary to use the rendering farm of the GPU (Graphics Processing Unit) cluster for drawing, which can greatly reduce the bandwidth required for lip movement animation synthesis.

[0099] In an embodiment of the present application, as Figure 4 shown, the second acquisition module 420 is specifically configured to input the sound file into a near-field recognition model, and extract the speech information in the sound file through the near-field recognition model to obtain the speech recognition feature.

[0100] In another embodiment of the present application, as Figure 5 shown, the animation synthesis device 500 may include: a first acquisition module 510, a second acquisition module 520, a third acquisition module 530, a generation module 540, and a training module 550. Among them, the training module 550 is used to obtain a sample face video and the corresponding sound; process the sound according to the near-field recognition model to obtain sample speech recognition features; process the sample face video according to the face modeling model to obtain target lip movement parameters; input the sample speech recognition features into the lip movement parameter acquisition model to generate predicted speech lip movement parameters; and generate a loss value according to the predicted speech lip movement parameters and the target lip movement parameters, and train the lip movement parameter acquisition model according to the loss value.

[0101] It should be noted that the first acquisition module 510 and the first acquisition module 410, the second acquisition module 520 and the second acquisition module 420, the third acquisition module 530 and the third acquisition module 430, and the generation module 540 and the generation module 440 described in the above embodiments may have the same functions and structures.

[0102] In an embodiment of the present application, as Figure 5 shown, the generation module 540 is further configured to obtain the migration weight of the face model, and according to the migration weight and the speech lip movement parameters, obtain the weighted lip movement parameters, and generate a lip movement animation according to the weighted lip movement parameters and the face model.

[0103] In an embodiment of the present application, as Figure 5 shown, the generation module 540 can obtain the weighted lip movement parameters through the following formula: bs2 = bs1 * W, where bs2 is the weighted lip movement parameter, bs1 is the speech lip movement parameter, and W is the migration weight.

[0104] It should be noted that the foregoing explanation of the embodiment of the animation synthesis method also applies to the animation synthesis device of this embodiment, and will not be repeated here.

[0105] The animation synthesis device of the embodiment of the present application obtains a sound file through the first acquisition module, obtains the speech recognition feature in the sound file through the second acquisition module, and processes the speech recognition feature according to the lip movement parameters through the third acquisition module to obtain the speech lip movement parameters; and generates a lip movement animation according to the speech lip movement parameters and the face model through the generation module. Thus, by using the speech recognition feature as the input feature for extracting the speech lip movement parameters, without the need for phoneme duration and being able to be independent of timbre, both the synthesis efficiency can be improved and the synthesis error can be reduced.

[0106] To implement the above embodiment, as Figure 6 shown, the present application also proposes a mobile terminal 600, including the above animation synthesis device 610.

[0107] It should be noted that the animation synthesis device 400, the animation synthesis device 500, and the animation synthesis device 610 described in the above embodiments may have the same functions and structures.

[0108] The mobile terminal of the embodiment of the present application, through the above animation synthesis device, uses the speech recognition feature as the input feature for extracting the speech lip movement parameters, without the need for phoneme duration and being able to be independent of timbre, both the synthesis efficiency can be improved and the synthesis error can be reduced.

[0109] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0110] Figure 7 FIG. shows a schematic block diagram of an exemplary electronic device 700 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0111] As Figure 7 shown, the electronic device includes: one or more processors 701, a memory 702, and interfaces for connecting the various components, including a high-speed interface and a low-speed interface. The various components are interconnected using different buses and can be mounted on a common motherboard or otherwise mounted as required. The processor can process instructions executed within the electronic device, including instructions stored in the memory or on the memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In other embodiments, multiple processors and / or multiple buses can be used in conjunction with multiple memories and multiple memories if desired. Similarly, multiple electronic devices can be connected, each device providing part of the necessary operations (e.g., as a server array, a set of blade servers, or a multi-processor system). Figure 7 In the example of one processor 701.

[0112] The memory 702 is the non-transitory computer-readable storage medium provided by the present application. Wherein, the memory stores instructions executable by at least one processor, so that the at least one processor executes the animation synthesis method provided by the present application. The non-transitory computer-readable storage medium of the present application stores computer instructions for causing a computer to execute the animation synthesis method provided by the present application.

[0113] The memory 702, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as program instructions / modules corresponding to the animation synthesis method in the embodiments of the present application (e.g., attached Figure 4The illustrated animation synthesis device 1000 may include: a first acquisition module 100, a second acquisition module 200, a third acquisition module 300, and a generation module 400). The processor 701 executes various functional applications and data processing of the server by running non-transitory software programs, instructions, and modules stored in the memory 702, that is, implements the animation synthesis method in the above method embodiments.

[0114] The memory 702 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the electronic device of the animation synthesis method, etc. In addition, the memory 702 may include a high-speed random access memory, and may also include a non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory 702 may optionally include a memory remotely set relative to the processor 701, and these remote memories may be connected to the electronic device of the animation synthesis method through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0115] The electronic device of the animation synthesis method may further include: an input device 703 and an output device 704. The processor 701, the memory 702, the input device 703, and the output device 704 may be connected through a bus or other means, Figure 7 Taking the connection through the bus as an example.

[0116] The input device 703 can receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the electronic device of the animation synthesis method, such as input devices like a touch screen, a keypad, a mouse, a trackpad, a touchpad, a pointing stick, one or more mouse buttons, a trackball, a joystick, etc. The output device 704 may include a display device, an auxiliary lighting device (e.g., an LED), and a tactile feedback device (e.g., a vibration motor), etc. The display device may include but is not limited to a liquid crystal display (LCD), a light-emitting diode (LED) display, and a plasma display. In some embodiments, the display device may be a touch screen.

[0117] The various embodiments of the systems and techniques described herein can be implemented in digital electronic circuitry, integrated circuit systems, off-the-shelf ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or a general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0118] These computer programs (also referred to as programs, software, software applications, or code) include machine instructions for a programmable processor, and can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. As used herein, the terms “machine-readable medium” and “computer-readable medium” refer to any computer program product, apparatus, and / or device (e.g., a disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0119] For purposes of providing an interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide an interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).

[0120] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0121] A computer system can include clients and servers. Clients and servers are generally far from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. A server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system, solving the defects of high management difficulty and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS").

[0122] According to the technical solution of the embodiment of the present application, by using the speech recognition feature as the input feature for speech lip movement parameter extraction, without requiring phoneme duration and being able to be independent of timbre, it can not only improve the synthesis efficiency but also reduce the synthesis error.

[0123] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present application can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution disclosed in the present application can be achieved, and no limitation is imposed herein.

[0124] The above specific embodiments do not constitute a limitation to the protection scope of the present application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present application shall be included within the protection scope of the present application.

Claims

1. An animation synthesis method, comprising: Obtaining a sound file, wherein the sound file is a sound file of a face model in a search virtual assistant APP replying to a user's query; Obtaining a speech recognition feature in the sound file, where the speech recognition feature is a near-field recognition feature of the speech; Processing the speech recognition feature by a lip movement parameter acquisition model according to the lip movement parameter to obtain a speech lip movement parameter; and Generating a lip movement animation according to the speech lip movement parameter and the face model; Wherein, obtaining the speech recognition feature in the sound file comprises: Inputting the sound file into a near-field recognition model; and Extracting speech information in the sound file through the near-field recognition model to obtain the speech recognition feature; Wherein, it further comprises: obtaining a migration weight of the face model, wherein the migration weight of the face model is determined according to different mobile terminals; Multiplying the migration weight and the speech lip movement parameter to obtain a weighted lip movement parameter; converting the face model into a face mesh model, and substituting the weighted lip movement parameter into the face mesh model for calculation to obtain a vector of feature points near the lips in the face mesh model, and generating a lip movement animation by executing the vector of feature points near the lips in the face mesh model.

2. The animation synthesis method according to claim 1, wherein, The lip movement parameter acquisition model is trained in the following manner: Obtaining a sample face video and a corresponding sound; Processing the sound according to the near-field recognition model to obtain a sample speech recognition feature; Processing the sample face video according to a face modeling model to obtain a target lip movement parameter; Inputting the sample speech recognition feature into the lip movement parameter acquisition model to generate a predicted speech lip movement parameter; And Generating a loss value according to the predicted speech lip movement parameter and the target lip movement parameter, and training the lip movement parameter acquisition model according to the loss value.

3. The animation synthesis method according to claim 1, wherein, The weighted lip movement parameter is obtained through the following formula: bs2 = bs1 * W, wherein, bs2 is the weighted lip movement parameter, bs1 is the speech lip movement parameter, and W is the migration weight.

4. The animation synthesis method according to claim 1, wherein The generating a lip movement animation according to the speech lip movement parameter and the face model comprises: Converting the face model into a face mesh model; and Substituting the speech lip movement parameter into the face mesh model for calculation to generate a lip movement animation.

5. An animation synthesis device, comprising: A first acquisition module, configured to acquire a sound file, wherein the sound file is a sound file of a face model in a search virtual assistant APP replying to a user's query; A second acquisition module, configured to acquire a speech recognition feature in the sound file, where the speech recognition feature is a near-field recognition feature of the speech; A third acquisition module, configured to process the speech recognition feature by a lip movement parameter acquisition model according to the lip movement parameter to obtain a speech lip movement parameter; and A generation module, configured to generate a lip movement animation according to the speech lip movement parameter and the face model; Wherein, the second acquisition module is specifically configured to: Input the sound file into a near-field recognition model; and Extract the speech information in the sound file through the near-field recognition model to obtain the speech recognition features; Wherein, the generation module is further configured to: Obtain the migration weight of the face model, wherein the migration weight of the face model is determined according to different mobile terminals; Multiply the migration weight and the speech lip movement parameters to obtain weighted lip movement parameters; convert the face model into a face mesh model, and substitute the weighted lip movement parameters into the face mesh model for calculation to obtain the vectors of the feature points near the lips in the face mesh model, and generate a lip movement animation by executing the vectors of the feature points near the lips in the face mesh model.

6. The animation synthesis device according to claim 5, further comprising: A training module, configured to obtain a sample face video and the corresponding sound; Process the sound according to the near-field recognition model to obtain sample speech recognition features; Process the sample face video according to the face modeling model to obtain target lip movement parameters; Input the sample speech recognition features into the lip movement parameter acquisition model to generate predicted speech lip movement parameters; And generate a loss value according to the predicted speech lip movement parameters and the target lip movement parameters, and train the lip movement parameter acquisition model according to the loss value.

7. The animation synthesis device according to claim 5, wherein, The generation module obtains the weighted lip movement parameters through the following formula: bs2 = bs1 * W, Wherein, the bs2 is the weighted lip movement parameter, the bs1 is the speech lip movement parameter, and the W is the migration weight.

8. The animation synthesis device according to claim 5, wherein, The generation module is specifically configured to: Convert the face model into a face mesh model; and Substitute the speech lip movement parameters into the face mesh model for calculation to generate a lip movement animation.

9. A mobile terminal, comprising the animation synthesis device according to any one of claims 5-8.

10. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the animation synthesis method according to any one of claims 1-4.

11. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the animation synthesis method according to any one of claims 1-4.

12. A computer program product, comprising a computer program, where the computer program implements the animation synthesis method according to any one of claims 1-4 when executed by a processor.

Citation Information

Patent Citations

  • Personalized voice and video generation system based on phoneme posterior probability

    CN110880315A

  • Video generation method and device, electronic equipment and readable storage medium

    CN111294665A