Methods, devices, terminal equipment, and readable storage media for generating facial landmarks

By processing audio signals using a neural network model, three-dimensional facial key points are generated, solving the problem that existing technologies cannot directly generate three-dimensional facial key points and improving the convenience of human-computer interaction.

CN115424309BActive Publication Date: 2026-04-03WUHAN TCL CORP RES CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-05-31
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Current technology cannot directly generate 3D facial key points from voice signals.

Method used

By acquiring the target audio signal and inputting it into a trained neural network model, the target weight vector is output, and the target average shape vector and feature vector are combined to calculate the target's 3D facial key points.

Benefits of technology

It enables the direct generation of 3D facial key points based on target audio signals, which is simple and convenient, and improves the ease of human-computer interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115424309B_ABST
    Figure CN115424309B_ABST
Patent Text Reader

Abstract

This application relates to the field of cross-modal generation technology, and provides a method, apparatus, terminal device, and readable storage medium for generating facial key points. The method includes: acquiring a target audio signal, inputting the target audio signal into a trained neural network model for processing, and outputting a target weight vector; acquiring a target average shape vector and a target feature vector; and calculating the target three-dimensional facial key points corresponding to the target audio signal based on the target average shape vector, target feature vector, and target weight vector. This application can, to some extent, solve the problem of not being able to directly generate three-dimensional facial key points from speech signals.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of cross-modal generation technology, and in particular relates to a method, apparatus, terminal device and readable storage medium for generating facial key points. Background Technology

[0002] Vision and hearing are the primary ways people perceive the external world. Research shows that combining visual and auditory information can help people better understand what the external world is trying to convey. For example, when people communicate with each other, seeing lip movements can significantly improve their understanding of spoken content.

[0003] Therefore, generating a talking face based on voice signals can help users better understand voice content, thereby improving the convenience of interpersonal communication and human-computer interaction.

[0004] Currently, methods for generating speaking faces based on speech signals mainly fall into two categories: shape-model-oriented methods and image-oriented methods. Shape-model-oriented methods typically employ deformable facial shape models, while image-oriented methods generally predict RGB face or mouth image sequences directly from speech.

[0005] However, none of these methods can currently generate 3D facial key points directly from speech signals. Summary of the Invention

[0006] This application provides a method, apparatus, terminal device, and readable storage medium for generating facial key points, which can solve the problem to some extent that it is impossible to directly generate three-dimensional facial key points based on voice signals.

[0007] In a first aspect, embodiments of this application provide a method for generating facial key points, including:

[0008] The target audio signal is acquired and input into a trained neural network model for processing, and the target weight vector is output.

[0009] Obtain the target average shape vector and target feature vector, and calculate the target 3D facial key points corresponding to the target audio signal based on the target average shape vector, target feature vector and target weight vector.

[0010] Secondly, embodiments of this application provide a facial key point generation apparatus, comprising:

[0011] The first acquisition module is used to acquire the target audio signal;

[0012] The processing module is used to input the target audio signal into the trained neural network model for processing and output the target weight vector.

[0013] The second acquisition module is used to acquire the target average shape vector and the target feature vector;

[0014] The calculation module is used to calculate the target three-dimensional facial key points corresponding to the target audio signal based on the target average shape vector, the target feature vector, and the target weight vector.

[0015] Thirdly, embodiments of this application provide a terminal device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps in the method provided in the first aspect above.

[0016] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the method provided in the first aspect above.

[0017] Fifthly, embodiments of this application provide a computer program product that, when run on a terminal device, causes the terminal device to execute the image recognition method provided in the first aspect above.

[0018] It is understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here.

[0019] The beneficial effects of this application embodiment compared with the prior art are as follows: In this application, the three-dimensional facial key points corresponding to the target audio signal can be obtained based on the target average shape vector, target feature vector and target weight vector, which is simple and convenient, and realizes the direct generation of three-dimensional facial key points based on the target audio signal. Attached Figure Description

[0020] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a flowchart illustrating a method for generating facial key points according to an embodiment of this application;

[0022] Figure 2 This is a schematic diagram of the structure of a neural network model to be trained provided in an embodiment of this application;

[0023] Figure 3This is a schematic diagram of the structure of another neural network model to be trained provided in an embodiment of this application;

[0024] Figure 4 This is a schematic diagram of the structure of a facial key point generation device provided in an embodiment of this application;

[0025] Figure 5 This is a schematic diagram of the structure of the terminal device provided in the embodiments of this application. Detailed Implementation

[0026] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0027] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0028] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0029] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."

[0030] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0031] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0032] The facial key point generation method provided in this application embodiment can be applied to terminal devices such as mobile phones, tablets, laptops, ultra-mobile personal computers (UMPCs), netbooks, and personal digital assistants (PDAs). This application embodiment does not impose any restrictions on the specific type of terminal device.

[0033] To illustrate the technical solution provided in this application, specific embodiments are described below.

[0034] Example 1

[0035] The following describes a method for generating facial key points according to Embodiment 1 of this application. Please refer to the appendix. Figure 1 The method includes:

[0036] Step S101: Obtain the target audio signal and input the target audio signal into the trained neural network model for processing, and output the target weight vector.

[0037] In step S101, the target audio signal can be acquired by the terminal device of this embodiment, or it can be acquired by other terminal devices and then sent to the terminal device of this embodiment for processing. In this embodiment, the terminal device for acquiring the target audio signal is not limited.

[0038] After acquiring the target audio signal, the terminal device inputs the target audio signal into a trained neural network model for processing, thereby obtaining the target weight vector.

[0039] It's important to note that the duration of the target audio signal input to the trained neural network model is the same as the duration of the initial audio signal. For example, if the initial audio signal is 280 milliseconds long, then the target audio signal input to the trained neural network model will also be 280 milliseconds long. Furthermore, a sliding time window can be set when inputting the target audio signal into the trained neural network model, allowing the user to obtain the desired video. For instance, if the user wants to obtain a video containing 25 frames per second, the first input target audio signal to the trained neural network model would be from millisecond 0 to 280, and the second input target audio signal would be from millisecond 40 to 320. In this case, a sliding time window is 40 milliseconds.

[0040] In addition, a voice buffer can be set, and the semantic buffer can be initialized to zero. When the terminal device acquires the target audio signal, it first stores the target audio signal in the voice buffer.

[0041] Step S102: Obtain the target average shape vector and target feature vector, and calculate the target three-dimensional facial key points corresponding to the target audio signal based on the target average shape vector, target feature vector and target weight vector.

[0042] In step S102, the target average shape vector is a vector composed of 3D facial key points. The terminal device acquires the target average shape vector and the target feature vector, and then substitutes the target average shape vector, the target feature vector, and the target weight vector into the following formula to obtain the target 3D facial key points corresponding to the target audio signal:

[0043]

[0044] in, Let w represent the target average shape vector, w represent the target weight vector, and S represent the target feature vector.

[0045] In some embodiments, the method further includes, before acquiring the target audio signal:

[0046] Acquire the initial audio signal and input it into the neural network model to be trained for processing, and output the initial weight vector;

[0047] Obtain the true weight vector corresponding to the initial audio signal, and calculate the target loss value based on the initial weight vector and the true weight vector;

[0048] If the target loss value does not meet the preset conditions, the network parameters of the neural network model to be trained are updated according to the target loss value, and the process returns to the step of obtaining the initial audio signal.

[0049] If the target loss value meets the preset conditions, training stops, and the trained neural network model is obtained.

[0050] In this embodiment, the neural network model to be trained is trained based on the true weight vector corresponding to the initial audio signal, thereby obtaining a trained neural network model. The length of the initial audio signal can be selected according to actual needs. For example, in this application, the initial audio signal is 280 milliseconds of audio, where 40 milliseconds correspond to one frame of image. This application does not impose any limitations on this.

[0051] After processing the initial audio signal into the neural network model to be trained and outputting the initial weight vector, the initial weight vector and the true weight vector are substituted into the following formula to calculate the target loss value:

[0052] L = ||w o -w t ||1

[0053] Where L represents the target loss value, w o Let w represent the initial weight vector. t Let ||1| represent the true weight vector, and ||1|| represent the 1-norm. It should be understood that the above formula is only one way to calculate the target loss value. In practical applications, users can choose other calculation methods to calculate the target loss value, and this application does not impose any limitations on this. It should be noted that the number of frames in the image corresponding to the initial audio signal is consistent with the number of initial weight vectors and the number of true weight vectors.

[0054] After obtaining the target loss value, it is determined whether the target loss value meets the preset conditions. If the target loss value does not meet the preset conditions, the network parameters of the neural network model to be trained are updated according to the target loss value, and the process returns to the step of obtaining the initial audio signal. If the target loss value meets the preset conditions, training stops, thus obtaining the trained neural network model.

[0055] In other embodiments, the method further includes, before obtaining the true weight vector corresponding to the initial audio signal:

[0056] Obtain the initial face image corresponding to the initial audio signal, and extract the initial two-dimensional face key points corresponding to the initial face image;

[0057] The initial 2D facial landmarks are converted into initial 3D facial landmarks based on the initial facial image, and an initial shape vector is constructed based on the initial 3D facial landmarks.

[0058] Principal component analysis is performed on the initial shape vector to obtain the true weight vector, the target average shape vector, and the target feature vector corresponding to the initial face image.

[0059] In this embodiment, after extracting the initial two-dimensional facial key points corresponding to the initial face image, the initial two-dimensional facial key points are first converted into initial three-dimensional facial key points based on the initial face image, and an initial shape vector is constructed based on the initial three-dimensional facial key points. The constructed initial shape vector F' i As shown below:

[0060] F' i =(x1) i ,y1 i ,z1 i x2 i ,y2 i z2 i ,…,x k i ,y k i ,z k i ) T

[0061] Where i = 1, 2, ..., n, n represents the number of frames in the initial face image, (x k i ,y k i ,z k i ) represents the coordinates of the k-th facial keypoint on the initial face image of the i-th frame, and T represents the transpose.

[0062] After obtaining the initial shape vector, an Active Shape Model (ASM) is generated. The ASM is a deformable shape model whose changes can be represented by a set of coefficients. These coefficients are weight vectors obtained by performing Principal Component Analysis (PCA) on the initial shape vector.

[0063] Therefore, after obtaining the initial shape vector, Principal Component Analysis (PCA) is performed on the initial shape vector to obtain the true weight vector, the target average shape vector, and the target feature vector corresponding to the initial face image. At this point, the initial shape vector can be expressed by the following formula:

[0064]

[0065] in, Represents the target average shape vector. Represents the true weight vector. Let p represent the target eigenvector, and p represent the number of principal component analyses. <n。

[0066] Since p < n, the time for training the neural network model to be trained according to the true weight vector is less than the time for training the neural network model to be trained according to the initial face image. Therefore, in the application, the training time of the neural network model to be trained can be reduced.

[0067] It should be noted that after obtaining the initial shape vectors, alignment operations can be performed on each initial shape vector to eliminate non-shape interference caused by external factors such as different angles, distances, and pose transformations of the faces in the initial face images. The spectral analysis method can be used to align the initial shape vectors.

[0068] In some other embodiments, performing principal component analysis on the initial shape vectors to obtain the true weight vector, target average shape vector, and target feature vector corresponding to the initial face image includes:

[0069] Determining a reference face image and a template face image according to the initial face image, and calculating a target shape vector according to the initial shape vector, the reference face image, and the template face image;

[0070] Performing principal component analysis on the target shape vector to obtain the true weight vector, target average shape vector, and target feature vector corresponding to the initial face image.

[0071] Since the facial shapes of different speakers are different, even after performing alignment operations on each initial shape vector, the mouths, noses, and eyes may still be inconsistent. Therefore, in order to more accurately obtain the true weight vector, target average shape vector, and target feature vector corresponding to the initial face image, this kind of difference existing in each initial shape vector can be removed.

[0072] Therefore, first determine the reference face image and the template face image according to the initial face image, and then substitute the initial shape vector, the reference face image, and the template face image into the following formula to calculate the target shape vector:

[0073] F” i =F' i -F r +F f

[0074] where F” i represents the target shape vector, F r represents the reference face image, and F​​​​Finally, principal component analysis is performed on the target shape vector to obtain the true weight vector, the target average shape vector, and the target feature vector corresponding to the initial face image. At this point, the target shape vector is expressed by the following formula:

[0076]

[0077] in, Represents the target average shape vector. Represents the true weight vector. Let p represent the target eigenvector, and p represent the number of principal component analyses. <n。

[0078] It should be understood that when removing the differences between the initial shape vectors, the target weight vector is substituted into the target weight vector after obtaining the target weight vector. The formula calculates the target 3D facial key points corresponding to the target audio signal.

[0079] In this embodiment, the differences between the initial shape vectors are removed, allowing for a more accurate determination of the true weight vector, target average shape vector, and target feature vector corresponding to the initial face image. Because the true weight vector corresponding to the initial face image is obtained more accurately, the accuracy of the training results of the neural network model under training can be improved, and the convergence speed of the neural network model under training can also be increased.

[0080] In other embodiments, determining a reference face image based on an initial face image includes:

[0081] Select the upper lip key points and lower lip key points from the initial 3D facial key points, and calculate the target difference between the coordinates of the upper lip key points and the coordinates of the lower lip key points;

[0082] The initial face images corresponding to the upper lip key points and lower lip key points with target differences less than a preset threshold are used as reference face images.

[0083] In this embodiment, upper lip key points and lower lip key points are selected from the initial 3D facial key points. Then, the target difference between the coordinates of the upper lip key points and the coordinates of the lower lip key points is calculated. If the target difference is less than a preset threshold, it indicates that the initial facial image corresponding to the upper lip key points and lower lip key points contains a closed mouth. Therefore, the initial facial images corresponding to the upper lip key points and lower lip key points with target differences less than the preset threshold are used as reference facial images.

[0084] In other embodiments, the neural network model to be trained includes a first preset number of convolutional layers and a second preset number of fully connected layers.

[0085] Currently, in this field, the neural network model to be trained generally employs a Long Short-Term Memory (LSTM) model. However, in this application, because the differences in the initial shape vectors are removed, the resulting true weight vector is more accurate, thus reducing the complexity of the neural network model to be trained. Therefore, in this application, a first preset number of convolutional layers and a second preset number of fully connected layers are used as the neural network model to be trained. The first and second preset numbers can be set by the user according to actual needs, and are not limited herein. It should be understood that each convolutional layer can be followed by an activation function layer.

[0086] In a specific application, the first preset quantity includes 4, and the second preset quantity includes 1, such as... Figure 2 As shown, each convolutional layer convolves the initial audio signal using a one-dimensional convolutional kernel. The number of convolutional kernels in each layer can be increased as training time decreases. Furthermore, the stride of each convolutional layer can be increased, thereby reducing training time.

[0087] In other embodiments, the second preset quantity includes 2. In this case, the initial audio signal is input into the neural network model to be trained for processing, and an initial weight vector is output, including:

[0088] The initial audio signal is input into the convolutional layer of the neural network model to be trained for processing, and the initial audio features are output.

[0089] The initial audio features are input into the first fully connected layer of the neural network model to be trained for processing, and an intermediate weight vector is output.

[0090] The intermediate weight vector and the historical weight vector are input into the second fully connected layer of the neural network model to be trained for processing, and the initial weight vector is output. The historical weight vector is the initial weight vector obtained in the previous training.

[0091] In this application, to make the generated speaking face video more natural, i.e., to make the transition between generated face image frames smoother, the initial weight vector obtained from the previous training is used as the constraint condition for the current training. At this time, the neural network model to be trained includes a first fully connected layer and a second fully connected layer. The initial audio features are processed in the first fully connected layer of the neural network model to be trained, and the output is an intermediate weight vector, not the initial weight vector. Then, the intermediate weight vector and the historical weight vector are input into the second fully connected layer for processing, and the initial weight vector is output.

[0092] For example, such as Figure 3As shown, the first preset number is 4. After obtaining the intermediate weight vector, the intermediate weight vector and the historical weight vector are input into the second fully connected layer for processing, thereby obtaining the initial weight vector.

[0093] In this embodiment, the neural network model to be trained includes a first fully connected layer and a second fully connected layer. After the initial audio features are input into the first fully connected layer of the neural network model to be trained for processing, an intermediate weight vector is obtained. Then, the intermediate weight vector and the historical weight vector are input together into the second fully connected layer for processing, and the initial weight vector is output. This ensures that the transition between face image frames obtained after training is smoother.

[0094] In this application, the 3D facial key points corresponding to the target audio signal can be obtained simply and conveniently based on the target average shape vector, target feature vector, and target weight vector. This enables the direct generation of 3D facial key points from the target audio signal.

[0095] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0096] Example 2

[0097] Figure 4 An example of a facial landmark generation apparatus is shown; for ease of illustration, only the parts relevant to the embodiments of this application are shown. The apparatus 400 includes:

[0098] The first acquisition module 401 is used to acquire the target audio signal.

[0099] The processing module 402 is used to input the target audio signal into a trained neural network model for processing and output the target weight vector.

[0100] The second acquisition module 403 is used to acquire the target average shape vector and the target feature vector.

[0101] The calculation module 404 is used to calculate the target three-dimensional facial key points corresponding to the target audio signal based on the target average shape vector, target feature vector and target weight vector.

[0102] Optionally, the device 400 further includes a training module, which specifically includes:

[0103] The signal acquisition unit is used to acquire the initial audio signal and input the initial audio signal into the neural network model to be trained for processing, and output the initial weight vector.

[0104] The weight vector acquisition unit is used to acquire the true weight vector corresponding to the initial audio signal and calculate the target loss value based on the initial weight vector and the true weight vector.

[0105] The return execution unit is used to update the network parameters of the neural network model to be trained according to the target loss value if the target loss value does not meet the preset conditions, and then return to the execution step of obtaining the initial audio signal.

[0106] The stop training unit is used to stop training if the target loss value meets a preset condition, thus obtaining the trained neural network model.

[0107] Optionally, the device 400 further includes:

[0108] The extraction module is used to obtain the initial face image corresponding to the initial audio signal and extract the initial two-dimensional face key points corresponding to the initial face image.

[0109] The conversion module is used to convert the initial two-dimensional facial key points into the initial three-dimensional facial key points based on the initial facial image, and to construct the initial shape vector based on the initial three-dimensional facial key points.

[0110] The analysis module is used to perform principal component analysis on the initial shape vector to obtain the true weight vector, the target average shape vector, and the target feature vector corresponding to the initial face image.

[0111] Optionally, the analysis module is specifically used to perform:

[0112] The reference face image and template face image are determined based on the initial face image, and the target shape vector is calculated based on the initial shape vector, the reference face image, and the template face image.

[0113] Principal component analysis is performed on the target shape vector to obtain the true weight vector, the target average shape vector, and the target feature vector corresponding to the initial face image.

[0114] Optionally, the analysis module is specifically used to perform:

[0115] Select the upper lip key points and the lower lip key points from the initial 3D facial key points, and calculate the target difference between the coordinates of the upper lip key points and the coordinates of the lower lip key points.

[0116] The initial face images corresponding to the upper lip key points and lower lip key points with target differences less than a preset threshold are used as reference face images.

[0117] Optionally, the neural network model to be trained includes a first preset number of convolutional layers and a second preset number of fully connected layers.

[0118] Optionally, the signal acquisition unit is specifically used to perform:

[0119] The initial audio signal is input into the convolutional layer of the neural network model to be trained for processing, and the initial audio features are output.

[0120] The initial audio features are input into the first fully connected layer of the neural network model to be trained for processing, and an intermediate weight vector is output.

[0121] The intermediate weight vector and the historical weight vector are input into the second fully connected layer of the neural network model to be trained for processing, and the initial weight vector is output. The historical weight vector is the initial weight vector obtained in the previous training.

[0122] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiment 1 of this application. For details on their specific functions and technical effects, please refer to a part of the method embodiment, which will not be repeated here.

[0123] Example 3

[0124] Figure 5 This is a schematic diagram of the terminal device provided in Embodiment 3 of this application. Figure 5 As shown, the terminal device 500 of this embodiment includes: a processor 501, a memory 502, and a computer program 503 stored in the memory 502 and executable on the processor 501. When the processor 501 executes the computer program 503, it implements the steps in the various method embodiments described above. Alternatively, when the processor 501 executes the computer program 503, it implements the functions of each module / unit in the various device embodiments described above.

[0125] For example, the computer program 503 described above can be divided into one or more modules / units. These modules / units are stored in the memory 502 and executed by the processor 501 to complete this application. The one or more modules / units can be a series of computer program instruction segments capable of performing specific functions, describing the execution process of the computer program 503 in the terminal device 500. For example, the computer program 503 can be divided into a first acquisition module, a processing module, a second acquisition module, and a calculation module, with the specific functions of each module as follows:

[0126] The target audio signal is acquired and input into a trained neural network model for processing, and the target weight vector is output.

[0127] Obtain the target average shape vector and target feature vector, and calculate the target 3D facial key points corresponding to the target audio signal based on the target average shape vector, target feature vector and target weight vector.

[0128] The aforementioned terminal device may include, but is not limited to, processor 501 and memory 502. Those skilled in the art will understand that... Figure 5 This is merely an example of terminal device 500 and does not constitute a limitation on terminal device 500. It may include more or fewer components than shown, or combine certain components, or different components. For example, the terminal device described above may also include input / output devices, network access devices, buses, etc.

[0129] The processor 501 may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware modules, etc. A general-purpose processor may be a microprocessor or any conventional processor.

[0130] The aforementioned memory 502 can be an internal storage unit of the terminal device 500, such as a hard disk or RAM of the terminal device 500. The aforementioned memory 502 can also be an external storage device of the terminal device 500, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card equipped on the terminal device 500. Furthermore, the aforementioned memory 502 can include both internal storage units and external storage devices of the terminal device 500. The aforementioned memory 502 is used to store the aforementioned computer program and other programs and data required by the terminal device. The aforementioned memory 502 can also be used to temporarily store data that has been output or will be output.

[0131] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the above device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0132] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0133] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0134] In the embodiments provided in this application, it should be understood that the disclosed devices / terminal equipment and methods can be implemented in other ways. For example, the device / terminal equipment embodiments described above are merely illustrative. For instance, the division of modules or units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or plug-ins may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0135] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0136] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0137] If the integrated modules / units described above are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the above method embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the above method embodiments. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately added or removed according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electrical carrier signals and telecommunication signals.

[0138] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A method for generating facial key points, characterized in that, include: The target audio signal is acquired and input into a trained neural network model for processing, and the target weight vector is output. Obtain the target average shape vector and the target feature vector, and calculate the target three-dimensional facial key points corresponding to the target audio signal based on the target average shape vector, the target feature vector and the target weight vector; Prior to acquiring the target audio signal, the method further includes: Acquire an initial audio signal and input the initial audio signal into the neural network model to be trained for processing, and output an initial weight vector; Obtain the true weight vector corresponding to the initial audio signal, and calculate the target loss value based on the initial weight vector and the true weight vector. If the target loss value does not meet the preset conditions, update the network parameters of the neural network model to be trained based on the target loss value, and return to the step of obtaining the initial audio signal. If the target loss value meets the preset conditions, stop training and obtain the trained neural network model. Before obtaining the true weight vector corresponding to the initial audio signal, the method further includes: Obtain the initial face image corresponding to the initial audio signal, and extract the initial two-dimensional face key points corresponding to the initial face image; The initial two-dimensional facial key points are converted into initial three-dimensional facial key points based on the initial facial image, and an initial shape vector is constructed based on the initial three-dimensional facial key points. Principal component analysis is performed on the initial shape vector to obtain the true weight vector, target average shape vector, and target feature vector corresponding to the initial face image.

2. The method as described in claim 1, characterized in that, The step of performing principal component analysis on the initial shape vector to obtain the true weight vector, target average shape vector, and target feature vector corresponding to the initial face image includes: A reference face image and a template face image are determined based on the initial face image, and a target shape vector is calculated based on the initial shape vector, the reference face image, and the template face image. Principal component analysis is performed on the target shape vector to obtain the true weight vector, target average shape vector, and target feature vector corresponding to the initial face image.

3. The method as described in claim 2, characterized in that, Determining the reference face image based on the initial face image includes: Select the upper lip key point and the lower lip key point from the initial three-dimensional facial key points, and calculate the target difference between the coordinates of the upper lip key point and the coordinates of the lower lip key point; The initial face images corresponding to the upper lip key points and lower lip key points whose target difference is less than a preset threshold are used as reference face images.

4. The method according to any one of claims 1 to 3, characterized in that, The neural network model to be trained includes a first preset number of convolutional layers and a second preset number of fully connected layers.

5. The method as described in claim 4, characterized in that, The step of inputting the initial audio signal into the neural network model to be trained for processing and outputting an initial weight vector includes: The initial audio signal is input into the convolutional layer of the neural network model to be trained for processing, and the initial audio features are output. The initial audio features are input into the first fully connected layer of the neural network model to be trained for processing, and an intermediate weight vector is output. The intermediate weight vector and the historical weight vector are input into the second fully connected layer of the neural network model to be trained for processing, and the initial weight vector is output. The historical weight vector is the initial weight vector obtained in the previous training.

6. A device for generating facial key points, characterized in that, include: The first acquisition module is used to acquire the target audio signal; The processing module is used to input the target audio signal into a trained neural network model for processing and output a target weight vector; The second acquisition module is used to acquire the target average shape vector and the target feature vector; The calculation module is used to calculate the target three-dimensional facial key points corresponding to the target audio signal based on the target average shape vector, the target feature vector, and the target weight vector; The device further includes: an extraction module, used to acquire an initial face image corresponding to the initial audio signal, and extract initial two-dimensional face key points corresponding to the initial face image; The conversion module is used to convert the initial two-dimensional facial key points into the initial three-dimensional facial key points based on the initial facial image, and to construct the initial shape vector based on the initial three-dimensional facial key points; The analysis module is used to perform principal component analysis on the initial shape vector to obtain the true weight vector, the target average shape vector, and the target feature vector corresponding to the initial face image. The training module specifically includes: The signal acquisition unit is used to acquire the initial audio signal and input the initial audio signal into the neural network model to be trained for processing, and output the initial weight vector. The weight vector acquisition unit is used to acquire the true weight vector corresponding to the initial audio signal, and calculate the target loss value based on the initial weight vector and the true weight vector; The execution unit is returned to update the network parameters of the neural network model to be trained according to the target loss value if the target loss value does not meet the preset conditions, and then returns to execute the step of obtaining the initial audio signal. The training stop unit is used to stop training if the target loss value meets a preset condition, thereby obtaining the trained neural network model.

7. A terminal device, characterized in that, The terminal device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method as described in any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Virtual image synthesis method and device, electronic equipment and storage medium

    CN112465935A

  • Single-image three-dimensional face reconstruction method and system based on convolutional neural network

    CN112734911A

  • Method and apparatus for detecting and tracking lips

    US20140050392A1