Method and system for speech 3D driven digital face based on deep regression network

By directly predicting 3D face motion coefficients through deep regression networks and combining speech feature processing and expression control modules, the problems of data acquisition difficulties and stiff expressions after speaker changes in voice-driven 3D digital faces are solved, achieving robustness and expression diversity.

CN118230756BActive Publication Date: 2025-11-21BEIJING UNISOUND INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410329686.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-21
Publication Date
2025-11-21
Estimated Expiration
2044-03-21

AI Technical Summary

Technical Problem

Existing voice-driven 3D digital face technology requires data to be re-collected after the speaker is changed. The correspondence between phoneme information and lip shape information is complex, which makes it difficult to implement, has poor robustness, and results in stiff expressions and monotonous lip shapes.

Method used

A deep regression network is used to directly predict 3D face motion coefficients through speech. Combined with a speech feature processing module, a 3D face blendshape parameter conversion module, and a face expression control module, facial expressions are controlled by speech to capture subtle expressions.

Benefits of technology

It achieves robustness among any speakers, solves the problem of stiff facial expressions, and improves the diversity and naturalness of lip movements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118230756B_ABST
    Figure CN118230756B_ABST
Patent Text Reader

Abstract

The method and system for voice 3D driving digital face based on deep regression network, the method calls a voice feature processing module, adopts pre-emphasis, frame division, windowing, Fourier transform, mel frequency filter, logarithmic operation, discrete cosine transform mode to process the voice signal, and obtains the mel cepstrum feature of the voice; the mel cepstrum feature of the voice is processed by adopting the model extraction mode of ASR, and the identity information in the mel cepstrum feature of the voice is removed; a 3D face blendshape parameter conversion module is called, a network based on transformer structure is adopted to extract voice feature parameters, and the voice feature parameters are converted into 3D face blendshape parameters; a face expression control module is called, the voice feature parameters and the corresponding emotional features predicted by the voice model are combined, the face expression information is controlled, and the corresponding 3D face motion information is output. The application solves the problems of difficulty in implementation, single mouth shape, poor robustness and stiff face expression.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of voice-driven digital face, in particular to a method and system for voice 3D-driven digital face based on deep regression network. BACKGROUND

[0002] Generally, the method for voice-driven 3D digital person is to pre-process the voice, extract the phoneme information corresponding to the audio, and then control the speaking of the 3D character model according to the correspondence between the phonemes and the mouth shapes.

[0003] At present, this method can only be used for a single speaker, and after changing the speaker, the data needs to be re-collected. The correspondence between the phoneme information and the mouth shape information is complex, and it is difficult to realize the smoothing between different mouth shapes. At the same time, the speaker's expression is stiff, the phonemes only correspond to the mouth shapes, and the expression is lacking.

[0004] Therefore, it is urgent to solve the problems of complex implementation, single mouth shape, poor robustness, and stiff facial expression. SUMMARY

[0005] Therefore, the present application provides a method and system for voice 3D-driven digital face based on deep regression network, which uses deep regression network to directly predict 3D face motion coefficients through voice, solves the problem of difficult implementation, uses a statistical-based method to support voice input of any person, solves the problem of algorithm robustness, and controls the expression of the face through voice, which can capture the subtle expression of the face and solve the problem of stiff facial expression.

[0006] In order to achieve the above purpose, the present application provides the following technical scheme: a method for voice 3D-driven digital face based on deep regression network, comprising:

[0007] The voice feature processing module is called, the voice signal is processed by using pre-emphasis, the high frequency of the audio in the voice signal is restored, the voice signal is processed by using weighted window, the voice signal is smoothed by using sliding window, the windowed voice signal is processed by using Fourier transform, the frequency domain signal of the voice signal is obtained, the voice signal is filtered by using mel frequency filter, the high frequency information in the voice signal is removed, the voice signal with the removed high frequency information is logarithmically transformed to obtain the mel spectrum cepstrum of the voice signal, the voice signal with the mel feature is processed by using discrete cosine transform to obtain the mel cepstrum feature, and the mel cepstrum feature of the voice is processed by using the model extraction method of ASR to remove the identity information in the mel cepstrum feature of the voice;

[0008] Call 3D face blendshape parameter conversion module, adopt the network based on transformer structure to extract speech feature parameter, convert speech feature parameter into 3D face blendshape parameter;

[0009] Call face expression control module, adopt speech model to predict the emotional characteristics of input speech, and automatically mark the speech data;Adopt the speech model to extract the emotional characteristics in the speech;The speech feature parameters and the corresponding emotional characteristics predicted by the speech model are merged, the face expression information is controlled, and the corresponding 3D face motion information is output.

[0010] As a preferred scheme of the method for driving a digital face by speech 3D in a deep regression network, the formula for expressing the mel frequency filter is:

[0011]

[0012] In the formula, f is the input frequency, and m is the mel frequency.

[0013] As a preferred scheme of the method for driving a digital face by speech 3D in a deep regression network, the mel cepstrum feature of the speech is processed by adopting the model extraction mode of ASR, the semantic feature is extracted, and the timbre and pitch information is removed.

[0014] As a preferred scheme of the method for driving a digital face by speech 3D in a deep regression network, after the speech feature parameters are converted into 3D face blendshape parameters, the wingloss mode is adopted to process the jitter of the face key points in the regression process, and the expression formula of the wingloss mode is:

[0015]

[0016]

[0017] In the formula, w is to limit the nonlinear part in the [-w, w] space, is the curvature of the nonlinear region, x is the independent variable, and C is a constant for representing the connection between the linear part and the nonlinear part.

[0018] As a preferred scheme of the method for driving a digital face by speech 3D in a deep regression network, after the corresponding 3D face motion information is output, the face motion information smoothing mode is used to keep the face smooth in a short frame, and the single-person feature fine-tuning mode is used to collect data of a single person and fine-tune the face expression model.

[0019] The application also provides a processing system for driving a digital face by speech 3D based on a deep regression network, which comprises:

[0020] The voice feature processing module is configured to process the voice signal in a pre-emphasis manner to restore high frequencies of audio in the voice signal, perform frame processing on the voice signal in a weighted window manner, perform smoothing on the voice signal in a sliding window manner, process the windowed voice signal in a Fourier transform manner to obtain a frequency domain signal of the voice signal, filter the voice signal in a Mel frequency filter to remove high frequency information in the voice signal, perform logarithmic transformation processing on the voice signal from which the high frequency information is removed to obtain a Mel spectral cepstrum of the voice signal, perform discrete cosine transformation processing on the voice signal with the Mel feature to obtain a Mel cepstrum feature, and process the Mel cepstrum feature of the voice in a model extraction manner of ASR to remove identity information in the Mel cepstrum feature of the voice.

[0021] The 3D face blendshape parameter conversion module is configured to extract voice feature parameters by using a network based on a transformer structure, and convert the voice feature parameters into 3D face blendshape parameters.

[0022] The face expression control module is configured to automatically label voice data by using a voice model to predict emotional features of input voice, extract emotional features in the voice by using the voice model, combine the voice feature parameters and corresponding emotional features predicted by the voice model, control face expression information, and output corresponding 3D face motion information.

[0023] As an optimal solution of the processing system of the voice 3D driven digital face based on the deep regression network, in the voice feature processing module,

[0024] The expression formula of the Mel frequency filter is:

[0025]

[0026] In the formula, f is an input frequency, and m is a Mel frequency.

[0027] As an optimal solution of the processing system of the voice 3D driven digital face based on the deep regression network, in the voice feature processing module,

[0028] The Mel cepstrum feature of the voice is processed in the model extraction manner of ASR to extract semantic features and remove timbre and pitch information.

[0029] As an optimal solution of the processing system of the voice 3D driven digital face based on the deep regression network, in the 3D face blendshape parameter conversion module,

[0030] After converting the speech feature parameters into 3D face blendshape parameters, the wingloss method is used to process the shaking of the face key points in the regression process, and the expression formula of the wingloss method is as follows:

[0031]

[0032]

[0033] In the formula, w is to limit the nonlinear part in the [-w, w] space; ∈ is the curvature of the nonlinear region; x is the independent variable; and C is a constant for representing the connection between the linear part and the nonlinear part.

[0034] As an optimal solution of the speech 3D driving digital face processing system based on a deep regression network, in the face expression control module:

[0035] After outputting the corresponding 3D face motion information, the face motion information smoothing method is used to keep the face smooth within a short frame; and the single-person feature fine-tuning method is used for data collection of a single person to fine-tune the face expression model.

[0036] The application has the following advantages: the speech feature processing module is called, the speech signal is processed by using the pre-emphasis method to restore the high frequency of the audio in the speech signal; the speech signal is processed by using the weighted window method, and the sliding window method is used to make the speech signal frames smooth; the windowed speech signal is processed by using the Fourier transform method to obtain the frequency domain signal of the speech signal; the speech signal is filtered by using the Mel frequency filter to remove the high frequency information in the speech signal; the speech signal with the removed high frequency information is logarithmically transformed to obtain the Mel spectrum cepstrum of the speech signal; the speech signal with the Mel feature is processed by using the discrete cosine transform to obtain the Mel cepstrum feature; the Mel cepstrum feature of the speech is processed by using the model extraction method of the ASR to remove the identity information in the speech Mel cepstrum feature; the 3D face blendshape parameter conversion module is called, the network based on the transformer structure is used to extract the speech feature parameters, and the speech feature parameters are converted into 3D face blendshape parameters; the face expression control module is called, the speech model is used to predict the emotional features of the input speech, and the speech data is automatically labeled; the speech model is used to extract the emotional features of the speech; the speech feature parameters and the corresponding emotional features predicted by the speech model are merged to control the face expression information, and the corresponding 3D face motion information is output. The application uses a deep regression network to directly predict the 3D face motion coefficient through the speech; the face expression is controlled through the speech, the subtle expression of the face can be captured, and the problems of implementation difficulty, single mouth shape, poor robustness and rigid face expression are solved. BRIEF DESCRIPTION OF DRAWINGS

[0037] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings needed to be used in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only exemplary, and for those skilled in the art, other drawings can be derived from the provided drawings without creative labor.

[0038] The structures, proportions, sizes, etc. shown in the specification are only used to cooperate with the content disclosed in the specification, to be understood and read by those skilled in the art, and do not have technical significance to limit the conditions that the present application can be implemented. Any modification of structure, change of proportion relationship or adjustment of size, without affecting the effects and purposes that the present application can produce, should still fall within the scope of the technical content disclosed by the present application.

[0039] Figure 1 A method flow diagram of video-driven 3D digital face based on a deep regression network provided in Embodiment 1 of the present application;

[0040] Figure 2 A processing system architecture diagram of video-driven 3D digital face based on a deep regression network provided in Embodiment 2 of the present application. DETAILED DESCRIPTION

[0041] The embodiments of the present application will be described below by specific specific embodiments, and those skilled in the art can easily understand other advantages and effects of the present application from the content disclosed in the specification. Obviously, the described embodiments are part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of the present application.

[0042] Embodiment 1

[0043] Reference Figure 1 Embodiment 1 of the present application provides a method of video-driven 3D digital face based on a deep regression network, comprising the following steps:

[0044] S1, call the voice feature processing module, adopt pre-emphasis method to process the voice signal, restore the high frequency of the audio in the voice signal; adopt weighted window method to process the frame of the voice signal, adopt sliding window method to make the voice signal frame smooth; adopt Fourier transform method to process the windowed voice signal, obtain the frequency domain signal of the voice signal; adopt mel frequency filter to filter process the voice signal, remove the high frequency information in the voice signal; logarithmic transform the voice signal without high frequency information to obtain the mel frequency cepstrum of the voice signal; perform discrete cosine transform on the voice signal with mel feature to obtain mel cepstrum feature; adopt ASR model extraction method to process the mel cepstrum feature of the voice to remove the identity information in the mel cepstrum feature of the voice;

[0045] Specifically, the collected voice signal is preprocessed to eliminate the influence of the lips, and then pre-emphasis processing is performed to solve the problem that the high frequency component of the voice is weakened during the process of vocal cords and lips, and to restore the high frequency part of the audio. After processing, the voice signal is relatively stable in a short time, and the weighted window is used to process the frame of the voice signal. At the same time, the sliding window method can ensure the smoothness of the language information before and after the frame. The high and low frequency of the voice signal can better reflect the original characteristics of the voice signal, and the original information of the voice signal is expressed through the structural characteristics of the frequency. In order to obtain the spectral characteristics of the voice signal, Fourier transform (FFT) is needed on the windowed voice signal to obtain the frequency domain signal of the voice information. The mel frequency filter is used to filter process the voice signal, and as the input frequency increases, the output frequency tends to be stable, which is consistent with the auditory effect of human ear. The logarithmic transformation of the voice signal obtains the mel frequency cepstrum of the voice signal. Discrete cosine transform (DCT) is performed on the mel feature voice signal to obtain the mel cepstrum feature (MFCC). The mel cepstrum feature represents the structural characteristics of the voice signal, so as to more accurately express the original characteristics of the voice signal.

[0046] S2, call the 3D face blendshape parameter conversion module, adopt the network based on transformer structure to extract the voice feature parameter, and convert the voice feature parameter into 3D face blendshape parameter;

[0047] Wherein, tansformer is a network structure commonly used in deep learning, which has achieved good results in most applications at present, which can make the network have stronger robustness, and may extract more effective features. Through the deep network, the voice feature is processed into high-dimensional data feature, and then the regression method is used to correspond to the face blendshape feature one by one.

[0048] S3, calling a facial expression control module, using a speech model to predict the emotional characteristics of the input speech, automatically labeling the speech data; using the speech model to extract emotional characteristics in the speech; merging the speech feature parameters and the corresponding emotional characteristics predicted by the speech model, controlling the facial expression information, and outputting the corresponding 3D facial motion information.

[0049] Wherein, merging the speech feature parameters and the corresponding emotional characteristics predicted by the speech model means splicing the two kinds of information into one dimension, and the formulaic expression is shown in formula (1):

[0050]

[0051] In the formula, R full is the complete information after splicing; R mouth is the information of the mouth shape; R head is the information of the head; t is a certain time; is splicing and fusion.

[0052] In this embodiment, the mel frequency filter expression formula is shown in formula (2):

[0053]

[0054] In the formula, f is the input frequency; and m is the mel frequency.

[0055] In this embodiment, the mel cepstrum feature of the speech is processed by using the model extraction mode of ASR to extract semantic features and remove timbre and pitch information.

[0056] In this embodiment, after the speech feature parameters are converted into 3D face blendshape parameters, the wingloss mode is used to process the jitter of the face key points in the regression process, and the expression formula of the wingloss mode is:

[0057]

[0058]

[0059] In the formula, w is to limit the nonlinear part in the [-w, w] space; ∈ is the curvature of the nonlinear region; x is the independent variable; and C is a constant used to represent the connection between the linear part and the nonlinear part.

[0060] Specifically, in Wingloss, ∈ is a very small number, so the change of ∈ will cause the stability of network training. In the initial training process, the linear part can ensure the training of the network. In the nonlinear part, the loss is limited in the nonlinear range by w, and at the same time the loss is amplified, the speech information is deeply mined, and the performance of the network is improved.

[0061] In this embodiment, after outputting the corresponding 3D face motion information, the face motion information smoothing mode is used to keep the face smooth within a short time frame; and the single-person feature fine-tuning mode is used for data collection of a single person to fine-tune the face expression model.

[0062] Specifically, the information smoothing mode is to average the blendshapes corresponding to the face features within a few frames by using a sliding window. The sliding window is a technique that generates smoothed output by applying a fixed-size window over a series of consecutive data points and performing averaging or other processing within the window. In this case, for each time point of face features and corresponding BlendShapes parameters, a sliding window, for example, a window of size N, can be used to average the last N frames of data, thereby obtaining the smoothed result.

[0063] To sum up, the speech feature processing module is called, the speech signal is processed in a pre-emphasis manner to restore the high frequency of the audio in the speech signal, the speech signal is processed in a frame manner by using a weighting window, the sliding window is used to smooth the front and rear frames of the speech signal, the windowed speech signal is processed by using the Fourier transform manner to obtain the frequency domain signal of the speech signal, the speech signal is filtered by using the Mel frequency filter to remove the high frequency information in the speech signal, the speech signal with the removed high frequency information is processed by using the logarithmic transformation to obtain the Mel spectral cepstrum of the speech signal, the speech signal with the Mel feature is processed by using the discrete cosine transform to obtain the Mel cepstrum feature, the Mel cepstrum feature of the speech is processed by using the model extraction manner of the ASR to remove the identity information in the speech Mel cepstrum feature, the 3D face blendshape parameter conversion module is called, the network based on the transformer structure is used to extract the speech feature parameter, and the speech feature parameter is converted into the 3D face blendshape parameter, the face expression control module is called, the emotion feature of the input speech is predicted by using the speech model, and the speech data is automatically labeled, the emotion feature in the speech is extracted by using the speech model, the speech feature parameter and the corresponding emotion feature predicted by using the speech model are merged, the face expression information is controlled, and the corresponding 3D face motion information is output. The deep regression network is used to directly predict the 3D face motion coefficient through the speech, the face expression is controlled through the speech, the subtle expression of the face can be captured, the problems of difficulty in implementation, single mouth shape, poor robustness and rigid face expression are solved.

[0064] It should be noted that the method of the embodiments of the present disclosure can be executed by a single device, such as a computer or a server, etc. The method of the embodiments can also be applied to a distributed scenario, and be completed by multiple devices cooperating with each other. In the case of such a distributed scenario, one of the multiple devices can only execute one or more steps in the method of the embodiments of the present disclosure, and the multiple devices can interact with each other to complete the method.

[0065] It should be noted that some embodiments of the present disclosure are described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than the order described above and still achieve desirable results. Additionally, the processes depicted in the figures do not necessarily require the particular order shown, or sequential order, to achieve the desired results. In certain implementations, multitasking and parallel processing can be advantageous.

[0066] Embodiment 2

[0067] Referring to Figure 2 Embodiment 2 of the present disclosure also provides a processing system for voice 3D driven digital face based on deep regression network, comprising:

[0068] The speech feature processing module 1 is configured to process the speech signal in a pre-emphasis manner to restore the high frequencies of the audio in the speech signal, perform frame processing on the speech signal in a weighted window manner, smooth the speech signal in a sliding window manner, process the windowed speech signal in a Fourier transform manner to obtain the frequency domain signal of the speech signal, filter process the speech signal using a Mel filter to remove high frequency information in the speech signal, perform logarithmic transformation processing on the speech signal from which the high frequency information has been removed to obtain the Mel spectral cepstrum of the speech signal, perform discrete cosine transformation processing on the speech signal with Mel features to obtain the Mel cepstrum features, and process the Mel cepstrum features of the speech using an ASR model extraction manner to remove identity information in the Mel cepstrum features of the speech.

[0069] The 3D face blendshape parameter conversion module 2 is configured to extract speech feature parameters using a network based on a transformer structure, and convert the speech feature parameters into 3D face blendshape parameters.

[0070] The face expression control module 3 is configured to predict the emotional features of the input speech using a speech model, automatically label the speech data, extract the emotional features in the speech using the speech model, merge the speech feature parameters and the corresponding emotional features predicted by the speech model, control the face expression information, and output the corresponding 3D face motion information.

[0071] In this embodiment, the speech feature processing module 1 includes:

[0072] The expression formula of the mel frequency filter is:

[0073]

[0074] In the formula, f is the input frequency; and m is the mel frequency.

[0075] In this embodiment, the speech feature processing module 1 includes:

[0076] The mel frequency cepstrum feature of the speech is processed by using the model extraction mode of the ASR, the semantic feature is extracted, and the timbre and pitch information is removed.

[0077] In this embodiment, the 3D face blendshape parameter conversion module 2 includes:

[0078] After the speech feature parameter is converted into the 3D face blendshape parameter, the wingloss mode is used to process the shaking of the face key point in the regression process, and the expression formula of the wingloss mode is:

[0079]

[0080]

[0081] In the formula, w is used to limit the nonlinear part in the [-w, w] space; ∈ is the curvature of the nonlinear region; x is the independent variable; and C is a constant used to represent the connection between the linear part and the nonlinear part.

[0082] In this embodiment, the face expression control module 3 includes:

[0083] After the corresponding 3D face motion information is output, the face motion information smoothing mode is used to keep the face smooth in the short time frame; and the single person feature fine-tuning mode is used to collect data of a single person and fine-tune the face expression model.

[0084] It should be noted that the information interaction and execution process between the modules of the system described above are based on the same concept as the method embodiment in Embodiment 1 of the present application, and the technical effects brought by the method embodiment are the same as those of the method embodiment of the present application. The specific content can be referred to the description of the method embodiment in the foregoing method embodiment of the present application, and will not be repeated here.

[0085] Embodiment 3

[0086] The embodiment 3 of the present application provides a non-transitory computer readable storage medium, which stores program codes of a voice 3D driven digital face method based on a depth regression network, and the program codes include instructions for executing the voice 3D driven digital face method based on the depth regression network in the embodiment 1 or any possible implementation manner thereof.

[0087] The computer readable storage medium can be any available medium or a data storage device such as a server, data center, etc. integrated with one or more available medium. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)), etc.

[0088] Embodiment 4

[0089] The embodiment 4 of the present application provides an electronic device, which comprises a memory and a processor.

[0090] The processor and the memory complete mutual communication through a bus; the memory stores program instructions executable by the processor, and the processor calling the program instructions can execute the voice 3D driven digital face method based on the depth regression network in the embodiment 1 or any possible implementation manner thereof.

[0091] Specifically, the processor can be implemented by hardware or software, when implemented by hardware, the processor can be a logic circuit, an integrated circuit, etc.; when implemented by software, the processor can be a general-purpose processor, which is implemented by reading software codes stored in a memory, and the memory can be integrated in the processor or exist independently outside the processor.

[0092] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable systems. The computer instructions can be stored in a computer readable storage medium or transferred from one computer readable storage medium to another, for example, the computer instructions can be transferred from one website, computer, server or data center to another through a wired (for example, coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (for example, infrared, wireless, microwave, etc.) manner.

[0093] It should be apparent to those skilled in the art that the modules or steps of the application described above can be implemented with a general purpose computing system, which can be centralized on a single computing system or distributed on a network of multiple computing systems, and optionally, they can be implemented with program codes executable by a computing system, so that they can be stored in a storage system and executed by a computing system, and in some cases, the steps shown or described herein can be executed in a different order, or they can be made into individual integrated circuit modules, or multiple modules or steps can be made into a single integrated circuit module. Thus, the present application is not limited to any particular combination of hardware and software.

[0094] Although the present application has been described in detail with general description and specific embodiments above, some modifications or improvements can be made to the present application, which is obvious to those skilled in the art. Therefore, these modifications or improvements made without departing from the spirit of the present application are within the scope of the present application.

Claims

1. A method for voice-driven 3D digital face recognition based on deep regression networks, characterized in that, include: The speech feature processing module is invoked. Pre-emphasis is used to process the speech signal to restore high frequencies. A weighted window approach is used for frame processing, and a sliding window approach is used to smooth the speech signal between frames. Fourier transform is applied to the windowed speech signal to obtain its frequency domain signal. A Mel frequency filter is used to filter the speech signal, removing high-frequency information. A logarithmic transform is applied to the high-frequency-removed speech signal to obtain its Mel-frequency cepstrum. Discrete cosine transform is applied to the speech signal with Mel features to obtain its Mel-frequency cepstrum features. Finally, an ASR model extraction method is used to process the Mel-frequency cepstrum features, removing identity information from them. The 3D face blendshape parameter conversion module is called, and a network based on the transformer structure is used to extract speech feature parameters and convert the speech feature parameters into 3D face blendshape parameters. The facial expression control module is invoked, and the emotional features of the input speech are predicted using a speech model to automatically label the speech data; the emotional features in the speech are extracted using the speech model; the speech feature parameters and the corresponding emotional features predicted by the speech model are merged to control the facial expression information and output the corresponding 3D facial motion information.

2. The method for voice-driven 3D digital face recognition based on deep regression networks according to claim 1, characterized in that, The formula for the Mel frequency filter is as follows: In the formula, f is the input frequency; m is the Mel frequency.

3. The method for voice-driven 3D digital face recognition based on deep regression networks according to claim 1, characterized in that, The ASR model extraction method is used to process the Mel cepstral features of speech, extract semantic features, and remove timbre and pitch information.

4. The method for voice-driven 3D digital face recognition based on deep regression networks according to claim 1, characterized in that, After converting the speech feature parameters into 3D face blendshape parameters, the wing loss method is used to process the jitter of facial key points during the regression process. The formula for the wing loss method is as follows: In the formula, w restricts the nonlinear part to the space [-w, w]; ∈ is the curvature of the nonlinear region; x is the independent variable; C is a constant used to characterize the connection between the linear and nonlinear parts.

5. The method for voice-driven 3D digital face based on deep regression network according to claim 1, after outputting the corresponding 3D face motion information, maintains the smooth transition of the face within a short frame by smoothing the face motion information; and performs data collection on a single person and fine-tunes the facial expression model by fine-tuning the single-person feature.

6. A voice-driven 3D digital face processing system based on deep regression networks, characterized in that, include: The speech feature processing module is used to process the speech signal using a pre-emphasis method to recover the high frequencies of the audio; to perform frame processing on the speech signal using a weighted window method and to smooth the speech signal between frames using a sliding window method; to process the windowed speech signal using a Fourier transform method to obtain the frequency domain signal of the speech signal; to filter the speech signal using a Mel frequency filter to remove high-frequency information from the speech signal; to perform a logarithmic transform on the speech signal with high-frequency information removed to obtain the Mel frequency cepstrum of the speech signal; to perform a discrete cosine transform on the speech signal with Mel features to obtain the Mel frequency cepstrum features; and to process the Mel frequency cepstrum features of the speech using an ASR model extraction method to remove identity information from the Mel frequency cepstrum features of the speech. The 3D face blendshape parameter conversion module is used to extract speech feature parameters using a transformer-based network and convert the speech feature parameters into 3D face blendshape parameters. The facial expression control module is used to predict the emotional features of the input speech using a speech model, automatically label the speech data, and extract the emotional features from the speech using the speech model. The speech feature parameters and the corresponding emotional features predicted by the speech model are merged to control facial expression information and output corresponding 3D facial motion information.

7. The voice-driven 3D digital face processing system based on deep regression networks according to claim 6, characterized in that, In the speech feature processing module: The formula for the Mel frequency filter is as follows: In the formula, f is the input frequency; m is the Mel frequency.

8. The voice-driven 3D digital face processing system based on deep regression networks according to claim 6, characterized in that, In the speech feature processing module: The ASR model extraction method is used to process the Mel cepstral features of speech, extract semantic features, and remove timbre and pitch information.

9. The voice-driven 3D digital face processing system based on deep regression networks according to claim 6, characterized in that, In the 3D face blendshape parameter conversion module: After converting the speech feature parameters into 3D face blendshape parameters, the wing loss method is used to process the jitter of facial key points during the regression process. The formula for the wing loss method is as follows: In the formula, w restricts the nonlinear part to the space [-w, w]; ∈ is the curvature of the nonlinear region; x is the independent variable; C is a constant used to characterize the connection between the linear and nonlinear parts.

10. The voice-driven 3D digital face processing system based on deep regression networks according to claim 6, characterized in that, In the facial expression control module: After outputting the corresponding 3D face motion information, the face is kept to transition smoothly within a short frame by smoothing the face motion information; data is collected from a single person by fine-tuning the individual features to fine-tune the facial expression model.

Citation Information

Patent Citations

  • Identification method and system based on multi-modal data fusion

    CN116934926A

  • Method and device for generating broadcast video of digital human based on driving text

    CN117711042A