Model training method, facial feature extraction method, equipment, medium and software
By performing spectral conversion on speech data, extracting style and content features, and jointly training model parameters, the problem of lack of personalized style in large speech models is solved, a high degree of adaptation of digital human facial features and speaker style is achieved, and the accuracy of facial driving is improved.
Patent Information
- Application Number
- CN202510331586.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-09-23
AI Technical Summary
Existing large speech models lack personalized style when driving digital human facial movements, resulting in the feature vector being unable to accurately reflect the characteristics of a specific speaker.
By performing spectral conversion on the speech data, style features and content features are extracted separately, and the model parameters are adjusted by joint training to ensure that the style features are personalized and the content features are not personalized. Facial features are generated by combining the style features and content features.
It achieves a high degree of adaptation between the digital human’s facial features and the speaker’s style, improving the accuracy and personalized performance of facial driving.
Smart Images

Figure CN120689914A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of Internet technology, and in particular to a model training method, a facial feature extraction method, a computer device, a storage medium (computer-readable storage medium), and computer software. Background Art
[0002] Digital human technology is increasingly being applied in industries such as virtual reality and filmmaking, and using speech to drive facial movements in digital humans is gaining increasing attention. Related technologies use large speech models to directly learn the mapping between a speaker's speech and feature vectors, which are used to drive the facial movements of the digital human. However, because large speech models can only capture and mimic averaged features to accommodate speech input from multiple speakers, the feature vectors inferred by these large speech models lack personalized style for a specific speaker. Summary of the Invention
[0003] Embodiments of the present application provide a model training method, a facial feature extraction method, a computer device, a storage medium (computer-readable storage medium), and computer software.
[0004] In a first aspect, an embodiment of the present application provides a model training method, the method comprising:
[0005] Performing spectrum conversion processing on the first voice data to obtain first spectrum data of the first voice data;
[0006] Inputting the first spectrum data into a first model to perform first feature extraction to obtain a style feature of the first speech data;
[0007] Inputting the first spectrum data into a second model to perform second feature extraction to obtain content features of the first speech data;
[0008] reconstructing a spectrum of the first speech data according to the style feature and the content feature of the first speech data;
[0009] Adjust parameters of the first model and parameters of the second model according to a first loss value extracted from the first feature, a second loss value extracted from the second feature, and a spectrum reconstruction loss value of the spectrum reconstruction.
[0010] In a second aspect, an embodiment of the present application provides a facial feature extraction method, the method comprising:
[0011] performing spectrum conversion processing on the third speech data to obtain second spectrum data;
[0012] Inputting the second spectrum data into the first model to perform first feature extraction to obtain style features;
[0013] Inputting the third speech data into the speech model to perform second feature extraction to obtain content features unrelated to the style features;
[0014] The style features and the content features are input into a mapping model for feature mapping processing to obtain facial features.
[0015] In a third aspect, an embodiment of the present application also provides a computer device comprising a processor and a memory; the memory stores a plurality of instructions; the processor loads instructions from the memory to execute the steps of any model training method provided in the embodiment of the present application, or to execute the steps of the facial feature extraction method provided in the embodiment of the present application.
[0016] Fourthly, an embodiment of the present application also provides a computer-readable storage medium, which stores multiple instructions, and the instructions are suitable for a processor to load to execute the steps of any model training method provided in the embodiment of the present application, or to execute the steps of the facial feature extraction method provided in the embodiment of the present application.
[0017] In a fifth aspect, an embodiment of the present application also provides a computer program product, including a computer program or instructions, which, when executed by a processor, implements the steps of any model training method provided in the embodiment of the present application, or executes the steps of the facial feature extraction method provided in the embodiment of the present application.
[0018] The embodiment of the present application can decouple the style features from the content features during the training phase, and then jointly train the first model and the second model to improve the accuracy of model training. Since the content features output by the second model and the style features output by the first model are used to reconstruct the spectral data, and the content features output by the second model do not contain personalized style, in order to ensure the accuracy of spectral reconstruction, the first model can only output style features indicating personalized style, thereby constraining the style feature extraction of the first model and further improving the accuracy of the first model. In addition, the embodiment of the present application extracts style features through the first model on the one hand, and extracts content features that are unrelated to style features through the large speech model on the other hand, and then combines the style features and content features to determine facial features. When determining facial features, the embodiment of the present application uses style features as one of the bases, and can integrate the speaker's personalized style into the facial features to achieve a high degree of adaptation between the facial features and the speaker's style. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0020] Figure 1 is a schematic diagram of a drive system provided in an embodiment of the present application;
[0021] Figure 2 This is a flowchart of a facial feature extraction method provided by an embodiment of the present application;
[0022] Figure 3 is a schematic diagram of a facial feature extraction process provided in an embodiment of the present application;
[0023] Figure 4 This is a flow chart of a model training method provided in an embodiment of the present application;
[0024] Figure 5 is a schematic diagram of a feature decoupling module provided in an embodiment of the present application;
[0025] Figure 6 This is a schematic diagram of a model joint training process provided in an embodiment of the present application;
[0026] Figure 7 This is a flowchart of another model training method provided in an embodiment of the present application;
[0027] Figure 8 This is a schematic diagram of another model joint training process provided in an embodiment of the present application. DETAILED DESCRIPTION
[0028] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of this application. At the same time, in the description of the embodiments of the present application, the terms "first", "second", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more features. In the description of the embodiments of the present application, the meaning of "multiple" is two or more, unless otherwise clearly and specifically defined.
[0029] See also Figure 1 , Figure 1This is a schematic diagram of a drive system provided in an embodiment of the present application. Figure 1 As shown, the driving system includes: a server 10 and a terminal 20.
[0030] In the embodiments of the present application, facial features can be determined based on the speaker's voice data, and then the facial movements of the digital human can be driven based on the facial features. The speaker's voice data can be processed using an artificial intelligence model, such as the style encoding model, voice macro model, and mapping model described in the following embodiments, to determine the facial features.
[0031] Server 10 is used to train the artificial intelligence model. Since model training generally requires a large amount of data and computing power, the embodiment of the present application performs model training through server 10. Optionally, server 10 can be a single server or a computer cluster consisting of multiple servers.
[0032] Terminal 20 is used to collect the speaker's voice data. Terminal 20 can be a smartphone, wearable device, tablet computer, desktop computer, intelligent robot, virtual reality device, smart home device, smart car device, etc. The user can interact with Terminal 20, such as recording audio and video, so that Terminal 20 can collect the speaker's voice data.
[0033] In some embodiments, the server can determine facial features based on the speaker's voice data. Thus, terminal 20 is further configured to transmit the speaker's voice data to server 10; server 10 is further configured to process the speaker's voice data using a trained artificial intelligence model to determine facial features and transmit the facial features to terminal 20.
[0034] In some embodiments, the terminal can also determine facial features based on the speaker's voice data. Thus, the server 10 can deploy the trained artificial intelligence model to the terminal 20; the terminal 20 processes the speaker's voice data using the trained artificial intelligence model to determine facial features.
[0035] like Figure 1 As shown, the server 10 and the terminal 20 can communicate via a network. The network can be a wireless network or a wired network.
[0036] The following describes the model training method and facial feature extraction method provided in the embodiments of the present application. It should be understood that, for ease of distinction, the embodiments of the present application refer to the voice data used in the model training method as first voice data and second voice data, and the spectral data obtained in the model training method as first spectral data; the voice data used in the facial feature extraction method as third voice data, and the spectral data obtained in the facial feature extraction method as second spectral data; this does not constitute a limitation of the present application.
[0037] See also Figure 2 , Figure 2 This is a flow chart of a facial feature extraction method provided by an embodiment of the present application. The facial feature extraction method can be applied to Figure 1 In the drive system shown, for example, Figure 1 Executed by the server 10 in Figure 1 The terminal 20 in the process executes the Figure 2 As shown, the facial feature extraction method includes the following steps:
[0038] Step 110: performing spectrum conversion processing on the third speech data to obtain second spectrum data;
[0039] Step 120: Input the second spectrum data into the first model to extract the first feature to obtain the style feature;
[0040] Step 130: Input the third speech data into the speech model to perform second feature extraction to obtain content features unrelated to the style features;
[0041] Step 140: Input the style features and content features into the mapping model for feature mapping processing to obtain facial features.
[0042] Stylistic features are those related to the speaker's style, while content features are those unrelated to the speaker's style. In other words, content features are unrelated to stylistic features. A speaker's style is a unique characteristic that distinguishes different speakers. Typically, a speaker's style doesn't change with their mood, the context of their speech, or other factors. Instead, it remains consistent across different scenarios.
[0043] In some embodiments, the stylistic features include at least one of the following: a speaker's mouth shape and a speaker's timbre. Mouth shape refers to the shape of the mouth used to pronounce a sound, which varies between speakers. Timbre refers to the texture of the sound, which varies with the waveform of the sound.
[0044] In some embodiments, content features include at least one of the following: a speaker's pitch feature, a speaker's pitch variation feature, a speaker's volume feature, and a background sound feature. Pitch features, also known as tone features, refer to the highness of a sound; the higher the frequency of a sound, the higher the pitch or tone; the lower the frequency of a sound, the lower the pitch or tone. Pitch variation features refer to the regularity or pattern of pitch or tone changes over time and / or context. Volume features refer to the intensity or loudness of a sound; the greater the amplitude of a sound, the louder the volume; the smaller the amplitude of a sound, the lower the volume. Background sound features refer to sounds in the background environment, such as noise.
[0045] The style features are extracted by the first model. In step 110, the third speech data is converted into second spectrum data; in step 120, the second spectrum data is input into the first model, so that the first model outputs the style features. Optionally, the spectrum conversion processing in step 110 can be Mel spectrum conversion, Fourier transform, cepstral coefficient conversion, etc. Optionally, the first model can be composed of a two-dimensional convolutional neural network and an attention mechanism. Of course, it can also be other network structures, such as recurrent neural networks, transformers, etc. In actual applications, the corresponding network can be flexibly selected in combination with computational efficiency and accuracy requirements to construct the first model.
[0046] Content features are extracted by the speech model. In step 130, the third speech data is input into the speech model, which then outputs content features. Optionally, the speech model can be a model such as Wav2Vec or HuberT. The speech model can be a trained speech model from related art. That is, in this embodiment of the present application, the parameters of the speech model are frozen, eliminating the need for further training.
[0047] In step 140, the style features and content features are input into the mapping model, which then performs feature mapping and outputs facial features. These facial features can be used to drive the facial motion of the digital human. In the embodiments of the present application, the facial features can be latent space vectors of motion (Motion), such as facial representations in BlendShape, BFM, or Flame head models. Alternatively, the mapping model can be composed of a one-dimensional convolution and an attention mechanism.
[0048] In some embodiments, after step 140, the facial feature extraction method may further include: driving the facial movement of the digital human based on the facial features. Driving the facial movement of the digital human based on the facial features includes: inputting the facial features into a third model for parameter mapping processing to obtain facial parameters; and driving the facial movement of the digital human based on the facial parameters. The third model is used to map facial features into facial parameters. Optionally, the third model can be constructed based on a Transformer architecture, or the third model can be a FLAME model, etc. Facial parameters include, but are not limited to, facial expressions (such as smiling, laughing, frowning, crying, anger, fear, etc.) and facial movements (such as lowering the head, turning left, turning right, shaking the head, nodding, etc.).
[0049] See also Figure 3 , Figure 3 This is a schematic diagram of a facial feature extraction method provided in an embodiment of the present application. The facial feature extraction method can be applied to Figure 1 In the drive system shown, for example, Figure 1 Executed by the server 10 in Figure 1 The terminal 20 in the process executes the command.
[0050] like Figure 3 As shown, the third speech data is spectrally converted to obtain second spectral data, which is then processed by the first model to obtain style features. Furthermore, the third speech data is processed by the frozen large speech model to obtain content features. The content features and style features are then input into the mapping model, which outputs facial features. These facial features can then be used to drive the facial movements of the digital human.
[0051] In summary, the facial feature extraction method provided in the embodiment of the present application, on the one hand, extracts style features through a first model, and on the other hand, extracts content features unrelated to the style features through a large speech model, and then combines the style features and content features to determine the facial features. When determining facial features, the embodiment of the present application uses style features as one of the bases, and can integrate the speaker's personalized style into the facial features to achieve a high degree of adaptation of the facial features to the speaker's style. In addition, the facial features extracted by the embodiment of the present application can be used to drive the facial movement of a digital human, thereby improving the accuracy of the digital human's facial drive and achieving style adaptation of the digital human's facial drive to any speaker.
[0052] See also Figure 4 , Figure 4 This is a flow chart of a model training method provided in an embodiment of the present application. The model training method can be applied to Figure 1 In the drive system shown, for example, Figure 1The server 10 in the embodiment is executed. Figure 4 As shown, the model training method includes the following steps:
[0053] Step 410: performing spectrum conversion processing on the first speech data to obtain first spectrum data of the first speech data;
[0054] Step 420: Input the first spectrum data into the first model to perform first feature extraction to obtain the style feature of the first speech data;
[0055] Step 430: Input the first spectrum data into the second model to perform second feature extraction to obtain content features of the first speech data;
[0056] Step 440: reconstructing the spectrum of the first speech data according to the style characteristics and content characteristics of the first speech data;
[0057] Step 450: Adjust the parameters of the first model and the parameters of the second model according to the first loss value extracted from the first feature, the second loss value extracted from the second feature, and the spectrum reconstruction loss value of the spectrum reconstruction.
[0058] The first voice data may be the complete voice data of the speaker or a segment of the complete voice data of the speaker. For ease of distinction, the embodiment of the present application refers to the complete voice data of the speaker as the second voice data. The duration of the first voice data may be less than or equal to the duration of the second voice data.
[0059] The first voice data and the second voice data can be voice data of any one or more speakers, and the speakers of the first voice data and the second voice data may include the speaker of the third voice data, or the speakers of the first voice data and the second voice data may not include the speaker of the third voice data. To improve the accuracy of model training, the speakers of the first voice data and the second voice data can be multiple speakers. The embodiment of the present application constructs a training data set based on the first voice data and the second voice data. Optionally, each training sample in the training data set includes a second voice data and facial features corresponding to the second voice data.
[0060] In some embodiments, for any second voice data, random resampling can be performed along the time dimension to obtain at least two first voice data. Optionally, the duration of any two first voice data may be the same or different; the duration of any first voice data may be the same or different from the duration of the corresponding second voice data. The embodiment of the present application does not limit the specific method of data resampling processing. It should be understood that the data resampling processing can arbitrarily adjust the content of the second voice data without changing the speaker's style implied in the second voice data.
[0061] In step 410, each first voice data is subjected to spectrum conversion processing to obtain the first spectrum data of the first voice data. In step 420, the first spectrum data of each first voice data is processed by the first model to obtain the style features of the first voice data. In step 430, the first spectrum data of each first voice data is processed by the second model to obtain the content features of the first voice data. Then in step 440, for each first voice data, the spectrum of the first voice data is reconstructed according to the style features and content features of the first voice data. Optionally, the second model can be composed of a two-dimensional convolutional neural network and an attention mechanism. Of course, it can also be other network structures, such as a recurrent neural network, a transformer, etc. In actual applications, the corresponding network can be flexibly selected to construct the second model in combination with computational efficiency, accuracy requirements, etc. Optionally, the second model can have the same network structure as the first model.
[0062] Based on step 420, a first loss value for first feature extraction can be constructed, based on step 430, a second loss value for second feature extraction can be constructed, and based on step 440, a spectrum reconstruction loss value for spectrum reconstruction can be constructed. Optionally, the method for determining the first loss value for first feature extraction includes: determining the first loss value for first feature extraction based on the style features of the first speech data. For example, the first loss value for first feature extraction can be determined based on the style features of at least two first speech data corresponding to the same second speech data. Optionally, the method for determining the second loss value for second feature extraction includes: inputting the first speech data into a large speech model for second feature extraction to obtain reference features for the first speech data; determining a second loss value for the first speech data based on the content features and the reference features of the first speech data; and determining a second loss value for the second feature extraction based on the second loss value of the first speech data. The reference feature is a universal feature of the speaker extracted by the large speech model, does not contain the speaker's personalized style, is a deep-level feature of the audio, and can serve as a reference for the true value of the content feature. For example, for each first speech data, a reference feature of the first speech data is obtained using a large speech model, and then a second loss value of the first speech data is determined based on the content feature and the reference feature of the first speech data; and then a second loss value for second feature extraction is determined based on the second loss values of at least two first speech data. Optionally, the loss value in the embodiment of the present application can be calculated using the L1 norm, Euclidean distance, or cosine similarity.
[0063] In order to improve the accuracy of model training, the embodiment of the present application adopts a joint training method. In step 450, the parameters of the first model and the parameters of the second model are adjusted according to the first loss value extracted by the first feature, the second loss value extracted by the second feature, and the spectrum reconstruction loss value of the spectrum reconstruction. For example, the joint training loss value can be determined based on the first loss value extracted by the first feature, the second loss value extracted by the second feature, and the spectrum reconstruction loss value of the spectrum reconstruction; then the parameters of the first model and the parameters of the second model are adjusted to converge the joint training loss value. When the joint training loss value converges, the first model and the second model that have completed training can be obtained. Optionally, the first loss value extracted by the first feature, the second loss value extracted by the second feature, and the spectrum reconstruction loss value of the spectrum reconstruction can be summed or weighted to determine the joint training loss value.
[0064] In some embodiments, spectrum reconstruction can be achieved through a decoding model; based on this, the above-mentioned step 450 includes: inputting the style features and content features of the first speech data into the decoding model for spectrum reconstruction to obtain reconstructed spectrum data of the first speech data. Based on this, the above-mentioned method for determining the spectrum reconstruction loss value may include: determining the spectrum reconstruction loss value of the first speech data based on the first spectrum data and the reconstructed spectrum data of the first speech data; and determining the spectrum reconstruction loss value of the spectrum reconstruction based on the spectrum reconstruction loss value of the first speech data. For example, for each first speech data, the spectrum reconstruction loss value of the first speech data is determined based on the first spectrum data and the reconstructed spectrum data of the first speech data; and then the spectrum reconstruction loss value of the spectrum reconstruction is determined based on the spectrum reconstruction loss values of at least two first speech data. Based on this, the above-mentioned step 450 may include: adjusting the parameters of the first model, the parameters of the second model, and the parameters of the decoding model to converge the joint training loss value. Optionally, the decoding model may be composed of a deconvolution layer, a pooling layer, and an upsampling layer.
[0065] In the embodiment of the present application, the first model, the second model and the decoding model can exist independently or be integrated into one module, for example, into a feature decoupling module. Figure 5 , Figure 5 Schematic diagram of a feature decoupling module provided in an embodiment of the present application. Figure 5As shown, the feature decoupling module includes a first model 510, a second model 520, a decoding model 530 and a frozen large speech model 540. For each speech data, the feature decoupling module first performs spectrum conversion processing on the speech data to obtain spectrum data; then, on the one hand, the spectrum data is input into the first model 510 to obtain style features, and on the other hand, the spectrum data is input into the second model 520 to obtain content features; thereafter, the style features and content features are input into the decoding model 530 to obtain reconstructed spectrum data; in addition, in the feature decoupling module, the speech data is also input into the frozen large speech model 540 to obtain reference features. Based on Figure 5 The feature decoupling module shown can determine a first loss value for first feature extraction based on style features; a second loss value for second feature extraction based on content features and reference features; and a spectral reconstruction loss value for spectral reconstruction based on spectral data and reconstructed spectral data. The first loss value for first feature extraction, the second loss value for second feature extraction, and the spectral reconstruction loss value for spectral reconstruction are combined to further determine a joint training loss value. The parameters of the first model 510, second model 520, and decoding model 530 in the feature decoupling module are then adjusted to achieve convergence of the joint training loss value. When the joint training loss value converges, the trained first model 510, second model 520, and decoding model 530 are obtained.
[0066] In some embodiments, the method further includes: inputting the second speech data into a large speech model to extract a second feature, thereby obtaining reference features for the second speech data; and inputting the style features of the first speech data and the reference features of the second speech data into a mapping model for feature mapping to obtain facial features for the first speech data. Optionally, the mapping model can be composed of a one-dimensional convolution and an attention mechanism, which can learn the mapping relationship between the style features, the reference features (content features), and the facial features, so that the style features and the reference features (content features) can be input into the mapping model to map the corresponding facial features. Based on this, step 450 includes adjusting the parameters of the first model, the second model, and the mapping model based on the first loss value of the first feature extraction, the second loss value of the second feature extraction, the spectral reconstruction loss value of the spectral reconstruction, and the feature mapping loss value of the feature mapping process. For example, a joint training loss value can be determined based on the first loss value of the first feature extraction, the second loss value of the second feature extraction, the spectral reconstruction loss value of the spectral reconstruction, and the feature mapping loss value of the feature mapping process; and then adjusting the parameters of the first model, the second model, and the mapping model to achieve convergence of the joint training loss value. Optionally, the method for determining the feature mapping loss value of feature mapping processing includes: determining the feature mapping loss value of the first speech data based on the facial features of the first speech data and the facial features of the second speech data; and determining the feature mapping loss value of the feature mapping processing based on the feature mapping loss value of the first speech data. For example, for each first speech data, the feature mapping loss value of the first speech data is determined based on the facial features of the first speech data and the facial features of the second speech data corresponding to the first speech data; and then determining the feature mapping loss value of the feature mapping processing based on the feature mapping loss values of at least two first speech data. The facial features of the second speech data may be included in the training dataset.
[0067] See also Figure 6 , Figure 6 This is a schematic diagram of a model joint training process provided by an embodiment of the present application. Figure 6 As shown, for each speech data, the speech data is input into the feature decoupling module 610, and the style feature, content feature, reference feature, spectrum data and reconstructed spectrum data can be obtained. The structure of the feature decoupling module 610 can refer to Figure 5As shown. The style features and reference features are then input into the mapping model 620 so that the mapping model 620 outputs facial features. Based on this, the first loss value for first feature extraction can be determined based on the style features; the second loss value for second feature extraction can be determined based on the content features and the reference features; the spectrum reconstruction loss value for spectrum reconstruction can be determined based on the spectrum data and the reconstructed spectrum data; and the feature mapping loss value for feature mapping processing can be determined based on the facial features and the facial features corresponding to the speech data in the training dataset. The first loss value for first feature extraction, the second loss value for second feature extraction, the spectrum reconstruction loss value for spectrum reconstruction, and the feature mapping loss value for feature mapping processing can be combined to further determine a joint training loss value, and then the parameters of the first model, second model, decoding model, and mapping model 620 in the feature decoupling module 610 are adjusted to converge the joint training loss value. When the joint training loss value converges, the trained first model, second model, decoding model, and mapping model can be obtained.
[0068] In summary, the model training method provided by the embodiment of the present application can decouple the style features from the content features during the training phase, and then jointly train the first model and the second model to improve the accuracy of the model training. In addition, the embodiment of the present application can use the reference features output by the large speech model as the true value to supervise the content features output by the second model, thereby ensuring that the content features output by the second model do not contain a personalized style. In addition, since the content features output by the second model and the style features output by the first model are used to reconstruct the spectral data, and the content features output by the second model do not contain a personalized style, in order to ensure the accuracy of the spectral reconstruction, the first model can only output style features indicating a personalized style, thereby constraining the style feature extraction of the first model and further improving the accuracy of the first model.
[0069] In addition, the embodiment of the present application performs data resampling processing on each second voice data to obtain the first voice data, and then performs model training based on the facial features corresponding to the first voice data and the second voice data. On the one hand, the content of the second voice data can be arbitrarily adjusted through data resampling processing without changing the speaker's style contained in the second voice data, thereby effectively expanding the sample data and improving the diversity and richness of the training samples. On the other hand, based on different first voice data of the same second voice data, the predicted style features should be consistent, so that the loss value is constructed and integrated into the joint training loss value to perform model joint training, which can improve the accuracy of model training.
[0070] Below, an example is used to introduce and illustrate the model training process in the embodiment of the present application.
[0071] See also Figure 7 , Figure 7This is a flow chart of a model training method provided in an embodiment of the present application. The model training method can be applied to Figure 1 In the drive system shown, for example, Figure 1 The server 10 in the embodiment is executed. Figure 7 As shown, the model training method includes the following steps:
[0072] Step 701: resampling the second voice data to obtain first voice data;
[0073] Step 702: Perform spectrum conversion processing on the first speech data to obtain first spectrum data of the first speech data;
[0074] Step 703: Input the first spectrum data into the first model to perform first feature extraction to obtain the style feature of the first speech data;
[0075] Step 704: determining a first loss value for first feature extraction based on the style feature of the first speech data;
[0076] Step 705: Input the first spectrum data into the second model to perform second feature extraction to obtain content features of the first speech data;
[0077] Step 706: Input the first speech data into the speech model to perform second feature extraction to obtain reference features of the first speech data;
[0078] Step 707: Determine a second loss value of the first speech data based on the content feature and the reference feature of the first speech data;
[0079] Step 708: Determine a second loss value for second feature extraction based on the second loss value of the first speech data;
[0080] Step 709: Inputting the style features and content features of the first speech data into the decoding model to perform spectrum reconstruction to obtain reconstructed spectrum data of the first speech data;
[0081] Step 710: Determine a spectrum reconstruction loss value of the first speech data based on the first spectrum data and the reconstructed spectrum data of the first speech data;
[0082] Step 711: determining a spectrum reconstruction loss value for spectrum reconstruction according to the spectrum reconstruction loss value of the first speech data;
[0083] Step 712: Input the second speech data into the speech model to extract the second feature, and obtain the reference feature of the second speech data;
[0084] Step 713: Inputting the style features of the first speech data and the reference features of the second speech data into the mapping model for feature mapping processing to obtain facial features of the first speech data;
[0085] Step 714: Determine a feature mapping loss value of the first speech data based on the facial features of the first speech data and the facial features of the second speech data;
[0086] Step 715: Determine a feature mapping loss value for feature mapping processing according to the feature mapping loss value of the first speech data;
[0087] Step 716: Determine a joint training loss value based on the first loss value of the first feature extraction, the second loss value of the second feature extraction, the spectrum reconstruction loss value of the spectrum reconstruction, and the feature mapping loss value of the feature mapping processing;
[0088] Step 717: Adjust the parameters of the first model, the second model, the decoding model, and the mapping model according to the joint training loss value.
[0089] See also Figure 8 , Figure 8 This is a schematic diagram of a model training process provided in an embodiment of the present application.
[0090] like Figure 8 As shown, for a second voice data in the training data set, data resampling processing is performed on the second voice data to obtain first voice data A and first voice data B.
[0091] For the first speech data A, the first speech data A is first subjected to spectrum conversion processing to obtain the first spectrum data of the first speech data A; the first spectrum data is input into the first model 810 to obtain the style features of the first speech data A; the first spectrum data is input into the second model 820 to obtain the content features of the first speech data A; the style features and content features of the first speech data A are input into the decoding model 830 to obtain the reconstructed spectrum data of the first speech data A; in addition, the first speech data A is input into the frozen speech large model 840 to obtain the reference features of the first speech data A.
[0092] For the first speech data B, the first speech data B is first subjected to spectrum conversion processing to obtain the first spectrum data of the first speech data B; the first spectrum data is input into the first model 810 to obtain the style features of the first speech data B; the first spectrum data is input into the second model 820 to obtain the content features of the first speech data B; the style features and content features of the first speech data B are input into the decoding model 830 to obtain the reconstructed spectrum data of the first speech data B; in addition, the first speech data B is input into the frozen speech large model 840 to obtain the reference features of the first speech data B.
[0093] In order to ensure that the first model only extracts style features, the embodiment of the present application can limit the maximum length of the latent of the first model, that is, limit the maximum number of style features, such as limiting the maximum number of style features to 30.
[0094] Taking the loss value calculated by L1 norm as an example, the second loss value and spectrum reconstruction loss value of the first speech data A can be shown in the following formulas 1 and 4, the second loss value and spectrum reconstruction loss value of the first speech data B can be shown in the following formulas 2 to 5, the second loss value of the second feature extraction can be shown in the following formula 3, and the spectrum reconstruction loss value of the spectrum reconstruction can be shown in the following formula 6.
[0095] Formula 1: Style-independent loss A=L1 (Style-independent latent A, Universal features Latent A)
[0096] Formula 2: Style-independent loss B=L1 (Style-independent latent B, Universal features Latent B)
[0097] Formula 3: Total Style-independent loss=0.5*Style-independent loss A+0.5*Style-independent loss B
[0098] Formula 4: Mel reconstruct loss A=L1(predict Mel A, ground truth Mel A)
[0099] Formula 5: Mel reconstruct loss B=L1(predict Mel B, ground truth Mel B)
[0100] Formula 6: Total Mel reconstruct loss=0.5*Mel reconstruct loss A+0.5*Mel reconstruct loss B
[0101] Among them, Style-independent loss A refers to the second loss value of the first speech data A, Style-independent latent A refers to the content feature of the first speech data A, and Universal features Latent A refers to the reference feature of the first speech data A; Style-independent loss B refers to the second loss value of the first speech data B, Style-independent latent B refers to the content feature of the first speech data B, and Universal features Latent B refers to the reference feature of the first speech data B; Total Style-independent loss refers to the second loss value of the second feature extraction; Mel reconstruct loss A refers to the spectrum reconstruction loss value of the first speech data A, predicted Mel A refers to the reconstructed spectrum data of the first speech data A, and ground truth Mel A refers to the first spectrum data of the first speech data A; Mel reconstruct loss B refers to the spectrum reconstruction loss value of the first speech data B, predicted Mel B refers to the reconstructed spectrum data of the first speech data B, and ground truth Mel B refers to the first spectrum data of the first speech data B; Total Mel reconstruct loss refers to the spectrum reconstruction loss value of spectrum reconstruction.
[0102] Because the first speech data A and the first speech data B are derived from the same second speech data, the speaker's style should be consistent between the first speech data A and the first speech data B. Taking the loss value calculated using the L1 norm as an example, the first loss value for the first feature extraction can be expressed as follows: Formula 7.
[0103] Formula 7: Total Style loss=L1(Style Latent A,Style Latent B)
[0104] Here, Style loss refers to the first loss value extracted by the first feature, Style Latent A refers to the style feature of the first speech data A, and Style Latent B refers to the style feature of the first speech data B.
[0105] In addition, if Figure 8As shown, the complete second speech data can be input into the frozen speech macro model 840 to obtain reference features of the second speech data. The style features of the first speech data A and the reference features of the second speech data are input into the mapping model 850 to obtain the facial features of the first speech data A; the style features of the first speech data B and the reference features of the second speech data are input into the mapping model 850 to obtain the facial features of the first speech data B.
[0106] Taking the loss value calculated by L1 norm as an example, the feature mapping loss value of the first speech data A can be shown as follows in Formula 8, the feature mapping loss value of the first speech data B can be shown as follows in Formula 9, and the feature mapping loss value of feature mapping processing can be shown as follows in Formula 10.
[0107] Formula 8: motion loss A=L1 (predict Motion A, ground truth Motion)
[0108] Formula 9: motion loss B=L1 (predict Motion B, ground truth Motion)
[0109] Formula 10: Total motion loss=0.5*motion loss A+0.5*motion loss B
[0110] Among them, motion loss A refers to the feature mapping loss value of the first voice data A, predicted Motion A refers to the facial features of the first voice data A; motion loss B refers to the feature mapping loss value of the first voice data B, predicted Motion B refers to the facial features of the first voice data B; ground truth Motion refers to the facial features corresponding to the second voice data; Total motion loss refers to the feature mapping loss value of feature mapping processing.
[0111] Based on this, the joint training loss value Total loss can be expressed as follows:
[0112] Formula 11: Total loss=Total Style loss+Total Style-independent loss+Total Mel reconstructive loss+2*Motion loss
[0113] Since the parameters of the large speech model 840 are frozen, no parameter adjustment is required during the training process. Therefore, the parameters of the first model 810, the second model 820, the decoding model 830, and the mapping model 850 are adjusted to converge the joint training loss. When the joint training loss converges, the first model 810, the second model 820, the decoding model 830, and the mapping model 850 are trained.
[0114] Based on the same inventive concept, an embodiment of the present application further provides a computer device, which may be a server, comprising a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps of the above-mentioned facial feature extraction method. This implements various functions, such as:
[0115] performing spectrum conversion processing on the third speech data to obtain second spectrum data;
[0116] Inputting the second spectrum data into the first model to perform first feature extraction to obtain style features;
[0117] Inputting the third speech data into the speech model to perform second feature extraction to obtain content features unrelated to the style features;
[0118] The style features and the content features are input into a mapping model for feature mapping processing to obtain facial features.
[0119] The computer device extracts style features using the first model and extracts content features unrelated to the style features using the large speech model. The style and content features are then combined to determine facial features. This embodiment of the present application uses style features as one of the criteria for determining facial features. This allows the speaker's personalized style to be integrated into the facial features, achieving a high degree of adaptation between the facial features and the speaker's style. Furthermore, the facial features extracted by this embodiment of the present application can be used to drive the facial movements of a digital human, thereby improving the accuracy of the digital human's facial drive and enabling adaptation of the digital human's facial drive to the style of any speaker.
[0120] Alternatively, the processor implements the steps of the above-mentioned model training method when executing the computer program, thereby achieving various functions, such as:
[0121] Performing spectrum conversion processing on the first voice data to obtain first spectrum data of the first voice data;
[0122] Inputting the first spectrum data into a first model to perform first feature extraction to obtain a style feature of the first speech data;
[0123] Inputting the first spectrum data into a second model to perform second feature extraction to obtain content features of the first speech data;
[0124] reconstructing a spectrum of the first speech data according to the style feature and the content feature of the first speech data;
[0125] Adjust parameters of the first model and parameters of the second model according to a first loss value extracted from the first feature, a second loss value extracted from the second feature, and a spectrum reconstruction loss value of the spectrum reconstruction.
[0126] The above-mentioned computer device can decouple the style features from the content features during the training phase, and then jointly train the first model and the second model to improve the accuracy of the model training. In addition, the embodiment of the present application can use the reference features output by the large speech model as the true value to supervise the content features output by the second model, thereby ensuring that the content features output by the second model do not contain a personalized style. In addition, since the content features output by the second model and the style features output by the first model are used to reconstruct the spectral data, and the content features output by the second model do not contain a personalized style, in order to ensure the accuracy of the spectral reconstruction, the first model can only output style features indicating a personalized style, thereby constraining the style feature extraction of the first model and further improving the accuracy of the first model.
[0127] The specific implementation of the above operations can be found in the previous embodiments and will not be repeated here.
[0128] Based on the same inventive concept, an embodiment of the present application also provides a computer-readable storage medium, which may include: a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk, etc.
[0129] Since the computer program stored in the computer-readable storage medium can execute any facial feature extraction method or model training method provided in the embodiments of the present application, the beneficial effects that can be achieved by any facial feature extraction method or model training method provided in the embodiments of the present application can be achieved. Please refer to the previous embodiments for details and will not be repeated here.
[0130] Based on the same inventive concept, embodiments of the present application further provide a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in various optional implementations of the above embodiments.
[0131] It should be noted that the object data (including but not limited to user device information, user personal information, etc.) and conversation data involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions. Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods.
[0132] Any reference to the memory, database or other media used in the various embodiments provided herein may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0133] The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may include, but are not limited to, general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic units, data processing logic units based on quantum computing, and the like.
[0134] In the above-mentioned embodiments of the model training method, facial feature extraction method, computer device, storage medium, and computer software, the description of each embodiment has its own focus. For parts not described in detail in a certain embodiment, reference can be made to the relevant description of other embodiments. Those skilled in the art will clearly understand that for the convenience and brevity of description, the specific working processes and beneficial effects of the above-mentioned model training method, facial feature extraction method, computer device, storage medium, computer software, and their corresponding units can be referred to the description of the model training method and facial feature extraction method in the above embodiments, and the details will not be repeated here.
[0135] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0136] The above is a detailed introduction to a model training method, facial feature extraction method, computer equipment, storage medium and computer software provided in the embodiments of the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for technical personnel in this field, based on the ideas of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A model training method, characterized in that: The method comprises: Performing spectrum conversion processing on the first voice data to obtain first spectrum data of the first voice data; Inputting the first spectrum data into a first model to perform first feature extraction to obtain a style feature of the first speech data; Inputting the first spectrum data into a second model to perform second feature extraction to obtain content features of the first speech data; reconstructing a spectrum of the first speech data according to the style feature and the content feature of the first speech data; Adjust parameters of the first model and parameters of the second model according to a first loss value extracted from the first feature, a second loss value extracted from the second feature, and a spectrum reconstruction loss value of the spectrum reconstruction.
2. The method according to claim 1, characterized in that The method for determining the second loss value extracted by the second feature includes: Inputting the first speech data into a large speech model to extract the second feature, thereby obtaining a reference feature of the first speech data; determining the second loss value of the first speech data according to the content feature and the reference feature of the first speech data; Determine the second loss value of the second feature extraction based on the second loss value of the first speech data.
3. The method according to claim 1, characterized in that The performing spectrum reconstruction on the first speech data according to the style feature and the content feature of the first speech data includes: Inputting the style feature and the content feature of the first speech data into a decoding model to perform spectrum reconstruction, thereby obtaining reconstructed spectrum data of the first speech data; The adjusting the parameters of the first model and the parameters of the second model includes: Adjust parameters of the first model, parameters of the second model, and parameters of the decoding model.
4. The method according to claim 3, characterized in that The method for determining the spectrum reconstruction loss value of the spectrum reconstruction includes: determining the spectrum reconstruction loss value of the first speech data according to the first spectrum data and the reconstructed spectrum data of the first speech data; The spectrum reconstruction loss value of the spectrum reconstruction is determined according to the spectrum reconstruction loss value of the first speech data.
5. The method according to claim 1, wherein The method further comprises: Inputting the second speech data into the speech model to extract the second feature, thereby obtaining a reference feature of the second speech data; Inputting the style feature of the first voice data and the reference feature of the second voice data into a mapping model for feature mapping processing to obtain facial features of the first voice data; The adjusting the parameters of the first model and the parameters of the second model according to the first loss value extracted from the first feature, the second loss value extracted from the second feature, and the spectrum reconstruction loss value of the spectrum reconstruction includes: Adjust the parameters of the first model, the parameters of the second model and the parameters of the mapping model according to the first loss value extracted by the first feature, the second loss value extracted by the second feature, the spectrum reconstruction loss value of the spectrum reconstruction and the feature mapping loss value of the feature mapping processing.
6. The method according to claim 5, characterized in that The method for determining the feature mapping loss value of the feature mapping processing includes: determining a feature mapping loss value of the first speech data based on the facial features of the first speech data and the facial features of the second speech data; The feature mapping loss value of the feature mapping processing is determined according to the feature mapping loss value of the first speech data.
7. A facial feature extraction method, characterized in that: The method comprises: performing spectrum conversion processing on the third speech data to obtain second spectrum data; Inputting the second spectrum data into the first model to perform first feature extraction to obtain style features; Inputting the third speech data into the speech model to perform second feature extraction to obtain content features unrelated to the style features; The style features and the content features are input into a mapping model for feature mapping processing to obtain facial features.
8. The method according to claim 7, characterized in that After inputting the style features and the content features into a mapping model for feature mapping processing to obtain facial features, the method further includes: Inputting the facial features into a third model for parameter mapping processing to obtain facial parameters; The facial movement of the digital human is driven according to the facial parameters.
9. A computer device, characterized in that: It includes a processor and a memory, wherein the memory stores multiple instructions; the processor loads instructions from the memory to execute the steps of the model training method according to any one of claims 1 to 6, or executes the steps of the facial feature extraction method according to claim 7 or 8.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores multiple instructions, which are suitable for loading by a processor to execute the steps of the model training method as described in any one of claims 1 to 6, or to execute the steps of the facial feature extraction method as described in claim 7 or 8.
11. A computer software, characterized in that The computer software includes a computer program, and the computer program is used by a processor to execute the steps of the model training method according to any one of claims 1 to 6, or to execute the steps of the facial feature extraction method according to claim 7 or 8.