Method and device for high-fidelity generation type amplification of audio driving lip shape amplitude

By amplifying the high-frequency information of audio features and using multi-layer convolutional encoder and generator for multi-scale fusion, the problem of uncontrollable opening amplitude of digital human lip shape is solved, and the naturalness and vividness of digital human expression is improved. It is suitable for virtual reality, film and television post-production and distance education.

CN120452468APending Publication Date: 2025-08-08BEIJING YUNSIZHIXUE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510543901.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

In the existing generative digital human solution, the opening range of the lip shape is uncontrollable, resulting in the speaker's weak expression and the inability to adjust the opening range of the mouth to meet user needs.

Method used

Audio features are extracted through speech pre-trained models, high-frequency audio features are amplified, and multi-scale fusion is used to control the degree of lip amplitude generation.

Benefits of technology

It realizes controllable adjustment of the lip opening range, improves the naturalness and vividness of digital people's expression, is simple and efficient, and does not increase the computing burden. It is suitable for virtual reality, film and television post-production and distance education.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120452468A_ABST
    Figure CN120452468A_ABST
Patent Text Reader

Abstract

The invention discloses a method and device for high-fidelity generation type amplification of the amplitude of a lip shape driven by audio, and the method comprises the steps: extracting audio features in audio information through a voice pre-training model; amplifying high-frequency audio features in the audio features to obtain high-frequency amplified audio information; inputting the high-frequency amplified audio information into a high-frequency audio feature extraction model, and extracting audio features of different sizes; and respectively fusing the audio features of different sizes into a generator of the voice-driven lip synthesis model, and respectively controlling the lip shape amplitude generation degree of the generator. According to the high-fidelity generation type method for amplifying the amplitude of the audio-driven lip shape, the problem of detail loss in voice-driven lip shape synthesis is solved through high-frequency feature amplification and multi-scale fusion, and the method can be widely applied to the fields of virtual reality, later movie and television and remote education.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of digital human technology, and in particular provides a method and device for high-fidelity generative amplification of audio-driven lip amplitude. Background Art

[0002] Digital Human (or Meta Human) is a digitized humanoid created through digital technology that closely resembles a human form. It is the product of the integration of information science and life science, using information science methods to simulate the human body at different levels of form and function.

[0003] The lip shape amplitude of the generated digital human is based on the implicit features of the audio and the input reference image and mask Figure 1 Using a face image or a mouth image of a person speaking as input, one or more models directly predict a face image of a certain size or a mouth image. We found that the generated talkers often suffer from the problem of "weak speaker expressiveness," which is mainly reflected in: 1. Clarity, that is, the degree of detail displayed; 2. The degree of lip opening. The degree of lip opening is uncontrollable in currently known solutions. Although some solutions (such as Musetalk) can adjust the lip opening to a certain extent by modifying the size of the face box, this is random and uncontrollable, and has no interpretable meaning.

[0004] Therefore, in current generative digital human solutions, the lip opening range in speech videos generated from audio information is directly generated, without a module to adjust the lip opening range. However, for users, adjusting the expression of speech is meaningful. In view of this, this patent is proposed to solve the problem of adjusting the mouth opening range in generative digital human solutions. Summary of the Invention

[0005] To address the above technical issues, the present invention proposes a method and device for high-fidelity generative amplification of audio-driven lip-shaping amplitude, thereby improving the lip-shaping amplitude in generative digital human solutions, making the digital human's speech more natural and vivid. Specifically, the following technical solutions are adopted:

[0006] In a first aspect, the present invention provides a method for high-fidelity generative amplification of audio-driven lip amplitude, comprising:

[0007] Extract audio features from audio information through a speech pre-training model;

[0008] Amplifying high-frequency audio features in the audio features to obtain high-frequency amplified audio information;

[0009] Input the high-frequency amplified audio information into the high-frequency audio feature extraction model to extract audio features of different sizes;

[0010] Audio features of different sizes are fused into the generator of the speech-driven lip synthesis model to control the degree of lip amplitude generation of the generator.

[0011] As an optional embodiment of the present invention, in a high-fidelity generative amplification method for audio-driven lip shape amplitude of the present invention, amplifying high-frequency audio features in the audio features to obtain high-frequency amplified audio information includes:

[0012] Converting the audio features in the audio information into audio information features in the frequency domain;

[0013] High-frequency audio features in the audio information features in the frequency domain are screened out through high-pass filtering, and the high-frequency audio features are multiplied by an amplification coefficient k (k>1) to obtain high-frequency amplified audio information.

[0014] As an optional embodiment of the present invention, in a method of high-fidelity generative amplification of audio driving lip amplitude of the present invention, the amplification coefficients k1, k2, ..., k corresponding to the high-frequency audio characteristics of different frequency bands are preset. n ;

[0015] Filter out high-frequency audio features in the audio information features in the frequency domain through high-pass filtering;

[0016] According to the frequency band of high-frequency audio characteristics, multiply it by the corresponding amplification factor k n Get high-frequency amplified audio information.

[0017] As an optional embodiment of the present invention, in a method of high-fidelity generative amplification of audio-driven lip shape amplitude of the present invention, the inputting of high-frequency amplified audio information into a high-frequency audio feature extraction model to extract audio features of different sizes includes:

[0018] The high-frequency audio feature extraction model is a multi-layer convolutional encoder having multiple layers of convolutional feature layers;

[0019] The high-frequency amplified audio information is input into a multi-layer convolutional encoder, and audio features of different sizes corresponding to different frequency bands are extracted through the multi-layer convolutional feature layers of the multi-layer convolutional encoder.

[0020] As an optional embodiment of the present invention, in a method for high-fidelity generative amplification of audio-driven lip amplitude of the present invention, the multi-layer convolution encoder includes n layers of multi-layer convolution feature layers L1, L2, ..., L n , the size of the corresponding output audio features is L1>L2>……>Ln ;

[0021] The multi-layer convolution feature layers L1, L2, ..., Ln respectively obtain lower-size audio features through downsampling. The multi-layer convolution feature layers L1, L2, ..., Ln n Audio features of the same size in [1] are connected through convolutional layers.

[0022] As an optional embodiment of the present invention, in a high-fidelity generative amplification method of audio-driven lip shape amplitude of the present invention, audio features of different sizes are respectively integrated into the generator of the speech-driven lip synthesis model, wherein the generator of the speech-driven lip synthesis model is a multi-layer convolution generator, and the multi-layer convolution generator has multi-layer convolution feature layers, and the multi-layer convolution feature layers of the multi-layer convolution generator correspond one-to-one to the multi-layer convolution feature layers of the multi-layer convolution encoder, and the corresponding output audio features have the same size.

[0023] As an optional embodiment of the present invention, in a method of high-fidelity generative amplification of audio-driven lip shape amplitude of the present invention, the extracting audio features in the audio information through a speech pre-training model includes: extracting audio features in the audio information through a HuBERT model.

[0024] In a second aspect, the present invention provides a device for high-fidelity generative amplification of audio-driven lip amplitude, comprising:

[0025] Audio feature extraction module, which extracts audio features from audio information through a speech pre-training model;

[0026] The audio feature amplification module amplifies the high-frequency audio features in the audio features to obtain high-frequency amplified audio information;

[0027] The speech-driven lip synthesis module inputs the high-frequency amplified audio information into the high-frequency audio feature extraction model to extract audio features of different sizes; the audio features of different sizes are respectively integrated into the generator of the speech-driven lip synthesis module to control the degree of lip shape amplitude generation of the generator.

[0028] In a third aspect, the present invention provides an electronic device comprising a processor and a memory, wherein the memory is used to store a computer executable program. When the computer program is executed by the processor, the processor executes the method of high-fidelity generative amplification of audio-driven lip amplitude.

[0029] In a fourth aspect, the present invention provides a computer-readable recording medium storing a computer-executable program, wherein when the computer-executable program is executed, the method for high-fidelity generative amplification of audio-driven lip amplitude is implemented.

[0030] Compared with the prior art, the present invention has the following beneficial effects:

[0031] The present invention provides a high-fidelity generative amplification method for audio-driven lip shape amplitude, which uses a scaler to adjust the opening and closing amplitude of the mouth. The specific idea is to amplify the high-frequency information in the audio characteristics to make the corresponding mouth opening movement larger while ensuring high fidelity with the original material.

[0032] Therefore, the method of the present invention for high-fidelity generative amplification of audio-driven lip amplitude has the following technical effects:

[0033] 1. The present invention provides a high-fidelity generative amplified audio-driven lip shape amplitude method that uses a scaler to adjust the mouth's opening and closing amplitude. This method can improve the lip opening amplitude in a plug-and-play manner without affecting the original digital human generation model. It is concise, clear, and effective, with almost no increase in computational burden.

[0034] 2. The method of the present invention for high-fidelity generative amplification of audio-driven lip-shaped amplitude has better generalization and robustness: The method of the present invention for high-fidelity generative amplification of audio-driven lip-shaped amplitude can be applied to any talker scheme. As long as there is a talker scheme for audio feature extraction, the scheme in this article can be used to improve the lip opening amplitude.

[0035] 3. The present invention provides a high-fidelity generative amplification method for audio-driven lip shape amplitude. This method solves the problem of detail loss in speech-driven lip shape synthesis through high-frequency feature amplification and multi-scale fusion. It can be widely used in virtual reality, film and television post-production, and distance education. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 A flowchart of a method for high-fidelity generative amplification of audio-driven lip-shaped amplitude according to an embodiment of the present invention;

[0037] Figure 2 A model structure diagram of a multi-layer convolutional encoder and generator in a method for high-fidelity generative amplification of audio-driven lip amplitude according to an embodiment of the present invention;

[0038] Figure 3 A comparison of the effects of a high-fidelity generative amplification method for audio-driven lip amplitude according to an embodiment of the present invention with the effects of existing solutions;

[0039] Figure 4 A schematic structural diagram of an electronic device according to an embodiment of the present invention;

[0040] Figure 5 Schematic diagram of a computer-readable recording medium according to an embodiment of the present invention. DETAILED DESCRIPTION

[0041] To make the purpose, technical solutions and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them.

[0042] Therefore, the following detailed description of the embodiments of the present invention is not intended to limit the scope of the claimed invention, but merely represents some embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative work are within the scope of protection of the present invention.

[0043] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features and technical solutions therein may be combined with each other.

[0044] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings.

[0045] In the description of the present invention, it should be noted that the terms "upper" and "lower" and the like indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, or the orientations or positional relationships in which the inventive product is typically placed when in use, or the orientations or positional relationships commonly understood by those skilled in the art. Such terms are intended solely to facilitate the description of the present invention and simplify the description, and are not intended to indicate or imply that the device or element referred to must have a specific orientation, be constructed, or operate in a specific orientation. Therefore, they should not be construed as limitations on the present invention. Furthermore, the terms "first" and "second" and the like are used solely for distinction and should not be construed as indicating or implying relative importance.

[0046] See also Figure 1 As shown, this embodiment provides a method for high-fidelity generative amplification of audio-driven lip amplitude, including:

[0047] Extract audio features from audio information through a speech pre-training model;

[0048] Amplifying high-frequency audio features in the audio features to obtain high-frequency amplified audio information;

[0049] Input the high-frequency amplified audio information into the high-frequency audio feature extraction model to extract audio features of different sizes;

[0050] Audio features of different sizes are fused into the generator of the speech-driven lip synthesis model to control the degree of lip amplitude generation of the generator.

[0051] This embodiment provides a high-fidelity generative amplification method for audio-driven lip shape amplitude, using a scaler to adjust the mouth opening and closing amplitude. The specific idea is to amplify the high-frequency information in the audio features to make the corresponding mouth opening movement larger while ensuring high fidelity with the original material.

[0052] Therefore, the high-fidelity generative amplification method for audio-driven lip-shaped amplitude in this embodiment has the following technical effects:

[0053] 1. This embodiment provides a high-fidelity generative amplified audio-driven lip shape amplitude method. Using a scaler to adjust the mouth's opening and closing amplitude, this method can improve the lip opening amplitude in a plug-and-play manner without affecting the original digital human generation model. This method is concise, clear, and effective, with virtually no increase in computational burden.

[0054] 2. The high-fidelity generative amplification method for audio-driven lip-shape amplitude in this embodiment has better generalization and robustness: The high-fidelity generative amplification method for audio-driven lip-shape amplitude in this embodiment can be applied to any talker solution. As long as there is a talker solution for audio feature extraction, the solution in this article can be used to improve the lip opening amplitude.

[0055] 3. This embodiment provides a high-fidelity generative amplification method for audio-driven lip shape amplitude. By amplifying high-frequency features and integrating multi-scale features, it solves the problem of detail loss in speech-driven lip shape synthesis and can be widely used in virtual reality, film and television post-production, and distance education.

[0056] As an optional implementation manner of this embodiment, in a high-fidelity generative amplification method for audio-driven lip shape amplitude of this embodiment, amplifying high-frequency audio features in the audio features to obtain high-frequency amplified audio information includes:

[0057] Converting the audio features in the audio information into audio information features in the frequency domain;

[0058] High-frequency audio features in the audio information features in the frequency domain are screened out through high-pass filtering, and the high-frequency audio features are multiplied by the amplification coefficient k (k>1) to obtain high-frequency amplified audio information.

[0059] This embodiment converts the audio features into the frequency domain by using a method such as fast Fourier transform (FFT), uses a high-pass filter to filter out high-frequency audio features, and then multiplies by a coefficient k greater than 1 to enhance the amplitude of the high-frequency audio features in the frequency domain.

[0060] Furthermore, in the method of high-fidelity generative amplification of audio driving lip amplitude in this embodiment, amplification coefficients k1, k2, ..., k corresponding to high-frequency audio features of different frequency bands are preset. n ;

[0061] Filter out high-frequency audio features in the audio information features in the frequency domain through high-pass filtering;

[0062] According to the frequency band of high-frequency audio characteristics, multiply it by the corresponding amplification factor k n Get high-frequency amplified audio information.

[0063] Specifically, the high-frequency component is extracted by a high-pass filter and multiplied by the amplification factor k: F high =H HPF ⊙F freq , F enhanced =k·F high , where H HPF : High-pass filter; k>1: dynamic amplification factor.

[0064] This embodiment sets different amplification factors for different frequency bands (such as 4kHz-8kHz, 8kHz-16kHz):

[0065] As an optional implementation of this embodiment, in a high-fidelity generative amplification method for audio-driven lip shape amplitude of this embodiment, the inputting of high-frequency amplified audio information into a high-frequency audio feature extraction model to extract audio features of different sizes includes:

[0066] The high-frequency audio feature extraction model is a multi-layer convolutional encoder having multiple layers of convolutional feature layers;

[0067] The high-frequency amplified audio information is input into a multi-layer convolutional encoder, and audio features of different sizes corresponding to different frequency bands are extracted through the multi-layer convolutional feature layers of the multi-layer convolutional encoder.

[0068] Specifically, the multi-layer convolution encoder includes n layers of multi-layer convolution feature layers L1, L2, ..., L n , the size of the corresponding output audio feature is L1>L2>……>L n ;

[0069] Multi-layer convolutional feature layer L1, L2, ..., L n Lower-size audio features are obtained by downsampling, and multi-layer convolution feature layers L1, L2, ..., L n Audio features of the same size in [1] are connected through convolutional layers.

[0070] This embodiment uses a multi-layer convolutional encoder to extract audio features of different sizes: F1, F2, ..., F n =Encoder(F enhanced), where Fi is the feature output of the i-th layer, and its size satisfies L1>L2>…>Ln.

[0071] In this embodiment, multi-layer convolutional encoders (Encoders) are connected across layers: low-frequency information is retained through residual connections to avoid gradient disappearance.

[0072] As an optional implementation manner of this embodiment, in a high-fidelity generative amplification method of audio-driven lip shape amplitude of this embodiment, audio features of different sizes are respectively integrated into the generator of the speech-driven lip synthesis model, wherein the generator of the speech-driven lip synthesis model is a multi-layer convolution generator, and the multi-layer convolution generator has multi-layer convolution feature layers, and the multi-layer convolution feature layers of the multi-layer convolution generator correspond one-to-one to the multi-layer convolution feature layers of the multi-layer convolution encoder, and the corresponding output audio features have the same size.

[0073] Specifically, the generator adopts a symmetrical structure and is connected to the encoder layer by layer: lip =Generator(F1,F2,...,F n ), where each layer of features controls the lip amplitude in different frequency bands, such as: F1 controls the global lip opening and closing; Fn controls subtle lip jitter.

[0074] As an optional implementation of this embodiment, in a method of high-fidelity generative amplification of audio-driven lip shape amplitude in this embodiment, the extracting audio features from audio information through a speech pre-training model includes: extracting audio features from audio information through a HuBERT model.

[0075] Specifically, a pre-trained speech model (HuBERT) is used to extract deep features of speech: F audio =HuBERT(x), Where x is the input speech signal (duration T, dimension D), and Faudio is the extracted audio feature matrix.

[0076] A specific example of a high-fidelity generative amplification method for audio-driven lip-shaped amplitude in this embodiment is as follows:

[0077] 1. Extract audio features from audio information through the HuBERT model.

[0078] 2. To extract high-frequency information of audio features, you can use fast Fourier transform (FFT) and other methods to convert it into the frequency domain.

[0079] 3. Enhance the high-frequency feature amplitude in the frequency domain. Specifically, use a high-pass filter to find the high-frequency information, and then multiply it by a coefficient k greater than 1.

[0080] 4. Then input the audio features into a new high-frequency extraction model, audio-encoder (audio encoder).

[0081] a) The function of the high-frequency extraction model audio-encoder is to extract audio features of different sizes, so that low-frequency, medium-frequency, and high-frequency audio features can be extracted. Then, these audio features of different sizes are respectively fused into the subsequent generator to control the generation degree of the generator respectively.

[0082] b) In the following comparative experiment, the model of the high-frequency extraction model audio-encoder is designed into 5 different sizes as follows Figure 2 shown: 512, 384, 256, 196, 96. Each size is obtained by downsampling to a lower size, and the same-sized layers are connected by Conv convolution. Of course, there will also be corresponding features of the same size in the subsequent generator, which is convenient for direct Cross-attention calculation or CAT calculation.

[0083] As Figure 3 shown, the effect of a method for high-fidelity generative amplification of audio-driven lip amplitude in this embodiment is compared with the effect diagram of the existing scheme. The mouth shapes of the same pronunciation before and after amplifying the high frequency are compared for the pronunciation of "xiang". From left to right, they are the baseline (standard model) and the effect after audio enhancement using the method for high-fidelity generative amplification of audio-driven lip amplitude in this embodiment. Among them, the Baseline (standard model) is a wav2lip structure with a size of 256. From Figure 3 the comparison of the effect diagrams, it can be intuitively seen that after audio enhancement using the method for high-fidelity generative amplification of audio-driven lip amplitude in this embodiment, the lip opening amplitude of the digital human is larger, and the high-fidelity with the original material is ensured.

[0084] This embodiment also provides a device for high-fidelity generative amplification of audio-driven lip amplitude, including:

[0085] An audio feature extraction module that extracts audio features in audio information through a speech pre-training model;

[0086] An audio feature amplification module that amplifies the high-frequency audio features in the audio features to obtain high-frequency amplified audio information;

[0087] A speech-driven lip synthesis module that inputs the high-frequency amplified audio information into a high-frequency audio feature extraction model to extract audio features of different sizes; fuses the audio features of different sizes into the generator of the speech-driven lip synthesis module respectively to control the lip amplitude generation degree of the generator respectively.

[0088] This embodiment provides a high-fidelity generative amplification device for audio-driven lip shape amplitude. The audio feature amplification module uses a scaler to adjust the opening and closing amplitude of the mouth. The specific idea is to amplify the high-frequency information in the audio features to make the corresponding mouth opening movement larger while ensuring high fidelity with the original material.

[0089] Therefore, the device for high-fidelity generative amplification of audio-driven lip amplitude in this embodiment has the following technical effects:

[0090] 1. This embodiment provides a high-fidelity generative amplification device for audio-driven lip shape amplitude. The audio feature amplification module adopts a scaler to adjust the opening and closing amplitude of the mouth. It can improve the opening amplitude of the lip shape in a plug-and-play manner without any impact on the original digital human generation model. It is concise, clear and effective, and hardly increases the computational burden.

[0091] 2. The high-fidelity generative amplification audio-driven lip-shaped amplitude device of this embodiment has better generalization and robustness: The high-fidelity generative amplification audio-driven lip-shaped amplitude device of this embodiment can be applied to any talker solution. As long as there is a talker solution for audio feature extraction, the solution in this article can be used to improve the lip opening amplitude.

[0092] 3. This embodiment provides a high-fidelity generative amplification device for audio-driven lip shape amplitude. By amplifying high-frequency features and integrating multi-scale features, it solves the problem of detail loss in speech-driven lip shape synthesis and can be widely used in virtual reality, film and television post-production, and distance education.

[0093] As an optional implementation manner of this embodiment, in a high-fidelity generative amplification device for audio-driven lip amplitude of this embodiment, the audio feature amplification module amplifies high-frequency audio features in the audio features to obtain high-frequency amplified audio information, including:

[0094] Converting the audio features in the audio information into audio information features in the frequency domain;

[0095] High-frequency audio features in the audio information features in the frequency domain are screened out through high-pass filtering, and the high-frequency audio features are multiplied by the amplification coefficient k (k>1) to obtain high-frequency amplified audio information.

[0096] The audio feature amplification module of this embodiment converts the audio features into the frequency domain by using fast Fourier transform (FFT) and other methods, uses a high-pass filter to filter out high-frequency audio features, and then multiplies them by a coefficient k greater than 1 to enhance the amplitude of the high-frequency audio features in the frequency domain.

[0097] Furthermore, in the embodiment of the present invention, a high-fidelity generative amplification device for audio driving lip amplitude is provided, wherein the audio feature amplification module presets amplification coefficients k1, k2, ..., k corresponding to high-frequency audio features of different frequency bands. n ;

[0098] Filter out high-frequency audio features in the audio information features in the frequency domain through high-pass filtering;

[0099] According to the frequency band of high-frequency audio characteristics, multiply it by the corresponding amplification factor k n Get high-frequency amplified audio information.

[0100] Specifically, the audio feature amplification module extracts high-frequency components through a high-pass filter and multiplies them by the amplification factor k: F high =H HPF ⊙F freq , F enhanced =k·F high , where H HPF : High-pass filter; k>1: dynamic amplification factor.

[0101] The audio feature amplification module of this embodiment sets different amplification factors for different frequency bands (such as 4kHz-8kHz, 8kHz-16kHz):

[0102] As an optional implementation of this embodiment, a high-fidelity generative amplification device for audio-driven lip shape amplitude in this embodiment, the speech-driven lip synthesis module inputs high-frequency amplified audio information into a high-frequency audio feature extraction model to extract audio features of different sizes including:

[0103] The high-frequency audio feature extraction model is a multi-layer convolutional encoder having multiple layers of convolutional feature layers;

[0104] The high-frequency amplified audio information is input into a multi-layer convolutional encoder, and audio features of different sizes corresponding to different frequency bands are extracted through the multi-layer convolutional feature layers of the multi-layer convolutional encoder.

[0105] Specifically, the multi-layer convolution encoder includes n layers of multi-layer convolution feature layers L1, L2, ..., L n , the size of the corresponding output audio feature is L1>L2>……>L n ;

[0106] Multi-layer convolutional feature layer L1, L2, ..., L n Lower-size audio features are obtained by downsampling, and multi-layer convolution feature layers L1, L2, ..., L n Audio features of the same size in [1] are connected through convolutional layers.

[0107] The speech-driven lip synthesis module of this embodiment uses a multi-layer convolutional encoder to extract audio features of different sizes: F1, F2, ..., F n =Encoder(F enhanced ), where Fi is the feature output of the i-th layer, and its size satisfies L1>L2>……>L n .

[0108] In this embodiment, multi-layer convolutional encoders (Encoders) are connected across layers: low-frequency information is retained through residual connections to avoid gradient disappearance.

[0109] As an optional implementation manner of this embodiment, this embodiment provides a high-fidelity generative amplification device for audio-driven lip shape amplitude, wherein the speech-driven lip synthesis module fuses audio features of different sizes into the generator of the speech-driven lip synthesis model, wherein the generator of the speech-driven lip synthesis model is a multi-layer convolution generator, and the multi-layer convolution generator has multi-layer convolution feature layers, and the multi-layer convolution feature layers of the multi-layer convolution generator correspond one-to-one to the multi-layer convolution feature layers of the multi-layer convolution encoder, and the corresponding output audio features have the same size.

[0110] Specifically, the generator adopts a symmetrical structure and is connected to the encoder layer by layer: lip =Generator(F1,F2,...,F n ), where each layer of features controls the lip amplitude in different frequency bands, such as: F1 controls the global lip opening and closing; Fn controls subtle lip jitter.

[0111] As an optional implementation of this embodiment, this embodiment provides a device for high-fidelity generative amplification of audio-driven lip amplitude, wherein the audio feature extraction module extracts audio features from audio information through a speech pre-training model, including: extracting audio features from audio information through a HuBERT model.

[0112] Specifically, a pre-trained speech model (HuBERT) is used to extract deep features of speech: F audio =HuBERT(x), Where x is the input speech signal (duration T, dimension D), and Faudio is the extracted audio feature matrix.

[0113] The following describes an electronic device embodiment of the present invention, which can be considered a specific physical implementation of the method and apparatus embodiments of the present invention described above. Details described in the electronic device embodiment of the present invention should be considered supplementary to the above-mentioned method or apparatus embodiments; details not disclosed in the electronic device embodiment of the present invention can be implemented with reference to the above-mentioned method or apparatus embodiments.

[0114] Figure 4 This is a structural diagram of an electronic device according to an embodiment of the present invention. The electronic device includes a processor and a memory, wherein the memory is used to store a computer executable program. When the computer program is executed by the processor, the processor executes the method of high-fidelity generative amplification of audio-driven lip amplitude.

[0115] like Figure 4 As shown, the electronic device is implemented as a general-purpose computing device. The processor may be one or multiple processors working in concert. The present invention also does not exclude distributed processing, meaning that the processors may be dispersed across different physical devices. The electronic device of the present invention is not limited to a single entity but may also be the sum of multiple physical devices.

[0116] The memory stores a computer executable program, typically a machine-readable code, which can be executed by the processor to enable the electronic device to perform the method of the present invention, or at least some of the steps in the method.

[0117] The memory includes a volatile memory, such as a random access memory unit (RAM) and / or a cache memory unit, and may also be a non-volatile memory, such as a read-only memory unit (ROM).

[0118] Optionally, in this embodiment, the electronic device further includes an I / O interface for exchanging data with an external device. The I / O interface may represent one or more of several types of bus structures, including a storage unit bus or storage unit controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus structures.

[0119] It should be understood that Figure 4 The electronic device shown is merely an example of the present invention. The electronic device of the present invention may also include elements or components not shown in the above examples. For example, some electronic devices also include display units such as screens, and some electronic devices also include human-computer interaction elements such as buttons and keyboards. As long as the electronic device can execute a computer-readable program stored in its memory to implement the method of the present invention or at least some of the steps of the method, it is considered an electronic device covered by the present invention.

[0120] Figure 5 FIG is a schematic diagram of a computer readable recording medium according to an embodiment of the present invention. Figure 5 As shown, a computer-readable recording medium stores a computer-executable program, and when the computer-executable program is executed, a method for high-fidelity generative amplification of audio-driven lip amplitude according to an embodiment of the present invention is implemented. The computer-readable recording medium may include a data signal propagated in baseband or as part of a carrier wave, which carries a readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable recording medium may also be any readable medium other than a readable recording medium, which can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the readable recording medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the above.

[0121] The program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, and the like, as well as conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a stand-alone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device may be connected to the user computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0122] Through the above description of the implementation mode, it is easy for those skilled in the art to understand that the present invention can be implemented by hardware capable of executing a specific computer program, such as the system of the present invention, and the electronic processing unit, server, client, mobile phone, control unit, processor, etc. contained in the system. The present invention can also be implemented by computer software that executes the method of the present invention, such as control software executed by a microprocessor, an electronic control unit, a client, a server, etc. However, it should be noted that the computer software that executes the method of the present invention is not limited to being executed by one or a specific hardware entity, and it can also be implemented in a distributed manner by unspecified specific hardware. For computer software, the software product can be stored in a computer-readable recording medium (which can be a CD-ROM, a USB flash drive, a mobile disk, etc.), or it can be distributed and stored on a network, as long as it enables an electronic device to execute the method according to the present invention.

[0123] The above embodiments are only used to illustrate the present invention and are not intended to limit the technical solutions described in the present invention. Although this specification has described the present invention in detail with reference to the above embodiments, the present invention is not limited to the above specific implementation methods. Therefore, any modification or equivalent replacement of the present invention; and all technical solutions and improvements thereof that do not depart from the spirit and scope of the invention are included in the scope of the claims of the present invention.

Claims

1. A method for high-fidelity generative amplification of audio-driven lip amplitude, characterized in that: include: Extract audio features from audio information through a speech pre-training model; Amplifying high-frequency audio features in the audio features to obtain high-frequency amplified audio information; Input the high-frequency amplified audio information into the high-frequency audio feature extraction model to extract audio features of different sizes; Audio features of different sizes are fused into the generator of the speech-driven lip synthesis model to control the degree of lip amplitude generation of the generator.

2. The method for high-fidelity generative amplification of audio-driven lip amplitude according to claim 1, characterized in that: The step of amplifying the high-frequency audio features in the audio features to obtain high-frequency amplified audio information includes: Converting the audio features in the audio information into audio information features in the frequency domain; High-frequency audio features in the audio information features in the frequency domain are screened out through high-pass filtering, and the high-frequency audio features are multiplied by an amplification coefficient k (k>1) to obtain high-frequency amplified audio information.

3. The method for high-fidelity generative amplification of audio-driven lip amplitude according to claim 2, characterized in that: The amplification coefficients k1, k2, ..., k corresponding to the high-frequency audio characteristics of different frequency bands are preset n ; Filter out high-frequency audio features in the audio information features in the frequency domain through high-pass filtering; According to the frequency band of high-frequency audio characteristics, multiply it by the corresponding amplification factor k n Get high-frequency amplified audio information.

4. The method for high-fidelity generative amplification of audio-driven lip amplitude according to claim 1, characterized in that: The high-frequency amplified audio information is input into the high-frequency audio feature extraction model to extract audio features of different sizes, including: The high-frequency audio feature extraction model is a multi-layer convolutional encoder having multiple layers of convolutional feature layers; The high-frequency amplified audio information is input into a multi-layer convolutional encoder, and audio features of different sizes corresponding to different frequency bands are extracted through the multi-layer convolutional feature layers of the multi-layer convolutional encoder.

5. The method for high-fidelity generative amplification of audio-driven lip amplitude according to claim 4, characterized in that: The multi-layer convolution encoder includes n layers of multi-layer convolution feature layers L1, L2, ..., L n , the size of the corresponding output audio feature is L1>L2>……>L n ; Multi-layer convolutional feature layer L1, L2, ..., L n Lower-size audio features are obtained by downsampling, and multi-layer convolution feature layers L1, L2, ..., L n Audio features of the same size in [1] are connected through convolutional layers.

6. The method for high-fidelity generative amplification of audio-driven lip amplitude according to claim 5, characterized in that: The audio features of different sizes are respectively integrated into the generator of the speech-driven lip synthesis model, wherein the generator of the speech-driven lip synthesis model is a multi-layer convolution generator, and the multi-layer convolution generator has a multi-layer convolution feature layer. The multi-layer convolution feature layer of the multi-layer convolution generator corresponds one-to-one to the multi-layer convolution feature layer of the multi-layer convolution encoder, and the corresponding output audio features have the same size.

7. The method for high-fidelity generative amplification of audio-driven lip amplitude according to claim 1, characterized in that: The extracting audio features from the audio information through the speech pre-training model includes: extracting audio features from the audio information through the HuBERT model.

8. A device for high-fidelity generative amplification of audio-driven lip amplitude, characterized in that: include: Audio feature extraction module, which extracts audio features from audio information through a speech pre-training model; The audio feature amplification module amplifies the high-frequency audio features in the audio features to obtain high-frequency amplified audio information; The speech-driven lip synthesis module inputs the high-frequency amplified audio information into the high-frequency audio feature extraction model to extract audio features of different sizes; Audio features of different sizes are fused into the generator of the speech-driven lip synthesis module to control the lip amplitude generation degree of the generator.

9. An electronic device comprising a processor and a memory, wherein the memory is used to store a computer executable program, wherein: When the computer program is executed by the processor, the processor performs the method for high-fidelity generative amplification of audio-driven lip amplitude as described in any one of claims 1 to 7.

10. A computer-readable recording medium storing a computer-executable program, characterized in that: When the computer executable program is executed, the method for high-fidelity generative amplification of audio-driven lip amplitude as described in any one of claims 1 to 7 is implemented.