Speech Enhancement Method, Apparatus, Electronic Device, and Computer-Readable Storage Medium

By combining video data and audio data to extract multimodal features and using transformer to fusion, the problem of poor speech enhancement robustness caused by visual feature instability is solved, and the stable speech enhancement effect is achieved in different noise environments.

CN114333863BActive Publication Date: 2025-07-18UNIV OF SCI & TECH OF CHINA +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111544776.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-16
Publication Date
2025-07-18
Estimated Expiration
2041-12-16

AI Technical Summary

Technical Problem

The existing speech enhancement technology is poorly robust in the case of unstable visual features, and the instability of visual information in the multimodal scheme leads to unstable speech enhancement effect.

Method used

By obtaining the target's video data and original audio data, visual features, semantic features and speech features are extracted, and feature fusion is used by transformers to improve the robustness of speech enhancement.

Benefits of technology

In the case of unstable visual features, semantic features assisted enhancement improves the robustness of speech enhancement, ensuring the stability of speech enhancement effect in different noise environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114333863B_ABST
    Figure CN114333863B_ABST
Patent Text Reader

Abstract

The present application discloses a voice enhancement method, apparatus, electronic device, and computer-readable storage medium. Among them, the method includes: obtaining video data and original audio data of a target, where the video data is obtained by shooting the target when obtaining the original audio data; extracting visual features using the video data, and extracting semantic features and voice features using the original audio data; performing voice enhancement processing based on the visual features, semantic features, and voice features to obtain enhanced audio data. Through the above method, the present application can improve the robustness of voice enhancement.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech technology, and particularly to a speech enhancement method, apparatus, electronic device, and computer-readable storage medium. Background Art

[0002] The purpose of speech enhancement is to remove background noise and interfering sounds from the speaker's speech. Essentially, it is also a separation task. There have been many studies on improving speech enhancement effects by combining auxiliary information outside of speech. For example, in current multimodal enhancement schemes, visual modality is used to assist speech enhancement. However, visual information has great instability. For example, due to problems such as lighting and equipment, visual information is unstable, which may cause the visual modality not to play a role, resulting in unstable speech enhancement effects. Summary of the Invention

[0003] The main technical problem to be solved by this application is to provide a speech enhancement method, apparatus, electronic device, and computer-readable storage medium that can improve the robustness of speech enhancement.

[0004] To solve the above technical problem, in the first aspect of this application, a speech enhancement method is provided, including: obtaining video data and original audio data of a target, where the video data is obtained by shooting the target when obtaining the original audio data; extracting visual features using the video data, and extracting semantic features and speech features using the original audio data; performing speech enhancement processing based on the visual features, semantic features, and speech features to obtain enhanced audio data.

[0005] To solve the above technical problem, in the second aspect of this application, a speech enhancement apparatus is provided, including: an obtaining module for obtaining video data and original audio data of a target, where the video data is obtained by shooting the target when obtaining the original audio data; a feature extraction module for extracting visual features using the video data, and extracting semantic features and speech features using the original audio data; a speech enhancement module for performing speech enhancement processing based on the visual features, semantic features, and speech features to obtain enhanced audio data.

[0006] To solve the above technical problem, in the third aspect of this application, an electronic device is provided, including a memory and a processor coupled to each other, where the memory is used to store program data, and the processor is used to execute the program data to implement the foregoing method.

[0007] To solve the above technical problem, in the fourth aspect of this application, a computer-readable storage medium is provided, in which program data is stored, and when the program data is executed by a processor, it is used to implement the foregoing method.

[0008] The beneficial effects of the present application are as follows: Different from the prior art, the present application obtains the video data and the original audio data of the target, where the video data is obtained by shooting the target when obtaining the original audio data. Then, visual features are extracted using the video data, and semantic features and speech features are extracted using the original audio data. Finally, speech enhancement processing is performed based on the visual features, semantic features, and speech features to obtain enhanced audio data. Among them, new semantic features are introduced, and multi-modal speech enhancement processing is performed by integrating visual features, semantic features, and speech features. In the case where the visual features are unstable, the semantic features can be used for auxiliary enhancement, which is beneficial to improving the robustness of speech enhancement. Description of the Drawings

[0009] In order to more clearly illustrate the technical solutions in the present application, the following will briefly introduce the drawings required in the description of the embodiments. Obviously, the following described drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. Among them:

[0010] Figure 1 is a flowchart of an embodiment of the speech enhancement method of the present application;

[0011] Figure 2 is another flowchart of an embodiment of the speech enhancement method of the present application;

[0012] Figure 3 is a flowchart of another embodiment of step S12 of the present application;

[0013] Figure 4 is a flowchart of yet another embodiment of step S12 of the present application;

[0014] Figure 5 is Figure 1 a flowchart of another embodiment of step S13 in

[0015] Figure 6 is Figure 1 another flowchart of another embodiment of step S13 in

[0016] Figure 7 is Figure 5 a flowchart of another embodiment of step S133 in

[0017] Figure 8 is a flowchart of another embodiment of the speech enhancement method of the present application;

[0018] Figure 9 is a structural schematic diagram of an embodiment of the speech enhancement device of the present application;

[0019] Figure 10It is a structural schematic block diagram of an embodiment of the electronic device of the present application;

[0020] Figure 11 It is a structural schematic block diagram of an embodiment of the computer-readable storage medium of the present application. Detailed implementation manners

[0021] Referring to "embodiment" in the present application means that the specific features, structures or characteristics described in combination with the embodiment can be included in at least one embodiment of the present application. The appearance of this phrase at various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0022] The terms "first" and "second" in the present application are only used for descriptive purposes, and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present application, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically defined. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices.

[0023] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0024] Currently, speech enhancement solutions can be divided into two types: single-modal solutions and multi-modal solutions. Among them, the single-modal solution is based on speech features, and the enhanced speech is obtained by learning the clean speech mask, or the voiceprint features are extracted to improve the speech enhancement effect of a specific speaker. The effect is relatively stable for a fixed speaker, but it can only be customized for the target person, and the generalization ability is weak, reducing the usability of the model. The multi-modal solution mainly adds visual features, but the visual elements are not as fine-grained as phonemes, and there is a certain instability in the visual features, and the model is not stable enough, resulting in poor robustness of speech enhancement. In addition, the feature fusion method adopted does not exploit the extraction ability of the visual modality for audio signals, and even has a counterproductive effect in a low-noise environment.

[0025] Based on this, the present application provides a voice enhancement method, which extracts multi-modal features by using video data and original audio data simultaneously to assist voice enhancement. Since audio and video can play good auxiliary effects in low-noise and high-noise scenarios respectively, fusing the two can improve the robustness in a noisy environment. Secondly, lip key-point information can enhance the extraction of visual features, and the introduction of semantic features improves the robustness of voice enhancement through information such as vocalization methods and semantic content. In addition, by effectively fusing features of different modalities through a transformer, it can be ensured that the system still has good voice enhancement results when the noise level is different and the visual features are unstable.

[0026] Please refer to Figures 1 to 2 , Figure 1 which is a schematic flowchart of an embodiment of the voice enhancement method of the present application, Figure 2 and which is another schematic flowchart of an embodiment of the voice enhancement method of the present application.

[0027] The method may include the following steps:

[0028] Step S11: Obtain video data and original audio data of a target, where the video data is obtained by shooting the target when obtaining the original audio data.

[0029] The target is any entity that can emit sound, such as a person, an anthropomorphic robot, etc. The target may include one or more entities. In one example, the original audio data may be the sound emitted when a person (one or more) is speaking.

[0030] The video data is a combination of image frames. Among them, the video data may include audio data or may not include audio data. In a related technology, only one face image is used to extract visual features. Although the computational complexity of the model is reduced, it is only effective for in-set data and the usability of the model is poor. In this embodiment, however, motion visual features are extracted through video data, and compared with extracting visual features only using one face image, the usability of the model is stronger.

[0031] In an application scenario, a camera may be used to shoot the target to obtain video data of the target, and at the same time, an audio collection device (such as a microphone) is used to collect the audio data of the target to obtain the original audio data. Among them, due to the influence of the collection environment and the like, the original audio data may be noisy audio data, and the voice enhancement method provided in this embodiment can remove the noise in the original audio data to obtain enhanced audio data. The noise may be, but is not limited to, non-target sounds such as noise, reverberation, and breathing sounds doped therein.

[0032] In some embodiments, the video data and the original audio data may also be preprocessed before step S12 to ensure the validity of the video data and the original audio.

[0033] In one example, the preprocessing may be to remove the image frames in the video data that do not contain the region of interest based on the facial key point information, where the region of interest includes the lip shape of the target. It can be understood that the video data is used to assist in speech enhancement, and the movement of the lip shape is closely related to pronunciation. Therefore, the result of speech enhancement can be improved by combining the movement information of the lip shape. Thus, the effective image frames in the video data need to contain the lip shape of the target. Specifically, the facial key point algorithm can be used to detect the facial key points of the image frames in the video data to obtain the facial key point information, and then based on the facial key point information, the image frames that do not contain the region of interest can be screened out, and thus this part of the image frames can be removed. In other cases, the region of interest may also be other parts related to pronunciation, not limited to the lip shape of the target. For example, when the target is an anthropomorphic robot, when the target pronounces, the eyes of the target (specifically, the eye region displayed on the display screen) will open, and when the target does not pronounce, the eyes of the target will close. At this time, the region where the eyes of the target are located can be used as the region of interest.

[0034] In another example, the preprocessing may be to remove the asynchronous data in the audio data and the video data to obtain synchronous audio data and video data. Since the present application is a multi-modal speech enhancement method, both audio data and video data are required, and good synchronization between the two is required. In some application scenarios, there may be a time delay or asymmetry in the acquisition of the audio data and the video data, resulting in the obtained audio data and video data being asynchronous.

[0035] In yet another example, the preprocessing may be to remove the image frames in the video data that do not contain the region of interest based on the facial key point information, and then remove the asynchronous data in the audio data and the video data. Among them, since some of the image frames in the video data are removed, resulting in the audio data and the video data being asynchronous, the audio data corresponding to this part of the image frames can be removed. In addition, the frame rate of the video data can be unified, for example, unified to 25 fps (Frames Per Second).

[0036] Step S12: Extract visual features using the video data, and extract semantic features and speech features using the original audio data.

[0037] In some embodiments, the visual features include the motion information of the lip shape of the target, such as opening and closing information. Semantics refers to the meaning of data, and the semantic features include the meaning expressed by the target in the original audio data. The speech features include the features of the sound emitted by the target in the original audio data, such as timbre, pitch, etc.

[0038] Step S13: Perform speech enhancement processing based on the visual features, semantic features, and speech features to obtain enhanced audio data.

[0039] As Figure 2 shown, specifically, the visual features, semantic features, and speech features can be fused to obtain enhanced features, and then based on the enhanced features, enhanced audio data can be obtained. In some embodiments, the enhanced features can be directly output as the enhanced audio data. In other embodiments, the enhanced features can be further processed to obtain the enhanced audio data. The specific content can be referred to the following embodiments respectively.

[0040] Due to problems such as illumination and equipment, the visual information is unstable, resulting in unstable speech enhancement effect. However, speech enhancement is essentially to extract the audio data of the target. Therefore, semantic features and speech features can be extracted from the audio data itself for its own speech enhancement. In this embodiment, the visual features, semantic features, and speech features are comprehensively used for multi-modal speech enhancement processing. Even when the visual features are unstable, high-quality speech can still be extracted according to the semantic features and speech features, which is beneficial to improving the robustness of speech enhancement.

[0041] In this embodiment, by obtaining the video data and the original audio data of the target, where the video data is obtained by shooting the target when obtaining the original audio data, then using the video data to extract the visual features, and using the original audio data to extract the semantic features and speech features, and finally performing speech enhancement processing based on the visual features, semantic features, and speech features to obtain the enhanced audio data. Among them, new semantic features are introduced. By comprehensively using the visual features, semantic features, and speech features for multi-modal speech enhancement processing, when the visual features are unstable, the semantic features can be used for auxiliary enhancement, which is beneficial to improving the robustness of speech enhancement.

[0042] Next, the acquisition methods of the visual features, speech features, and semantic features will be introduced respectively in combination with Figures 3 to 4 ...

[0043] Please refer to Figure 3 ... Figure 3 which is a schematic flowchart of another embodiment of step S12 of the present application.

[0044] In this embodiment, using the video data to extract the visual features may include steps S121 to S123.

[0045] Step S121: Use the video data to intercept the image of the target's attention area, where the attention area includes the lip shape of the target.

[0046] As Figure 2 shown, in some embodiments, the attention area image is an image centered on the lip shape (or lips). The types of attention area images may include, but are not limited to: RGB (color) images, grayscale images, and infrared images.

[0047] Step S122: Generate a key point mask image according to the key points of the lip shape.

[0048] Specifically, the attention area image can be processed to obtain the key points of the lip shape, so that a key point mask image (mask image) can be generated according to the key points of the lip shape for extracting the lip shape motion features, that is, visual features.

[0049] Step S123: Extract visual features based on the key point mask image and the attention area image.

[0050] In some embodiments, when extracting visual features, the sizes of the attention area image and the key point mask image need to be the same, for example, both are 64x64 pixels. In other embodiments, the sizes of the attention area image and the key point mask image may not be limited. In some other embodiments, even if the sizes of the attention area image and the key point mask image are different or do not meet the network input size, the image size can be adjusted as needed.

[0051] In some embodiments, a three-dimensional convolutional residual network can be used to perform a preset prediction on the key point mask image and the attention area image to obtain visual features. Specifically, a 64x64 color attention area image centered on the lip shape and the corresponding key point mask image can be merged into a 4-channel sequence as the network input. In this embodiment, the visual features have a certain temporal sequence, and each frame image has a certain connection with the adjacent front and rear images. Therefore, by using a neural network and adopting a 3D convolutional and residual network structure (such as Conv3D+resnet18), the feature changes between adjacent frames can be better integrated, and more effective features can be extracted.

[0052] Please refer to Figure 4 , Figure 4 which is a schematic flowchart of another embodiment of step S12 of the present application.

[0053] In this embodiment, extracting semantic features and speech features from the audio data may include steps S124 to S126. It should be noted that there is no certain sequence relationship between the above steps S121 to S123 and steps S124 to S126, and they can be executed or executed successively.

[0054] Step S124: Extract frequency-domain features from the audio data, and perform short-time Fourier transform on the audio data to obtain the mixed speech power spectrum.

[0055] On the one hand, a preset number of dimensions (e.g., 40 dimensions) of frequency-domain features can be extracted from the audio data as the input of the semantic extraction network for semantic feature extraction. The preset number can be set according to the needs of the semantic extraction network.

[0056] On the other hand, after performing short-time Fourier transform (Short-Time Fourier Transform, STFT) on the audio data, the mixed speech power spectrum and the mixed speech phase can be obtained. Among them, the mixed speech power spectrum is used for speech feature extraction. The mixed speech phase is used to predict the target speech phase subsequently. For specific details, please refer to the following embodiments.

[0057] Step S125: Process the frequency-domain features using the semantic extraction network to obtain semantic features.

[0058] In the semantic features, the distinguishable components in the audio data need to be extracted to remove the influence of noise, etc. In this embodiment, the pre-trained semantic extraction network is used to process the frequency-domain features (e.g., Fbank features) to obtain semantic features. Among them, the speech features output by the semantic extraction network are at the frame level.

[0059] Step S126: Process the mixed speech power spectrum using the enhancement network to obtain speech features.

[0060] Since speech enhancement needs to restore the target speech, in this embodiment, the mixed speech power spectrum is used as the initial speech feature and input into the pre-trained enhancement network to perform enhancement processing on the initial speech feature through the enhancement network, thereby obtaining speech features.

[0061] In some embodiments, the frame lengths and step sizes of the speech features and the semantic features are the same. For example, the frame length is 25 ms and the step size is 10 ms; the visual features, for example, have a frame length of 40 ms. In order to align with the speech features, the visual features and the speech features can be aligned by means of bilinear interpolation. Thus, it is convenient to perform feature fusion on the visual features, semantic features, and speech features subsequently.

[0062] In this embodiment, the semantic extraction network and the enhancement network can be obtained by training a deep neural network (Deep Neural Networks, DNN). In other examples, other types of neural networks can also be used.

[0063] Please refer to Figures 5 to 6 , Figure 5 isFigure 1 Flow diagram of another embodiment of step S13 Figure 6 It is Figure 1 Another process of another embodiment of step S13

[0064] In this embodiment, step S13 may include sub-steps S131 to S133

[0065] Step S131: Merge visual features and semantic features to obtain auxiliary features

[0066] Since the visual features have obvious enhancement effects when the noise is large, and the semantic features can keep the clean speech undistorted when the noise is small. In order to effectively integrate the advantages of the two features, first, the two features are merged in the time dimension to obtain auxiliary features (denoted as F).

[0067] Step S132: Fuse the auxiliary features and speech features to obtain enhanced features

[0068] Since this application involves a multi-modal speech enhancement scheme, the selection of the feature fusion method between different modalities also affects the final speech enhancement effect. In this regard, this embodiment also provides a feature fusion method applicable to this method, that is, using the attention mechanism for fusion, which can

[0069] After obtaining the auxiliary features, the attention mechanism can be used to fuse the auxiliary features and speech features to obtain enhanced features. Specifically, as Figure 6 shown, the auxiliary features and speech features can be input into the transformer model to fuse the auxiliary features and speech features based on the multi-head attention mechanism (MHA). Among them, the corresponding features can be extracted from the speech features according to the feature key extracted from the auxiliary features, and then the enhanced features, that is, the features of the clean speech, can be obtained after three layers of MHA

[0070] Step S133: Based on the enhanced features, obtain enhanced audio data

[0071] Since the enhanced features obtained after feature fusion have removed noise, in this embodiment, the enhanced features can be directly used as the enhanced audio data

[0072] However, in some other embodiments, in order to strengthen the connection between adjacent frames and avoid speech distortion problems, further post-processing can also be performed on the enhanced features

[0073] Please refer to Figure 7 , Figure 7 It is Figure 5 Flow diagram of another embodiment of step S133

[0074] In this embodiment, step S133 may further include sub-steps S1331 to S1333.

[0075] Step S1331: Based on the enhanced features, the target speech power spectrum is extracted.

[0076] Specifically, a Long Short-Term Memory (LSTM) neural network can be used to process the enhanced features to obtain the target speech power spectrum. The Long Short-Term Memory neural network can be, but is not limited to, any one of a Bi-directional Long Short-Term Memory (BiLSTM) neural network and a unidirectional Long Short-Term Memory (unidirectional LSTM) neural network. Among them, BiLSTM is composed of a forward LSTM and a backward LSTM combined.

[0077] In some embodiments, a BiLSTM can be used to process the enhanced features to obtain the target speech power spectrum. The inventors of this application found in experiments that the speech enhancement result of BiLSTM is significantly improved compared with that of the unidirectional LSTM.

[0078] Step S1332: A preset operation is performed on the target speech power spectrum and the mixed speech power spectrum to obtain the target speech logarithmic power spectrum, where the mixed speech power spectrum is obtained by performing a short-time Fourier transform on the original audio data.

[0079] Optionally, the preset operation can be a dot product (inner product, scalar product of vectors) or other mathematical operation methods that can play a similar role.

[0080] Step S1333: Based on the target speech phase, a preset transform is performed on the target speech logarithmic power spectrum to obtain the enhanced audio data.

[0081] Optionally, the preset transform is an Inverse Short-Time Fourier Transform (iSTFT), but is not limited thereto.

[0082] Since the audio data is divided into a power spectrum and a phase part after a short-time Fourier transform, and among them, the phase also affects the speech enhancement effect. In some embodiments, the mixed speech phase obtained by performing a short-time Fourier transform on the original audio data can be used as the target speech phase, and an inverse short-time Fourier transform is performed in combination with the target speech logarithmic power spectrum.

[0083] In some other embodiments, it is also possible to predict the phase corresponding to the target speech power spectrum. For example, before step S1333, it may further include: performing preset processing on the mixed speech phase and the target speech logarithmic power spectrum using a phase prediction network to obtain the target speech phase, where the mixed speech phase is obtained by performing short-time Fourier transform on the original audio data. Specifically, the mixed speech phase and the target speech logarithmic power spectrum are stacked in the channel dimension and then used as the input of the phase prediction network, so that the deviation between the target speech phase and the mixed speech phase can be obtained. Further, adding this deviation to the mixed speech phase can obtain the target speech phase. Thus, according to this target speech phase, the speech enhancement effect can be further improved.

[0084] In some embodiments, in order to better predict the target speech phase, a phase-related loss function may also be added during the training process of the phase prediction network. Specifically, the phase prediction network can be trained using the phase loss function, and then the weight value corresponding to the target point in the target speech logarithmic power spectrum in the phase loss function is increased, where the value of the target point in the target speech logarithmic power spectrum is greater than a preset threshold. The preset threshold can be set according to the actual situation. Among them, by increasing the phase weight of the points with larger values in the power spectrum, the phase prediction of important points can be made more accurate.

[0085] In some embodiments, the formula of the phase loss function is as follows:

[0086]

[0087] In the above formula (1), L phase is the phase loss value, is the target speech logarithmic power spectrum, is the target speech phase, φ tf is the predicted speech phase. Among them, the power spectrum can be regarded as a picture with dimensions of T*F, and tf represents the power spectrum index. A corresponding phase weight can be set for each sampling point in the target speech logarithmic power spectrum. By adjusting the phase weight of the sampling points with larger values corresponding to tf, the phase prediction of important points can be made more accurate.

[0088] If the loss value does not meet the training stop condition of the model, a new sample is selected to continue training the model. If the loss value meets the training stop condition of the model, the currently trained network can be used as the phase prediction network for corresponding service scenarios.

[0089] Please refer to Figure 8 , Figure 8 which is a schematic flowchart of another embodiment of the speech enhancement method of the present application.

[0090] (1) Obtain the video data and the original audio data of the target, where the video data is obtained by shooting the target when obtaining the original audio data.

[0091] (2) Extract visual features using the video data, and extract semantic features and speech features using the original audio data.

[0092] Among them, extracting visual features includes: processing the video data to obtain a key point mask image and a region of interest image, and then inputting the key point mask image and the region of interest image into a three-dimensional convolutional residual network (including Conv3D + resnet18) to obtain visual features.

[0093] Among them, extracting semantic features includes: extracting frequency domain features using the audio data, and then inputting them into a semantic extraction network to output semantic features.

[0094] Among them, extracting speech features includes: performing a short-time Fourier transform on the audio data to obtain a mixed speech power spectrum and a mixed speech phase, and then inputting the mixed speech power spectrum into an enhancement network to obtain speech features.

[0095] (3) Perform speech enhancement processing based on the visual features, semantic features, and speech features to obtain enhanced audio data.

[0096] Among them, first merge the visual features and semantic features in the time dimension to obtain auxiliary features, and then input the auxiliary features and speech features into a feature fusion module to be fused using a transformer to obtain enhanced features. In some embodiments, the enhanced features have already removed noise and can be output as enhanced audio data.

[0097] In other embodiments, post-processing can also be performed on the enhanced features to strengthen the connection between adjacent frames and avoid speech distortion problems. This includes: inputting the enhanced features into a BiLSTM to obtain a target speech power spectrum (not shown in the figure), and then performing a dot product calculation on the target speech power spectrum and the mixed speech power spectrum using a Mask (mask) to obtain a target speech logarithmic power spectrum. Further, the target speech logarithmic power spectrum and the mixed speech phase can be input into a phase prediction network to obtain a target speech phase, and finally, the target speech logarithmic power spectrum and the target speech phase are subjected to an inverse short-time Fourier transform to obtain enhanced audio data.

[0098] On this basis, this embodiment further provides a training method. First, the visual extraction network (3D convolutional residual network) and the semantic extraction network can be pre-trained. Then, during the overall training process, by fixing the parameters of the semantic feature extraction network, due to the differences between modalities, the parameters of the visual extraction network can be fixed first. After the entire model converges, the parameters of the visual extraction network are uniformly optimized and updated, that is, the parameters of the visual extraction network are updated after the model converges, and the effect will be more stable. It should be noted that the speech extraction network (enhancement network) here does not need pre-training and is trained during the training of the entire model. Finally, by calculating the loss value (such as L2-loss) between the enhanced speech data and the corresponding clean audio data, the parameters of the entire model are updated according to the loss value.

[0099] Please refer to Figure 9 , Figure 9 which is a schematic structural diagram of an embodiment of the speech enhancement device of the present application.

[0100] In this embodiment, the speech enhancement device 100 may include an acquisition module 110, a feature extraction module 120, and a speech enhancement module 130. Among them, the acquisition module 110 is used to acquire the video data and the original audio data of the target, where the video data is obtained by shooting the target when acquiring the original audio data; the feature extraction module 120 is used to extract visual features using the video data, and extract semantic features and speech features using the original audio data; the speech enhancement module 130 is used to perform speech enhancement processing based on the visual features, semantic features, and speech features to obtain enhanced audio data.

[0101] In some embodiments, the speech enhancement module 130 is further used to merge the visual features and the semantic features to obtain auxiliary features; fuse the auxiliary features and the speech features to obtain enhanced features; and obtain enhanced audio data based on the enhanced features.

[0102] In some embodiments, the speech enhancement module 130 is further used to fuse the auxiliary features and the speech features using the attention mechanism to obtain enhanced features, and based on the enhanced features, extract the target speech power spectrum; perform a preset operation on the target speech power spectrum and the mixed speech power spectrum to obtain the target speech logarithmic power spectrum, where the mixed speech power spectrum is obtained by performing short-time Fourier transform on the original audio data; and perform a preset transformation on the target speech logarithmic power spectrum based on the target speech phase to obtain enhanced audio data.

[0103] In some embodiments, the preset operation is dot multiplication, and / or, the speech enhancement module 130 is further used to process the enhanced features using a long short-term memory neural network to obtain the target speech power spectrum.

[0104] In some embodiments, the preset transformation is an inverse short-time Fourier transform, and / or the speech enhancement module 130 is further configured to perform a preset process on the mixed speech phase and the target speech log power spectrum by using a phase prediction network to obtain the target speech phase, where the mixed speech phase is obtained by performing a short-time Fourier transform on the original audio data.

[0105] In some embodiments, when training the phase prediction network, it further includes training the phase prediction network by using a phase loss function; increasing the weight value corresponding to the target point in the phase loss function in the target speech log power spectrum, where the value of the target point in the target speech log power spectrum is greater than a preset threshold.

[0106] In some embodiments, the feature extraction module 120 is further configured to intercept an image of the target's attention area by using video data, where the attention area includes the lip shape of the target; generate a key point mask image according to the key points of the lip shape; and extract visual features based on the key point mask image and the attention area image.

[0107] In some embodiments, the feature extraction module 120 is further configured to perform a preset prediction on the key point mask image and the attention area image by using a three-dimensional convolutional residual network to obtain visual features.

[0108] In some embodiments, the feature extraction module 120 is further configured to extract frequency domain features by using audio data, and perform a short-time Fourier transform on the audio data to obtain a mixed speech power spectrum; process the frequency domain features by using a semantic extraction network to obtain semantic features; and process the mixed speech power spectrum by using an enhancement network to obtain speech features.

[0109] In some embodiments, before extracting visual features by using video data, and extracting semantic features and speech features by using the original audio data, the acquisition module 110 is further configured to remove the image frames in the video data that do not include the attention area based on the face key point information, where the attention area includes the lip shape of the target; and / or remove the asynchronous data in the audio data and the video data.

[0110] In this embodiment, the speech enhancement device is used to implement the speech enhancement method in the above embodiment. Therefore, the description of the above steps can be referred to the method embodiment correspondingly, and will not be repeated here.

[0111] Please refer to Figure 10 , Figure 10 which is a schematic structural diagram of an embodiment of the electronic device of the present application.

[0112] In this embodiment, the electronic device 200 may include a memory 210 and a processor 220 that are coupled to each other. The memory 210 is used to store program data, and the processor 220 is used to execute the program data to implement the steps in any of the above method embodiments. Specifically, the electronic device 200 may include, but is not limited to: desktop computers, laptop computers, servers, mobile phones, tablet computers, etc., which are not limited herein.

[0113] Specifically, the processor 220 is used to control itself and the memory 210 to implement the steps in any of the above method embodiments. The processor 220 may also be referred to as a CPU (Central Processing Unit). The processor 220 may be an integrated circuit chip with signal processing capabilities. The processor 220 may also be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Additionally, the processor 220 may be implemented jointly by multiple integrated circuit chips.

[0114] Please refer to Figure 11 , Figure 11 which is a structural schematic diagram of an embodiment of the computer-readable storage medium of the present application.

[0115] In this embodiment, the computer-readable storage medium 300 stores program data 310. When the program data 310 is executed by the processor, it is used to implement the steps in any of the above method embodiments.

[0116] The computer-readable storage medium 300 may be a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc., which can store computer programs, or it may be a server storing the computer program. The server may send the stored computer program to other devices for running, or it may also run the stored computer program itself.

[0117] In several embodiments provided in the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.

[0118] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0119] In addition, each functional unit in various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0120] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0121] The above are only the embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.

Claims

1. A voice enhancement method, characterized in that, Including: Obtain the video data and the original audio data of the target, where the video data is obtained by shooting the target when obtaining the original audio data; Extract visual features using the video data, and extract semantic features and speech features using the original audio data; Perform speech enhancement processing based on the visual features, the semantic features, and the speech features to obtain enhanced audio data.

2. The method according to claim 1, wherein The performing speech enhancement processing based on the visual features, the semantic features, and the speech features to obtain enhanced audio data includes: Merge the visual features and the semantic features to obtain auxiliary features; Fuse the auxiliary features and the speech features to obtain enhanced features; Based on the enhanced features, obtain the enhanced audio data.

3. The method according to claim 2, wherein The fusing the auxiliary features and the speech features to obtain enhanced features includes: Use the attention mechanism to fuse the auxiliary features and the speech features to obtain the enhanced features; The obtaining the enhanced audio data based on the enhanced features includes: Extract the target speech power spectrum based on the enhanced features; Perform a preset operation on the target speech power spectrum and the mixed speech power spectrum to obtain the target speech logarithmic power spectrum, where the mixed speech power spectrum is obtained by performing short-time Fourier transform on the original audio data, and the preset operation is dot multiplication; Based on the target speech phase, perform a preset transform on the target speech logarithmic power spectrum to obtain the enhanced audio data, where the preset transform is inverse short-time Fourier transform.

4. The method according to claim 3, characterized in that, The extracting the target speech power spectrum based on the enhanced features includes: Use a long short-term memory neural network to process the enhanced features to obtain the target speech power spectrum.

5. The method according to claim 3, wherein Before performing the preset transform on the target speech logarithmic power spectrum based on the target speech phase to obtain the enhanced audio data, it further includes: Use a phase prediction network to perform a preset process on the mixed speech phase and the target speech logarithmic power spectrum to obtain the target speech phase, where the mixed speech phase is obtained by performing short-time Fourier transform on the original audio data.

6. The method according to claim 5, wherein The method further includes: Train the phase prediction network using a phase loss function; Increase the weight value corresponding to the target point in the target speech logarithmic power spectrum in the phase loss function, where the value of the target point in the target speech logarithmic power spectrum is greater than a preset threshold.

7. The method according to claim 1, wherein The extracting visual features using the video data includes: Use the video data to intercept the attention area image of the target, and the attention area includes the lip shape of the target; Generate a key point mask image according to the key points of the lip shape; Based on the key point mask image and the attention area image, extract the visual features.

8. The method according to claim 7, wherein The extracting the visual features based on the key point mask image and the attention area image includes: Performing a preset prediction on the key point mask image and the region of interest image by using a three-dimensional convolutional residual network to obtain the visual feature.

9. The method according to claim 1, wherein: The extracting the semantic feature and the speech feature by using the original audio data includes: Extracting frequency domain features by using the original audio data, and performing a short-time Fourier transform on the original audio data to obtain a mixed speech power spectrum; Processing the frequency domain features by using the semantic extraction network to obtain the semantic feature; Processing the mixed speech power spectrum by using the enhancement network to obtain the speech feature.

10. The method according to claim 1, wherein: Before the extracting the visual feature by using the video data, and the extracting the semantic feature and the speech feature by using the original audio data, further includes: Based on the face key point information, removing the image frames in the video data that do not include the region of interest, and the region of interest includes the lip shape of the target; and / or Removing the asynchronous data in the original audio data and the video data.

11. A voice enhancement device, characterized in that, Includes: An acquisition module, configured to acquire video data and original audio data of a target, wherein the video data is obtained by shooting the target when acquiring the original audio data; A feature extraction module, configured to extract a visual feature by using the video data, and extract a semantic feature and a speech feature by using the original audio data; A speech enhancement module, configured to perform speech enhancement processing based on the visual feature, the semantic feature and the speech feature to obtain enhanced audio data.

12. An electronic device, characterized in that, The electronic device includes a memory and a processor coupled to each other, the memory is configured to store program data, and the processor is configured to execute the program data to implement the method according to any one of claims 1-10.

13. A computer-readable storage medium, characterized in that, Program data is stored in the computer-readable storage medium, and when the program data is executed by a processor, it is used to implement the method according to any one of claims 1-10.

Citation Information

Patent Citations

  • Audio and video speech enhancement processing method and model

    CN112951258A

  • Audio-visual speech enhancement method and system capable of fully utilizing vision and speech connection

    CN113470671A