Audio and video speech recognition method and system based on audio enhancement
Through the audio enhancement method driven by visual feature, the noise reduction mask is generated using visual context information, which solves the problem of insufficient robustness of the audio speech recognition system in noisy environments, and achieves more efficient noise suppression and recognition performance improvement.
Patent Information
- Application Number
- CN202510363857.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-08
AI Technical Summary
The prior art fails to fully utilize visual context information to enhance audio features, resulting in insufficient robustness of audio speech recognition systems in noisy environments.
Through visual feature extraction, audio feature extraction, visual context-driven audio enhancement and audio video feature fusion, a cross-modal attention mechanism is used to generate a noise reduction mask, suppress audio noise and enhance audio signals.
It improves the robustness and recognition performance of the audio and video voice recognition system in noisy environments, effectively suppressing noise interference.
Smart Images

Figure CN120279925A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of audio - visual speech recognition, and specifically to an audio - visual speech recognition method and system based on audio enhancement. Background Art
[0002] Audio - Visual Speech Recognition (AVSR) is a technology that combines audio and video information for speech recognition. In a noisy environment, the audio signal may be severely interfered with, resulting in a significant decline in the performance of traditional Audio Speech Recognition (ASR) systems. However, visual information (such as lip movement) is not affected by audio noise. Therefore, the robustness of speech recognition can be improved through the fusion of audio and video.
[0003] In the prior art, although there have been some methods attempting to use visual information to assist audio enhancement or speech recognition, these methods often have the following deficiencies: they fail to fully utilize visual context information to enhance audio features. Summary of the Invention
[0004] The technical task of the present invention is to address the above - mentioned deficiencies and provide an audio - visual speech recognition method and system based on audio enhancement, which can enhance the audio signal in a noisy environment and apply it to an end - to - end audio - visual speech recognition system to improve the robustness and recognition performance of the system.
[0005] The technical solution adopted by the present invention to solve its technical problems is as follows:
[0006] An audio - visual speech recognition method based on audio enhancement realizes audio - visual speech recognition based on vision - context - driven audio enhancement. The implementation of this method includes the following steps:
[0007] 1) Visual feature extraction: Extract visual features from the input lip video.
[0008] 2) Audio feature extraction: Extract audio features from the input audio signal.
[0009] 3) Vision - context - driven audio enhancement: Through a cross - modal attention mechanism, use visual features to generate visual context information related to audio features; generate a noise reduction mask according to the visual context information and apply it to the audio features to enhance the audio signal and suppress noise.
[0010] 4) Audio - visual feature fusion: Fuse the enhanced audio features with visual features to generate joint features for speech recognition.
[0011] By introducing visual context information, this method can effectively suppress noise in the audio signal and improve the robustness of the audio-visual speech recognition system in a noisy environment.
[0012] Furthermore, the visual feature extraction includes local and global lip movement information.
[0013] Furthermore, extracting audio features from the input audio signal includes a log Mel spectrogram.
[0014] Furthermore, the visual feature extraction specifically includes:
[0015] Preprocess the input lip video, including cropping, scaling, and grayscaling;
[0016] Use a 3D convolutional layer and a ResNet-18 network to extract visual feature f v , with the output feature dimension being T×D1, where T is the number of video frames and D1 is the feature dimension.
[0017] Furthermore, the audio feature extraction specifically includes:
[0018] Preprocess the input audio signal, including framing, windowing, and Mel spectrogram calculation;
[0019] Use a 2D convolutional layer and a ResBlock network to extract audio feature f a , with the output feature dimension being T×D1, where T is the number of audio frames, which is consistent with the frame rate of the visual feature.
[0020] Furthermore, the visual context-driven audio enhancement specifically includes:
[0021] 3.1) Calculate the visual context information:
[0022]
[0023] where C t represents the visual context information calculated at time step t, which is used to enhance the audio feature at the current moment and has the shape of a D2-dimensional vector;
[0024] Q t is obtained by multiplying the current audio feature at time step t with the learnable weight matrix W q and has the shape of a D2-dimensional vector:
[0025]
[0026] K is composed of the visual features f v at all time steps and the learnable weight matrix W kIt is obtained by multiplication to form a matrix of shape T×D2:
[0027] K = f v ·W k ;
[0028] V is obtained by multiplying all the visual features f at each time step v by the learnable weight matrix W v to form a matrix of shape T×D2:
[0029] V = f v ·W v ;
[0030] 3.2) Generate a noise reduction mask:
[0031] m = Sigmoid(Conv(ReLU(Conv(C; θ1)); θ2));
[0032] where Conv represents a one-dimensional convolution operation, and θ1 and θ2 are the weight parameters of the convolutional layer;
[0033] 3.3) Apply the noise reduction mask to enhance the audio features:
[0034]
[0035] where f a is the original audio feature, is the enhanced audio feature, and ⊙ represents element-wise multiplication, with a shape of T×D1.
[0036] Furthermore, the audio-visual feature fusion:
[0037] Concatenate the enhanced audio feature and the visual feature:
[0038]
[0039] where ∥ represents feature concatenation, and W f and b f are the weight and bias parameters of the fusion module.
[0040] The present invention also claims a audio-visual speech recognition system based on audio enhancement, including:
[0041] A visual feature extraction module for extracting visual features from the input lip video;
[0042] An audio feature extraction module: for extracting audio features from the input audio signal;
[0043] Vision context-driven audio enhancement module: Through a cross-modal attention mechanism, visual features are used to generate visual context information related to audio features; a noise reduction mask is generated based on the visual context information and applied to the audio features to enhance the audio signal and suppress noise;
[0044] Audio-visual feature fusion module: Fuse the enhanced audio features with visual features to generate joint features for speech recognition;
[0045] This system can implement the above-mentioned audio-visual speech recognition method based on audio enhancement.
[0046] The present invention also claims to protect an audio-visual speech recognition device based on audio enhancement, including: at least one memory and at least one processor;
[0047] The at least one memory is used to store machine-readable programs;
[0048] The at least one processor is used to call the machine-readable program to implement the above method.
[0049] The present invention also claims to protect a computer-readable medium, on which computer instructions are stored, and when the computer instructions are executed by a processor, the above method is implemented.
[0050] Compared with the prior art, the audio-visual speech recognition method and system based on audio enhancement of the present invention have the following beneficial effects:
[0051] By introducing visual context information, the present invention can effectively suppress noise in the audio signal, enhance the audio signal in a noisy environment, and apply it to an end-to-end audio-visual speech recognition system, improving the robustness and recognition ability of the audio-visual speech recognition system in a noisy environment. Description of the Drawings
[0052] Figure 1 is a flowchart of the audio-visual speech recognition method based on audio enhancement provided by an embodiment of the present invention. Detailed Embodiments
[0053] The following further illustrates the present invention with specific embodiments.
[0054] The technical solution adopted by the embodiment of the present invention to solve its technical problems is:
[0055] An audio-visual speech recognition method based on audio enhancement realizes audio-visual speech recognition based on vision context-driven audio enhancement. The implementation of this method includes the following steps:
[0056] 1. Visual feature extraction: Extract visual features from the input lip video, including local and global lip movement information.
[0057] 2. Audio feature extraction: Extract audio features from the input audio signal, such as log Mel spectrogram.
[0058] 3. Visual context-driven audio enhancement:
[0059] Generate visual context information related to audio features using visual features through a cross-modal attention mechanism;
[0060] Generate a noise reduction mask based on the visual context information and apply it to the audio features to enhance the audio signal and suppress noise.
[0061] 4. Audio-visual feature fusion: Fuse the enhanced audio features with visual features to generate joint features for speech recognition.
[0062] The specific implementation method is as follows:
[0063] The above-mentioned visual feature extraction:
[0064] Preprocess the input lip video, including cropping, scaling, and grayscaling.
[0065] Use a 3D convolutional layer and a ResNet-18 network to extract visual feature f v , and the output feature dimension is T×D1, where T is the number of video frames and D1 is the feature dimension.
[0066] The above-mentioned audio feature extraction:
[0067] Preprocess the input audio signal, including framing, windowing, and Mel spectrogram calculation.
[0068] Use a 2D convolutional layer and a ResBlock network to extract audio feature f a , and the output feature dimension is T×D1, where T is the number of audio frames, which is consistent with the frame rate of the visual features.
[0069] The above-mentioned visual context-driven audio enhancement specifically includes:
[0070] 3.1. Calculate visual context information:
[0071]
[0072] Among them, C t represents the visual context information calculated at time step t, which is used to enhance the audio features at the current moment and has the shape of a D2-dimensional vector;
[0073] Q tObtained by multiplying the current audio features at time step t by the learnable weight matrix W q to get a vector of shape D2 dimensions:
[0074]
[0075] K is obtained by multiplying the visual features f at all time steps v by the learnable weight matrix W k to get a matrix of shape T×D2:
[0076] K = f v ·W k ;
[0077] V is obtained by multiplying the visual features f at all time steps v by the learnable weight matrix W v to get a matrix of shape T×D2:
[0078] V = f v ·W v .
[0079] 3.2. Generate a noise reduction mask:
[0080] m = Sigmoid(Conv(ReLU(Conv(C; θ1)); θ2));
[0081] where Conv represents a one-dimensional convolution operation, and θ1 and θ2 are the weight parameters of the convolutional layer.
[0082] 3.3. Apply the noise reduction mask to enhance the audio features:
[0083]
[0084] where f a is the original audio feature, is the enhanced audio feature, ⊙ represents element-wise multiplication, and the shape is T×D1.
[0085] The audio-visual feature fusion:
[0086] Concatenate the enhanced audio feature and the visual feature:
[0087]
[0088] where ∥ represents feature concatenation, W f and b f are the weight and bias parameters of the fusion module.
[0089] By introducing visual context information, this method can effectively suppress noise in audio signals and improve the robustness of the audio-visual speech recognition system in noisy environments. This method can be applied to intelligent speech interaction in noisy scenarios.
[0090] An embodiment of the present invention also provides an audio-visual speech recognition system based on audio enhancement, including:
[0091] 1. A visual feature extraction module, configured to extract visual features from the input lip video, including local and global lip movement information.
[0092] 2. An audio feature extraction module, configured to extract audio features from the input audio signal, such as log Mel spectrogram.
[0093] 3. A visual context-driven audio enhancement module:
[0094] Through a cross-modal attention mechanism, visual context information related to the audio features is generated using the visual features.
[0095] A noise reduction mask is generated according to the visual context information and applied to the audio features to enhance the audio signal and suppress noise.
[0096] 4. An audio-visual feature fusion module: The enhanced audio features and visual features are fused to generate joint features for speech recognition.
[0097] This system can implement the audio-visual speech recognition method based on audio enhancement described in the above embodiment. The implementation method is as follows:
[0098] The visual feature extraction module:
[0099] Preprocess the input lip video, including cropping, scaling, and grayscaling.
[0100] Use a 3D convolutional layer and a ResNet-18 network to extract visual feature f v , and the output feature dimension is T×D1, where T is the number of video frames and D1 is the feature dimension.
[0101] The audio feature extraction module:
[0102] Preprocess the input audio signal, including framing, windowing, and Mel spectrogram calculation.
[0103] Use a 2D convolutional layer and a ResBlock network to extract audio feature f a , and the output feature dimension is T×D1, where T is the number of audio frames, which is consistent with the frame rate of the visual features.
[0104] The visual context-driven audio enhancement module specifically includes:
[0105] 3.1. Calculate visual context information:
[0106]
[0107] Among them, C t represents the visual context information calculated at time step t, which is used to enhance the audio features at the current moment and has the shape of a D2-dimensional vector;
[0108] Q t is obtained by multiplying the current audio features at time step t with the learnable weight matrix W q and has the shape of a D2-dimensional vector:
[0109]
[0110] K is obtained by multiplying the visual features f v at all time steps with the learnable weight matrix W k and has the shape of a T×D2 matrix:
[0111] K = f v ·W k ;
[0112] V is obtained by multiplying the visual features f v at all time steps with the learnable weight matrix W v and has the shape of a T×D2 matrix:
[0113] V = f v ·W v .
[0114] 3.2. Generate a noise reduction mask:
[0115] m = Sigmoid(Conv(ReLU(Conv(C; θ1)); θ2));
[0116] Among them, Conv represents a one-dimensional convolution operation, and θ1 and θ2 are the weight parameters of the convolutional layer.
[0117] 3.3. Apply the noise reduction mask to enhance the audio features:
[0118]
[0119] Among them, f a is the original audio feature, is the enhanced audio feature, ⊙ represents element-wise multiplication, and the shape is T×D1.
[0120] The audio-visual feature fusion module:
[0121] Concatenate the enhanced audio features and visual features:
[0122]
[0123] Among them, ∥ represents feature concatenation, and W f and b f are the weight and bias parameters of the fusion module.
[0124] An embodiment of the present invention also provides an audio-visual speech recognition device based on audio enhancement, including: at least one memory and at least one processor;
[0125] The at least one memory is used to store machine-readable programs;
[0126] The at least one processor is used to call the machine-readable program to implement the audio-visual speech recognition method based on audio enhancement described in the above embodiments.
[0127] An embodiment of the present invention also provides a computer-readable medium, on which computer instructions are stored. When the computer instructions are executed by a processor, the audio-visual speech recognition method based on audio enhancement described in the above embodiments is implemented. Specifically, a system or device equipped with a storage medium can be provided, on which software program codes for implementing the functions of any one of the above embodiments are stored, and the computer (or CPU or MPU) of the system or device is made to read and execute the program codes stored in the storage medium.
[0128] In this case, the program code read from the storage medium itself can implement the functions of any one of the above embodiments. Therefore, the program code and the storage medium storing the program code constitute a part of the present invention.
[0129] Embodiments of the storage medium for providing program codes include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), magnetic tapes, non-volatile memory cards, and ROMs. Optionally, the program code can be downloaded from a server computer through a communication network.
[0130] In addition, it should be clear that not only can the functions of any one of the above embodiments be implemented by executing the program code read by the computer, but also some or all of the actual operations can be completed by the operating system etc. operating on the computer based on the instructions of the program code, so as to implement the functions of any one of the above embodiments.
[0131] In addition, it can be understood that the program code read from the storage medium is written into the memory provided in the expansion board inserted into the computer or into the memory provided in the expansion unit connected to the computer, and then based on the instructions of the program code, the CPU or the like installed on the expansion board or the expansion unit is made to execute part or all of the actual operations, thereby realizing the functions of any one of the above embodiments.
[0132] The present invention has been described in detail above with reference to the accompanying drawings and preferred embodiments. However, the present invention is not limited to these disclosed embodiments. Based on the above-mentioned multiple embodiments, those skilled in the art can know that more embodiments of the present invention can be obtained by combining the code review means in the above different embodiments, and these embodiments are also within the protection scope of the present invention.
Claims
1. An audio - video speech recognition method based on audio enhancement, characterized in that, Implementing audio-visual speech recognition based on vision context-driven audio enhancement, the implementation of this method includes the following steps: 1) Visual feature extraction: Extract visual features from the input lip video; 2) Audio feature extraction: Extract audio features from the input audio signal; 3) Vision context-driven audio enhancement: Through a cross-modal attention mechanism, use visual features to generate visual context information related to audio features; Generate a noise reduction mask according to the visual context information and apply it to the audio features to enhance the audio signal and suppress noise; 4) Audio-visual feature fusion: Fuse the enhanced audio features with visual features to generate joint features for speech recognition.
2. The audio - video speech recognition method based on audio enhancement according to claim 1, wherein The visual feature extraction includes local and global lip movement information.
3. A method for audio - video speech recognition based on audio enhancement according to claim 1, wherein, The extraction of audio features from the input audio signal includes a log Mel spectrogram.
4. A method for audio - video speech recognition based on audio enhancement according to claim 1, wherein, The visual feature extraction specifically includes: Preprocess the input lip video, including cropping, scaling, and grayscaling; Use a 3D convolutional layer and a ResNet-18 network to extract visual features, and the output feature dimension is T×D1, where T is the number of video frames and D1 is the feature dimension.
5. A method for audio - video speech recognition based on audio enhancement according to claim 1, characterized in that, The audio feature extraction specifically includes: Preprocess the input audio signal, including framing, windowing, and Mel spectrogram calculation; Use a 2D convolutional layer and a ResBlock network to extract audio features, and the output feature dimension is T×D1, where T is the number of audio frames, which is consistent with the frame rate of the visual features.
6. A method for audio-visual speech recognition based on audio enhancement according to claim 1 or 4 or 5, characterized in that, The vision context-driven audio enhancement specifically includes: 3.1) Calculate visual context information: Among them, C t represents the visual context information calculated at time step t, which is used to enhance the audio features at the current moment and has the shape of a D2-dimensional vector; Q t It is obtained by multiplying the current audio features by a learnable weight matrix at time step t and has the shape of a D2-dimensional vector; K is obtained by multiplying the visual features of all time steps with a learnable weight matrix, and its shape is a T×D2 matrix; V is obtained by multiplying the visual features of all time steps with a learnable weight matrix, and its shape is a T×D2 matrix; 3.2) Generate a noise reduction mask: m = Sigmoid(Conv(ReLU(Conv(C; θ1)); θ2)); where Conv represents a one-dimensional convolution operation, and θ1 and θ2 are the weight parameters of the convolutional layer; 3.3) Apply the noise reduction mask to enhance audio features: where f a is the original audio feature, is the enhanced audio feature, and ⊙ represents element-wise multiplication with a shape of T×D1.
7. An audio-visual speech recognition method based on audio enhancement according to claim 6, characterized in that, The audio-visual feature fusion: Concatenate the enhanced audio features with visual features: Among them, ∥ represents feature concatenation, and W f and b f are the weight and bias parameters of the fusion module.
8. An audio-visual speech recognition system based on audio enhancement, characterized in that, Includes: A visual feature extraction module for extracting visual features from the input lip video; An audio feature extraction module: for extracting audio features from the input audio signal; A vision context-driven audio enhancement module: Through a cross-modal attention mechanism, use visual features to generate visual context information related to audio features; Generate a noise reduction mask according to the visual context information and apply it to the audio features to enhance the audio signal and suppress noise; An audio-visual feature fusion module: Fuse the enhanced audio features with visual features to generate joint features for speech recognition; This system implements the audio-visual speech recognition method based on audio enhancement described in any one of claims 1 to 7.
9. An audio - video speech recognition device based on audio enhancement, characterized in that, Includes: At least one memory and at least one processor; The at least one memory is used to store machine-readable programs; The at least one processor is configured to call the machine-readable program to implement the method according to any one of claims 1 to 7.
10. A computer-readable medium, characterized in that, The computer-readable medium stores computer instructions which, when executed by a processor, implement the method according to any one of claims 1 to 7.
Citation Information
Cited By
Novel optical fiber microphone sensor and audio transmission method thereof
CN121262503A