Microphone voice recognition system and method based on multi-mode audio-visual fusion

Through multimodal audio-visual fusion technology, combined with the three modes of audio, vision and spectrum, 3D CNN and bidirectional GRU are used to extract features, solving the problem of degradation in complex noise environments, and achieving a high robustness and low error rate speech recognition system.

CN120340463AActive Publication Date: 2025-07-18ZHONGSHAN FENGXU ELECTRONIC IND CO LTD

Patent Information

Application Number
CN202510521530.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-07-18
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

The performance of existing speech recognition technologies in complex noise environments has significantly decreased, resulting in a decrease in recognition accuracy and limiting their use range in practical applications.

Method used

The multimodal audio-visual fusion method is adopted, combining audio, vision and spectrum three-modal fusion, features are extracted through 3D CNN, dense spatiotemporal CNN and bidirectional GRU, and decoded using CTC decoder and Beam Search algorithm to improve the robustness and accuracy of the model.

Benefits of technology

It significantly improves the accuracy and robustness of speech recognition, can operate reliably in various complex noise environments, reduces error rates, and improves the practicality of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340463A_ABST
    Figure CN120340463A_ABST
Patent Text Reader

Abstract

The invention discloses a microphone voice recognition system and method based on multi-mode audio-visual fusion, and belongs to the technical field of artificial intelligence and voice interaction. Firstly, an audio module collects voice signals through a microphone, voice is converted into texts by means of a cloud voice recognition API, and words are further mapped into 300-dimensional semantic vectors through Word2Vec. A visual module extracts lip movement and log-Mel frequency spectrum features, after being subjected to Dlib detection and normalization processing, lip images are sent into a 3D CNN and a dense space-time CNN to extract space-time features, key areas are highlighted with the assistance of a space attention mechanism, and finally sequence visual features are extracted through bidirectional GRU. Meanwhile, a log-Mel spectrogram is generated from the audio signal, and the perception characteristic is enhanced through Mel filtering and logarithm processing. The audio word vector, the lip movement feature and the log-Mel feature are spliced into a multi-modal fusion vector, the multi-modal fusion vector is sent to a CTC decoder, and a text is predicted through Beam Search decoding. An Adam optimizer and a small-batch training strategy are used in the training process, and the model performance and generalization ability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of artificial intelligence and speech interaction, and particularly relates to a microphone speech recognition system and method based on multimodal audio-visual fusion. Background Art

[0002] With the rapid development of artificial intelligence and speech interaction technologies, automatic speech recognition (ASR) systems have been widely applied in fields such as smart homes, in-vehicle navigation, medical transcription, and robot interaction. However, the degradation of speech recognition performance in noisy environments has become a key bottleneck restricting its practical applications. Traditional speech recognition technologies mainly rely on a single audio modality, and their performance is significantly affected by factors such as background noise, reverberation, and speaker differences, resulting in a substantial reduction in recognition accuracy in complex scenarios (such as streets, factories, and public places). The speech recognition system with degraded performance can only be used in limited specific environments before being applied to the real-world environment, leading to low-quality services and a reduction in consumer expectations. Summary of the Invention

[0003] In order to solve the problem that the existing speech technologies in the background art significantly degrade in performance in complex noisy environments, the present invention aims to provide a microphone speech recognition system and method based on multimodal audio-visual fusion. Starting from the actual factors affecting speech recognition, a new multimodal audio-visual fusion method with strong robustness is proposed. Through innovative multi-feature fusion and network optimization, an efficient solution is provided for noise-robust speech recognition, promoting the practical application process of multimodal technologies in real scenarios. The technical solution of the present invention can not only significantly improve the accuracy of speech recognition, but also reliably run repeatedly in various application scenarios. Different from traditional solutions, through the fusion of three modalities: audio + vision + spectrum, combined with a deep learning optimized architecture, a speech recognition system with high robustness and low error rate is realized.

[0004] In order to solve the technical problem, the technical solution of the present invention is as follows:

[0005] A microphone speech recognition method based on multimodal audio-visual fusion, the method comprising:

[0006] Step S1: Real-time collect a speech signal through a microphone, and send it to a cloud speech recognition API to be converted into a text word sequence; then, use a pre-trained Word2Vec model to map each word into a 300-dimensional word vector. If the word recognition fails, fill it with a zero vector. The word vectors of multiple words are compressed through average pooling or a fully connected layer, and finally a one-dimensional audio semantic feature vector is generated;

[0007] Step S2: First, use Dlib to locate the face and lip regions, perform key-point alignment and normalization on the lips to generate a lip movement image sequence. The image sequence is input into a 3D CNN and a dense spatio-temporal CNN to extract lip movement visual features, and a spatial attention mechanism is used to focus on the key lip regions. Finally, a bidirectional GRU is used to extract temporal features. At the same time, the speech signal is converted into a log-Mel spectrogram to enhance the perceptual characteristics of the audio and generate log-Mel spectral features.

[0008] Step S3: Concatenate the audio semantic feature vector, lip movement visual features, and log-Mel spectral features into a fused feature vector, send it to a CTC decoder for decoding, and combine with the Beam Search algorithm to output the final text. During the training process, the Adam optimizer and the mini-batch training strategy are adopted to improve the accuracy and generalization ability of the model.

[0009] Furthermore, the step S1 includes:

[0010] S101: Input the audio signal, collect the speech signal in real time through the local microphone device, and transmit the speech data to the cloud speech recognition API for recognition and processing.

[0011] S102: Generate a word list. The API converts the speech into a text word sequence and outputs an identifiable word list.

[0012] S103: Word embedding vectorization. Use the Word2Vec model pre-trained on the Google News corpus. Words with similar contexts are close in the vector space. Map each word in the word list output by the API to a 300-dimensional vector. If a word is not recognized by the API, fill it with a zero vector. If the word list contains multiple words, compress them into a one-dimensional feature vector through average pooling or a fully connected layer as the final output of the audio module.

[0013] Furthermore, the step S2 includes:

[0014] S201: Lip movement image preprocessing. Use a Dlib linear classifier to locate the face, extract the face and lip regions to generate a lip sequence image. Align the lip region according to the key points to eliminate the influence of head pose changes. Normalize each frame of the image in terms of channels to reduce the interference of illumination changes and standardize the image data. During training, horizontally flip the sequence images to improve the generalization ability of the model.

[0015] S202: Generate log-Mel spectrogram. Frame the 3-second voice signal with a 25-ms window and a 10-ms step size, resulting in 750 frames of spectrogram. Filter each frame of spectrogram through a Mel triangular filter to convert linear frequency to Mel scale, enhancing the perceptual characteristics of the voice. Take the logarithm of the Mel spectrogram energy to obtain the log-Mel spectrogram, enhancing the distinction between high-frequency and low-frequency features.

[0016] S203: Construct a visual feature extraction network, including:

[0017] 3D CNN module. 3D CNN adds three-dimensional convolutional kernel parameters on the basis of 2D CNN, enabling the feature maps in consecutive frames to be associated with consecutive frames in the previous layer and integrated into a single frame, ultimately achieving the extraction of motion information. Process the lip movement sequence images, extract spatio-temporal features through a 3D convolutional kernel, and combine batch normalization, ReLU activation, and a 3D max pooling layer to output the initial feature map.

[0018]

[0019] Among them, is the weight parameter of the three-dimensional convolutional kernel at position (p, q, r);

[0020] Among them, is the input feature of the previous layer;

[0021] Dense spatio-temporal CNN module: Adopt a dense short connection structure to alleviate gradient disappearance through short path connections, reduce the number of parameters, and improve the training efficiency. Specifically include:

[0022] Dense block: Each block contains a 6-layer BN→ReLU→3D convolution structure, and the feature maps between layers are passed through splicing to enhance feature reuse;

[0023] Transition block: Contains BN→ReLU→3D convolution→3D average pooling, compressing the number of channels and reducing the size of the feature map to reduce the amount of calculation;

[0024] Spatial attention module: Aggregate channel information through average pooling and max pooling layers to generate a spatial attention weight map. The original feature map is multiplied element-wise with the attention map to highlight the key lip regions. The calculation formula is as follows:

[0025] M S (F) = σ(f 7×7 ([AvgPool(F); MaxPool(F)])):

[0026] Among them, σ is the Sigmoid function, and f 7×7 is a 7×7 convolution operation;

[0027] Bidirectional GRU Module: Processes temporal features, controls the information flow through update gates and reset gates, captures long-term temporal dependencies, and the concatenated hidden states of the output form a unified visual temporal feature vector, which serves as the final output of the visual module for subsequent multi-modal feature fusion.

[0028] Furthermore, step S3 includes:

[0029] Concatenate the audio semantic feature vector, lip motion visual feature, and log-Mel spectrogram feature to form a multi-modal fusion feature matrix, providing comprehensive information for subsequent decoding;

[0030] Use the connectionist temporal classification loss function to address the issue of inconsistent input and output sequence lengths and maximize the probability of the correct label sequence; perform decoding in combination with the beam search algorithm to generate the final text prediction result; during training, use the Adam optimizer with a learning rate set to 0.0001 and a mini-batch size of 8 to improve the stability of the model and prevent overfitting.

[0031] A microphone speech recognition system based on multi-modal audio-visual fusion, the system is applied to any of the above methods, and the system includes:

[0032] Audio Feature Extraction and Word Vector Generation Module: Real-time collect speech signals through a microphone and send them to the cloud speech recognition API to convert them into a text word sequence; then, use the pre-trained Word2Vec model to map each word to a 300-dimensional word vector. If the word recognition fails, fill it with a zero vector. The word vectors of multiple words are compressed through average pooling or a fully connected layer, and finally, a one-dimensional audio semantic feature vector is generated;

[0033] Visual Feature Extraction and log-Mel Spectrogram Generation Module: First, use Dlib to locate the face and lip regions, perform key point alignment and normalization processing on the lips to generate a lip motion image sequence; the image sequence is input into a 3D CNN and a dense spatio-temporal CNN to extract lip motion visual features, and a spatial attention mechanism is used to focus on the key lip regions. Finally, temporal features are extracted through a bidirectional GRU; at the same time, the speech signal is converted into a log-Mel spectrogram to enhance the perceptual characteristics of the audio and generate log-Mel spectrogram features;

[0034] Multi-modal Feature Fusion and Decoding Module: Concatenate the audio semantic feature vector, lip motion visual feature, and log-Mel spectrogram feature into a fusion feature vector, send it to the CTC decoder for decoding, and combine the Beam Search algorithm to output the final text; during the training process, use the Adam optimizer and the mini-batch training strategy to improve the accuracy and generalization ability of the model.

[0035] A computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements a microphone speech recognition method based on multi-modal audio-visual fusion as described in any one of the above.

[0036] A computer-readable storage medium stores a computer program, which implements a microphone speech recognition method based on multi-modal audio-visual fusion as described in any one of the above when executed by a processor.

[0037] Compared with the prior art, the advantages of the present invention are as follows:

[0038] Improve speech recognition accuracy: Complementarity is achieved through tri-modal fusion (audio word vectors, lip movements, log-Mel spectrograms) to make up for the deficiencies of a single modality and significantly improve speech recognition accuracy.

[0039] Strong information complementarity: By introducing the log-Mel spectrogram as the third modality, the audio signal is converted into a visualized spectrogram. Complementarity is achieved based on tri-modal fusion (audio word vectors, lip movements, log-Mel spectrograms) to enhance noise robustness, and dense spatio-temporal 3D CNN is used to efficiently extract spatio-temporal features.

[0040] Improve computational efficiency and real-time performance: Use spatial attention mechanism and bidirectional GRU to optimize feature extraction and temporal modeling, reduce redundant calculations, improve inference speed, and avoid overfitting at the same time.

[0041] Therefore, the object of this invention patent is to propose an innovative multi-modal audio-visual fusion method aiming at the technical deficiencies in the existing solutions to comprehensively improve the robustness and recognition accuracy of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 The technical roadmap of a microphone speech recognition method based on multi-modal audio-visual fusion of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0043] The following describes the specific embodiments of the present invention in conjunction with the embodiments:

[0044] It should be noted that the structures, ratios, sizes, etc. shown in this specification are only used to cooperate with the content disclosed in the specification for those skilled in the art to understand and read, and are not used to limit the implementation conditions of the present invention. Any modification of the structure, change of the proportional relationship, or adjustment of the size should still fall within the scope covered by the technical content disclosed in the present invention without affecting the effects that the present invention can produce and the purposes that can be achieved.

[0045] Meanwhile, terms such as "upper", "lower", "left", "right", "middle", and "one" cited in this specification are only for the convenience of clear narration and do not limit the scope of implementation of the present invention. The change or adjustment of their relative relationship shall also be regarded as the scope of implementation of the present invention without substantial change in technical content.

[0046] Embodiment 1:

[0047] This embodiment provides a microphone speech recognition method based on multimodal audio-visual fusion, aiming to solve the problem that the performance of traditional speech recognition systems significantly degrades in complex noise environments. Specifically, in traditional noise environments, the performance of speech recognition is unstable, and the audio signal is severely interfered by background noise, resulting in a significant reduction in recognition accuracy and limiting its practical application scope. The limitations of single modality: pure audio recognition relies on acoustic signals and is easily affected by noise, and cannot maintain stability under low signal-to-noise ratio (SNR) conditions; pure vision is not affected by noise, but has insufficient ability to distinguish homophones and is restricted by factors such as illumination and posture changes. The deficiencies of multimodal fusion: existing multimodal speech recognition technologies usually only combine audio and lip visual information, resulting in insufficient information complementarity and limited performance improvement in noise environments.

[0048] Therefore, this embodiment provides a microphone speech recognition system and method based on robust multimodal audio-visual fusion, which can effectively solve the problem of performance degradation of speech recognition systems in complex noise environments. This method is based on multimodal fusion and combines a deep learning optimization architecture to solve problems such as noise interference, homophone ambiguity, and computational efficiency, significantly improving the recognition stability in high-noise scenarios such as traffic and public places, and realizing a speech recognition system with high robustness and low error rate. As Figure 1 is the main technical solution process of the present invention. It can be seen from the figure that from the perspective of multimodal audio-visual fusion, this technical solution is mainly divided into the following three steps:

[0049] Step 1: Audio feature extraction and word vector generation

[0050] (1) Input audio signal: Real-time collect speech signals through a local microphone device and transmit the speech data to the cloud speech recognition API (Application Programming Interface) for recognition and processing.

[0051] (2) Generate a word list: The API converts the speech into a text word sequence and outputs an identifiable word list.

[0052] (3) Word embedding vectorization: Use the Word2Vec model pre-trained on the Google News corpus. Words in similar contexts are close in the vector space. Map each word in the word list output by the API to a 300-dimensional vector. If the API fails to recognize a word, fill it with a zero vector. If the word list contains multiple words, compress them into a one-dimensional feature vector through average pooling or a fully connected layer as the final output of the audio module.

[0053] Step 2: Visual feature extraction (lip movement and log-Mel spectrogram)

[0054] (1) Lip movement image preprocessing: Use the Dlib linear classifier to locate the face, extract the face and lip regions, and generate a lip sequence image (resolution 640×480, 30fps); align the lip region according to the key points to eliminate the influence of head pose changes, perform channel normalization on each frame of the image, set the mean to 0 and the variance to 1 to reduce the interference of light changes, standardize the image data, and perform horizontal flipping on the sequence image during training to improve the generalization ability of the model.

[0055] (2) log-Mel (logarithmic Mel) spectrogram generation: Divide a 3-second speech signal (sampling rate 16kHz, 48000 sampling points) into frames with a 25ms window and a 10ms step size, generating a total of 750 frames of spectrograms; filter each frame of the spectrogram through a Mel triangular filter to convert the linear frequency to the Mel scale and enhance the perceptual characteristics of the speech; take the logarithm of the Mel spectrogram energy to obtain the log-Mel spectrogram (size 40×750) to enhance the distinction between high-frequency and low-frequency features.

[0056] (3) Visual feature extraction network:

[0057] 3D CNN (3D convolutional neural network) module: Different from the 2D convolutional neural network, 3D CNN can effectively extract multi-dimensional features in the lip reading recognition task (such as the movement features of the lips, tongue, and teeth). It achieves this function by encoding motion information in multiple consecutive frames of images. Specifically, the formula is as follows. 3DCNN adds three-dimensional convolutional kernel parameters on the basis of 2D CNN, enabling the feature maps in consecutive frames to be associated with the consecutive frames of the previous layer and integrated into a single frame, ultimately realizing the extraction of motion information. Process the lip movement sequence image, extract spatio-temporal features through a 3D convolutional kernel (size 7×7×3), and combine batch normalization (BN), ReLU (rectified linear unit) activation, and a 3D max pooling layer to output the initial feature map.

[0058]

[0059] Among them, is the weight parameter of the three-dimensional convolutional kernel at position (p,q,r);

[0060] Among them, is the input feature of the previous layer.

[0061] Dense spatio-temporal CNN module: Adopting a dense short connection structure, it alleviates the vanishing gradient through short path connections, reduces the number of parameters, and improves the training efficiency. Specifically, it includes:

[0062] Dense block: Each block contains a 6-layer BN→ReLU→3D convolution structure (convolution kernel 3×3×3), and the feature maps between layers are passed through splicing to enhance feature reuse.

[0063] Transition block: It contains BN→ReLU→3D convolution→3D average pooling, compresses the number of channels and reduces the size of the feature map, and reduces the computational amount.

[0064] Spatial attention module: Aggregates channel information through average pooling and max pooling layers to generate a spatial attention weight map. The original feature map is multiplied element-wise with the attention map to highlight the key regions of the lips. The calculation formula is as follows:

[0065] M S (F) = σ(f 7×7 ([AvgPool(F); MaxPool(F)])):

[0066] Among them, σ is the Sigmoid function, and f 7×7 is a 7×7 convolution operation.

[0067] Bidirectional GRU (Gated Recurrent Unit) module: Processes temporal features, controls the information flow through update gates and reset gates, captures long temporal dependencies, concatenates the hidden states of the forward and backward GRUs, and outputs a visual feature vector.

[0068] Step 3: Multimodal feature fusion and decoding

[0069] (1) Feature splicing: Splice the audio word vector (300-dimensional), the lip movement feature vector (from the dense spatio-temporal CNN), and the log-Mel spectrogram feature vector into a fusion matrix.

[0070] (2) CTC decoding and training:

[0071] Use the Connectionist Temporal Classification (CTC) loss function to solve the problem of inconsistent input and output sequence lengths, maximize the probability of the correct label sequence, and decode in combination with the Beam Search algorithm to generate the final text prediction result. During training, the Adam (Adaptive Moment Estimation) optimizer (learning rate 0.0001) is used, and the mini-batch size is 8 to prevent overfitting.

[0072] It can be understood that:

[0073] Multimodal Audio-Visual Fusion Method: The present invention proposes an innovative multimodal audio-visual fusion method. This solution fuses three modalities: audio + vision + spectrum. The spectrum converts the speech signal into log-Mel frequency domain features, enhancing the representation of high-frequency and low-frequency details and compensating for the lack of frequency domain information in the pure visual modality. It aims to solve problems such as the limitations of a single modality, noise interference, and homophone ambiguity, and realizes a speech recognition system with high robustness and low error rate, thus forming a solution suitable for speech recognition systems in complex noise environments.

[0074] Efficient Network Architecture Design: The present invention adopts a dense spatio-temporal CNN and a spatial attention mechanism, significantly reducing parameter redundancy and computational complexity through short connections and bottleneck layers. Additionally, it combines a bidirectional gated recurrent unit (Bi-GRU) to process temporal data and uses a connectionist temporal classification (CTC) loss to solve the input-output alignment problem, improving the encoding accuracy.

[0075] The three-modal fusion method is the core innovation of the present invention. It introduces the log-Mel spectrogram as the third modality, converting the audio signal into a visual spectrogram, improving the information complementarity in complex noise environments, and thus enhancing the accuracy of the speech recognition system. This key point has important technical value and application prospects and is the most critical and unique innovative content of this invention patent.

[0076] In an alternative implementation, multimodal Transformer fusion: One alternative is to use a Transformer to encode audio (Log-Mel spectrum), lip movement video (spatio-temporal features), and text (ASR output) respectively, dynamically aligning the temporal relationship between audio and visual features through a cross-attention mechanism. For example, when there is noise interference, it enhances the weight of lip movement features and uses a large amount of unlabeled data for multimodal pre-training to improve the generalization of feature representation. However, the computational complexity is relatively high, and the model needs to be optimized for lightweight to adapt to edge devices.

[0077] In an alternative implementation, dynamic gating multimodal fusion: Another alternative is to design dynamic weight gates for each modality, adaptively adjusting the contribution ratio of audio, vision, and spectrogram according to the input signal-to-noise ratio (SNR), introducing a lightweight noise classifier to detect the type of ambient noise in real time (such as traffic, human voices), and dynamically switching the fusion strategy. However, an efficient noise classifier needs to be designed and real-time performance needs to be ensured.

[0078] Example 2:

[0079] The present invention provides a microphone speech recognition system based on multimodal audio-visual fusion. This system can be used to implement the above-mentioned microphone speech recognition method based on multimodal audio-visual fusion. Specifically, it includes:

[0080] Audio feature extraction and word vector generation module: The speech signal is collected in real time through a microphone and sent to the cloud speech recognition API to be converted into a text word sequence. Then, the pre-trained Word2Vec model is used to map each word into a 300-dimensional word vector. If the word recognition fails, it is filled with a zero vector. The word vectors of multiple words are compressed through average pooling or a fully connected layer, and finally, a one-dimensional audio semantic feature vector is generated.

[0081] Visual feature extraction and log-Mel spectrogram generation module: First, Dlib is used to locate the face and lip regions, and key point alignment and normalization processing are performed on the lips to generate a sequence of lip motion images. The image sequence is input into a 3D CNN and a dense spatio-temporal CNN to extract lip motion visual features, and the spatial attention mechanism is used to focus on the key lip regions. Finally, the temporal features are extracted through a bidirectional GRU. At the same time, the speech signal is converted into a log-Mel spectrogram to enhance the perceptual characteristics of the audio and generate log-Mel spectrogram features.

[0082] Multi-modal feature fusion and decoding module: The audio semantic feature vector, lip motion visual feature, and log-Mel spectrogram feature are concatenated into a fused feature vector and sent to the CTC decoder for decoding, and the final text is output in combination with the Beam Search algorithm. During the training process, the Adam optimizer and the mini-batch training strategy are adopted to improve the accuracy and generalization ability of the model.

[0083] Embodiment 3:

[0084] This embodiment provides a terminal device, which includes a processor and a memory. The memory is used to store a computer program, and the computer program includes program instructions. The processor is used to execute the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions to implement the corresponding method flow or corresponding function. The processor described in the embodiment of the present invention can be used for the operation of a microphone speech recognition method based on multi-modal audio-visual fusion, including the following steps:

[0085] Step S1: Real-time collect the voice signal through a microphone, and send it to the cloud speech recognition API to convert it into a text word sequence; then, use the pre-trained Word2Vec model to map each word into a 300-dimensional word vector. If the word recognition fails, fill it with a zero vector. The word vectors of multiple words are compressed through average pooling or a fully connected layer, and finally a one-dimensional audio semantic feature vector is generated;

[0086] Step S2: First, use Dlib to locate the face and lip regions, and perform key point alignment and normalization processing on the lips to generate a lip movement image sequence; the image sequence is input into a 3D CNN and a dense spatio-temporal CNN to extract lip movement visual features, and the spatial attention mechanism is used to focus on the key lip regions, and finally the temporal features are extracted through a bidirectional GRU; at the same time, the voice signal is converted into a log-Mel spectrogram to enhance the perceptual characteristics of the audio and generate log-Mel spectrogram features;

[0087] Step S3: Concatenate the audio semantic feature vector, lip movement visual feature, and log-Mel spectrogram feature into a fused feature vector, send it to the CTC decoder for decoding, and combine the Beam Search algorithm to output the final text; during the training process, the Adam optimizer and the mini-batch training strategy are adopted to improve the accuracy and generalization ability of the model.

[0088] Example 4:

[0089] This embodiment provides a storage medium, specifically a computer-readable storage medium (Memory). The computer-readable storage medium is a memory device in a terminal device, used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the terminal device and, of course, the extended storage medium supported by the terminal device. The computer-readable storage medium provides a storage space, and this storage space stores the operating system of the terminal. And, in this storage space, one or more instructions suitable for being loaded and executed by the processor are also stored. These instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory.

[0090] One or more instructions stored in the computer-readable storage medium can be loaded and executed by the processor to implement the corresponding steps of the method for microphone speech recognition based on multi-modal audio-visual fusion in the above embodiment; one or more instructions in the computer-readable storage medium are loaded and executed by the processor to perform the following steps:

[0091] Step S1: Real-time collect voice signals through a microphone and send them to the cloud speech recognition API to convert them into a text word sequence. Then, use the pre-trained Word2Vec model to map each word into a 300-dimensional word vector. If the word recognition fails, fill it with a zero vector. The word vectors of multiple words are compressed through average pooling or a fully connected layer, and finally, a one-dimensional audio semantic feature vector is generated.

[0092] Step S2: First, use Dlib to locate the face and lip regions, and perform key point alignment and normalization processing on the lips to generate a lip movement image sequence. The image sequence is input into a 3D CNN and a dense spatio-temporal CNN to extract lip movement visual features, and a spatial attention mechanism is used to focus on the key lip regions. Finally, temporal features are extracted through a bidirectional GRU. At the same time, the voice signal is converted into a log-Mel spectrogram to enhance the perceptual characteristics of the audio and generate log-Mel spectrogram features.

[0093] Step S3: Concatenate the audio semantic feature vector, lip movement visual features, and log-Mel spectrogram features into a fused feature vector, send it to a CTC decoder for decoding, and combine the Beam Search algorithm to output the final text. During the training process, the Adam optimizer and the mini-batch training strategy are adopted to improve the accuracy and generalization ability of the model.

[0094] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0095] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0096] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more of the processes and / or blocks Figure 1 in one or more of the processes and / or blocks Figure 1 specified in the function.

[0097] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, such that a series of operational steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more of the processes and / or blocks Figure 1 in one or more of the processes and / or blocks Figure 1 specified in the function.

[0098] The preferred embodiments of the present invention have been described in detail above, but the present invention is not limited to the above embodiments. Within the knowledge of those of ordinary skill in the art, various changes can be made without departing from the spirit of the present invention.

[0099] Many other changes and modifications can be made without departing from the concept and scope of the present invention. It should be understood that the present invention is not limited to the specific embodiments, and the scope of the present invention is defined by the appended claims.

Claims

1. A microphone speech recognition method based on multimodal audio-visual fusion, characterized in that, The method includes: Step S1: Real-time collect voice signals through a microphone and send them to the cloud speech recognition API to be converted into a text word sequence; then, use a pre-trained Word2Vec model to map each word into a 300-dimensional word vector. If the word recognition fails, fill it with a zero vector. The word vectors of multiple words are compressed through average pooling or a fully connected layer, and finally a one-dimensional audio semantic feature vector is generated; Step S2: First, use Dlib to locate the face and lip regions, and perform key point alignment and normalization processing on the lips to generate a lip movement image sequence; the image sequence is input into a 3D CNN and a dense spatio-temporal CNN to extract lip movement visual features, and a spatial attention mechanism is used to focus on the key lip regions, and finally temporal features are extracted through a bidirectional GRU; at the same time, the voice signal is converted into a log-Mel spectrogram to enhance the perceptual characteristics of the audio and generate log-Mel spectrogram features; Step S3: Concatenate the audio semantic feature vector, lip movement visual feature, and log-Mel spectrogram feature into a fused feature vector, send it to a CTC decoder for decoding, and combine the Beam Search algorithm to output the final text; during the training process, use an Adam optimizer and a mini-batch training strategy to improve the accuracy and generalization ability of the model.

2. The microphone speech recognition method based on multi-modal audio-visual fusion according to claim 1, wherein The said Step S1 includes: S101: Input the audio signal, real-time collect voice signals through a local microphone device, and transmit the voice data to the cloud speech recognition API for recognition and processing; S102: Generate a word list, the API converts the voice into a text word sequence and outputs it as a recognizable word list; S103: Word embedding vectorization, use a Word2Vec model pre-trained on the Google News corpus, words with similar contexts are close in the vector space, map each word in the word list output by the API into a 300-dimensional vector. If the API fails to recognize a certain word, fill it with a zero vector. If the word list contains multiple words, compress them through average pooling or a fully connected layer to form a one-dimensional feature vector, which is used as the final output of the audio module.

3. A microphone speech recognition method based on multimodal audio-visual fusion according to claim 1, characterized in that The said Step S2 includes: S201: Lip movement image preprocessing, use a Dlib linear classifier to locate the face, extract the face and lip regions, and generate a lip sequence image; align the lip region according to the key points to eliminate the influence of head pose changes, perform channel normalization on each frame of the image to reduce the interference of lighting changes, standardize the image data, and perform horizontal flipping on the sequence image during training to improve the generalization ability of the model; S202: Log-Mel spectrogram generation, divide the 3-second voice signal into frames with a 25ms window and a step size of 10ms, a total of 750 frames of spectrograms are generated; filter each frame of the spectrogram through a Mel triangular filter to convert the linear frequency to the Mel scale and enhance the perceptual characteristics of the voice; take the logarithm of the Mel spectrogram energy to obtain the log-Mel spectrogram and enhance the distinction between high-frequency and low-frequency features; S203: Construct a visual feature extraction network, including: 3D CNN module. 3D CNN adds three-dimensional convolutional kernel parameters on the basis of 2D CNN, enabling the feature maps in consecutive frames to be associated with consecutive frames in the previous layer and integrated into a single frame, ultimately achieving the extraction of motion information. It processes lip movement sequence images, extracts spatio-temporal features through 3D convolutional kernels, combines batch normalization, ReLU activation, and 3D max pooling layers, and outputs an initial feature map. Among them, is the weight parameter of the three-dimensional convolution kernel at position (p, q, r); Among them, is the input feature of the previous layer; Dense spatio-temporal CNN module: Adopts a dense short connection structure to alleviate gradient disappearance through short path connections, reduce the number of parameters, and improve the training efficiency. Specifically, it includes: Dense block: Each block contains a 6-layer structure of BN→ReLU→3D convolution. The feature maps between layers are passed through concatenation to enhance feature reuse. Transition block: Contains BN→ReLU→3D convolution→3D average pooling to compress the number of channels and reduce the size of the feature map, reducing the computational amount. Spatial attention module: Aggregates channel information through average pooling and max pooling layers to generate a spatial attention weight map. The original feature map is multiplied element-wise with the attention map to highlight the key regions of the lips. The calculation formula is as follows: M S (F) = σ(f 7×7 ([AvgFool(F); MaxPool(F)])): where σ is the Sigmoid function, and f 7×7 is a 7×7 convolution operation; Bidirectional GRU module: Processes temporal features, controls the information flow through update gates and reset gates, captures long temporal dependencies, and the output hidden states are concatenated to form a unified visual temporal feature vector, which is used as the final output of the visual module for subsequent multi-modal feature fusion.

4. A microphone speech recognition method based on multimodal audio-visual fusion according to claim 1, characterized in that The step S3 includes: Concatenate the audio semantic feature vector, lip movement visual feature, and log-Mel spectrogram feature to form a multi-modal fusion feature matrix, providing comprehensive information for subsequent decoding. Use the connectionist temporal classification loss function to solve the problem of inconsistent input and output sequence lengths and maximize the probability of the correct label sequence. Combine the beam search algorithm for decoding to generate the final text prediction result. During training, use the Adam optimizer with a learning rate of 0.0001 and a mini-batch size of 8 to improve the stability of the model and prevent overfitting.

5. A microphone speech recognition system based on multimodal audio-visual fusion, characterized in that, The system is applied to the method described in any one of claims 1-4. The system includes: Audio feature extraction and word vector generation module: Real-time collects voice signals through a microphone and sends them to the cloud speech recognition API to convert them into a text word sequence. Then, use the pre-trained Word2Vec model to map each word to a 300-dimensional word vector. If the word recognition fails, fill it with a zero vector. The word vectors of multiple words are compressed through average pooling or a fully connected layer, and finally, a one-dimensional audio semantic feature vector is generated. Visual feature extraction and log-Mel spectrogram generation module: First, use Dlib to locate the face and lip regions, perform key point alignment and normalization processing on the lips to generate a lip movement image sequence. The image sequence is input into 3D CNN and dense spatio-temporal CNN to extract lip movement visual features, and the spatial attention mechanism is used to focus on the key regions of the lips. Finally, temporal features are extracted through bidirectional GRU. At the same time, the voice signal is converted into a log-Mel spectrogram to enhance the perceptual characteristics of the audio and generate log-Mel spectrogram features. Multi-modal Feature Fusion and Decoding Module: Concatenate the audio semantic feature vector, lip movement visual feature, and log-Mel spectrum feature into a fused feature vector, feed it into the CTC decoder for decoding, and combine with the Beam Search algorithm to output the final text; during the training process, use the Adam optimizer and mini-batch training strategy to improve the accuracy and generalization ability of the model.

6. A computer device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements a microphone speech recognition method based on multi-modal audio-visual fusion according to any one of claims 1 to 4.

7. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the program is executed by the processor, it implements a microphone speech recognition method based on multi-modal audio-visual fusion according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Short text semantic similarity distinguishing method and system based on deep learning model Word2Vec

    CN106844346A

  • Noise-robust audio and video bimodal speech recognition method and system

    CN111754992A

  • Audio-visual bimodal speech recognition method based on convolutional block attention mechanism

    CN112216271A

  • License plate character recognition method with self-correction consciousness

    CN113591863A

  • Speech synthesis method and device, computing equipment, storage medium and program product

    CN114373443A

Cited By

  • Mandarin pronunciation real-time correction method and system based on multi-modal streaming learning

    CN121393456A