Cross-modal fusion speech recognition method and system based on image enhancement and large model enabling
By combining audio and video features and large language models, the problem of degradation in performance of traditional speech recognition systems in noisy environments is solved, and higher speech recognition accuracy and speech signal quality are achieved.
Patent Information
- Application Number
- CN202510270114.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-03
AI Technical Summary
In noisy and challenging environments, traditional speech recognition systems are susceptible to environmental noise interference, resulting in a degradation in speech recognition performance.
By obtaining the speaker's facial images and lip moving images in audio and video, encoding and cross-modal fusion are performed, the temporal dynamic relationship between audio and visual features is captured, and the speech recognition results are self-corrected and decoded using a large language model.
It improves the speech recognition accuracy in different noise environments, enhances the quality of speech signals, and obtains voice content more accurately through the context analysis capabilities of large language models.
Smart Images

Figure CN120089138A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing and the field of speech recognition, and particularly relates to a cross-modal fusion speech recognition method and system with image enhancement and large model empowerment. Background Art
[0002] The purpose of Automatic Speech Recognition (ASR) technology is to enable machines to "understand" human speech. It is an important branch of artificial intelligence and the basis for realizing natural, fluent, and efficient interaction between machines and humans, showing great value in multiple scenarios such as intelligent customer service, intelligent education, and smart home. Inspired by the multi-modal perception of the human brain (including images, audio, etc.), people have shown great interest in equipping machines with similar capabilities. With the rapid development of deep learning, the recognition accuracy of the latest speech recognition systems in quiet environments can exceed that of humans. However, in various noisy and challenging environments, the speech quality will significantly degrade in the noise signal, seriously affecting the speech recognition performance. However, human perception largely depends on hearing and vision to understand the external environment. For example, in a scenario with multiple speakers, humans can effortlessly identify the speaker and enhance the language of the target speaker through visual cues such as precise lip movements. Further, large language models have powerful language understanding and modeling capabilities, and can correct the often misrecognized problems in the speech decoding process by using language knowledge, becoming an important supplement to improve speech recognition performance. Therefore, how to use large language models to empower the speech recognition of audio-visual fusion based on images and audio in audio-visual to improve the speech recognition accuracy under complex background noise conditions is still a key technical problem to be solved urgently. Summary of the Invention
[0003] The technical problem to be solved by the present invention: Aiming at the above problems of the prior art, the present invention provides a cross-modal fusion speech recognition method and system with image enhancement and large model empowerment. The present invention aims to enhance the audio signal based on the speaker's facial image and the corresponding lip movement image in the audio-visual, capture the temporal dynamic relationship between audio and visual features to improve the accuracy of speech recognition, and make full use of the context comprehensive analysis ability of the large language model to self-correct and decode the speech recognition result, and more accurately obtain the speech content to improve the speech recognition accuracy in different noise environments.
[0004] To solve the above technical problems, the technical solution adopted by the present invention is as follows: A cross-modal fusion speech recognition method with image enhancement and large model empowerment, including the following steps: S1, obtaining the speaker's facial image frames in the audio-visual F and the corresponding lip movement image frames Land the audio signal A ; S2, respectively encode the speaker's facial image frames F and the corresponding lip movement image frames L and the audio signal A ; S3, perform cross-modal fusion on the encoded features to obtain context features Oa ; S4, input the context features Oa into the speech recognition decoder enhanced by the large model to obtain the final speech recognition result, including: S4.1, project the context features Oa onto the corresponding label category vector space through linear transformation by a multi-layer perceptron MLP; S4.2, use the Softmax activation function on the linearly transformed context features Oa to obtain the probability distribution of the categories ; S4.3, select and retain the Top-k words with high probabilities from the probability distribution of the categories ; S4.4, use the large language model to correct the Top-k words with high probabilities to obtain the final speech recognition result ; .
[0005] Optionally, when encoding the speaker's facial image frames F and the corresponding lip movement image frames L and the audio signal A in step S2, encoding the speaker's facial image frames F means decoding the speaker's facial image frames F using a facial feature decoder to extract visual features Hf .
[0006] Optionally, when encoding the speaker's facial image frames F and the corresponding lip movement image frames L and the audio signal A in step S2, encoding the lip movement image frames L means decoding the lip movement image frames L using a lip feature decoder to extract lip features Hl .
[0007] Optionally, when encoding the speaker's facial image frames F and the corresponding lip movement image frames L and the audio signal A in step S2, encoding the audio signal A means decoding the audio signal A using a speech feature decoder to extract audio features Ha .
[0008] Optionally, in step S3, the encoded features are cross-modally fused to obtain context features Oa including: S3.1, for the lip features Hl obtain the lip context correlation representation Fl , and use the modality interaction network for the obtained audio features Ha and visual features Hf to perform inter-modal correlation representation to obtain enhanced speech features Fa ; S3.2, use the modality interaction network for the obtained lip context correlation representation Fl and enhanced speech features Fa to perform interactive fusion to obtain context features Oa .
[0009] Optionally, the function expression for obtaining the lip context correlation representation for the lip features Hl is: Fl : , , , , wherein, is a feed-forward network, , and are all intermediate features, is a Mamba layer, is a convolutional layer, is a layer normalization operation; the function expression for using the modality interaction network to perform inter-modal correlation representation on the obtained audio features Ha and visual features Hf to obtain enhanced speech features Fa is: , , , , , wherein, and are respectively the audio features and visual features after passing through the feed-forward network and skip connection, and are all intermediate features, is a cross-modal attention operation, and there is: , , , , , Among them, is the splicing operation, ~ are the outputs of the 1st to attention heads, is the output weight matrix, is for any th attention head output, is the softmax activation function, , and are respectively the query, key, and value of the cross-modal attention operation, , and are respectively the weight parameters of the query, key, and value, is the dimension of the key.
[0010] Optionally, the function expression for the context feature Fl obtained by using the modality interaction network to interact and fuse the lip context correlation representation Fa and the enhanced speech feature Oa is: , , , , , Among them, and are respectively the lip context correlation representation and the enhanced speech feature after passing through the feed-forward network and skip connection, is the feed-forward network, and are both intermediate features, is the cross-modal attention operation, is the convolutional layer, is the layer normalization operation, and there are: , , , , , Among them, is a splicing operation, ~ is the output of the 1st to th attention heads, is the output weight matrix, is for any th attention head output, is the softmax activation function, , and are the query, key, and value of the cross-modal attention operation respectively, , and and are the weight parameters of the query, key, and value respectively, is the dimension of the key.
[0011] In addition, the present invention also provides a cross-modal fusion speech recognition system with image enhancement and large model empowerment, including a microprocessor and a memory connected to each other, and the microprocessor is programmed or configured to execute the cross-modal fusion speech recognition method of the present invention.
[0012] In addition, the present invention also provides a computer-readable storage medium, in which a computer program or instruction is stored, and the computer program or instruction is programmed or configured to execute the cross-modal fusion speech recognition method of the present invention through a processor.
[0013] In addition, the present invention also provides a computer program product, including a computer program or instruction, and the computer program or instruction is programmed or configured to execute the cross-modal fusion speech recognition method of the present invention through a processor.
[0014] Compared with the prior art, the present invention mainly has the following beneficial effects: In order to solve the problem that traditional audio-based automatic speech recognition systems are vulnerable to environmental noise interference when performing speech interaction in a noisy environment, on the one hand, the present invention enhances the audio signal based on the speaker's facial image and the corresponding lip movement image in the audio-visual, separates the target speech from the background noise through facial visual information to improve the quality of the speech signal, further combines the enhanced audio signal with the lip visual information to improve the accuracy of speech recognition, and captures the temporal dynamic relationship between the audio and visual features to more accurately obtain the speech content expressed by the user; on the other hand, by making full use of the context comprehensive analysis ability of the large language model to self-correct and decode the speech recognition result, it can more accurately obtain the speech content to improve the speech recognition accuracy in different noise environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 It is a schematic diagram of the basic process of the method according to the embodiment of the present invention.
[0016] Figure 2 It is a schematic diagram of the process of step S4 in the embodiment of the present invention.
[0017] Figure 3 It is a schematic diagram of the overall network structure in the embodiment of the present invention. Detailed implementation manners
[0018] In order to enable those skilled in the art to better understand the technical solutions of the present invention, the technical solutions of the present invention will be further described in detail below with reference to the accompanying drawings in the embodiments of the present invention.
[0019] As Figure 1 shown, this embodiment provides a cross-modal fusion speech recognition method with image enhancement and large model empowerment, including the following steps: S1, obtaining the speaker's facial image frames F and corresponding lip movement image frames L and audio signals A ; S2, respectively encoding the speaker's facial image frames F and corresponding lip movement image frames L and audio signals A ; S3, performing cross-modal fusion on the encoded features to obtain context features Oa ; S4, inputting the context features Oa into a large model-enhanced speech recognition decoder to obtain the final speech recognition result.
[0020] As Figure 2 shown, in this embodiment, step S4 inputs the context features Oa into a large model-enhanced speech recognition decoder to obtain the final speech recognition result, including: S4.1, performing a linear transformation on the context features Oa through a multi-layer perceptron MLP to project them into the corresponding label category vector space; for the video, the output context features Oa constitute the output target sequence Y ={ y i | i =, 1, 2, 3,..., L}, where L is the length of the output sequence, and performing a linear transformation on each context feature Y in the target sequence Oa through a multi-layer perceptron MLP to project them into the corresponding label category vector space; S4.2, Obtain the probability distribution of the category by using the Softmax activation function (normalized exponential function) for the context features after linear transformation Oa which can be expressed as: , where Oa is the context feature with respect to the th category and is the Softmax activation function (normalized exponential function); S4.3, Select and retain the Top-k words with the highest probabilities from the probability distribution of the category which can be expressed as: where represents selecting the top k probabilities after sorting based on the probability distribution of the category ; S4.4, Use the large language model to correct the Top-k words with the highest probabilities to obtain the final speech recognition result which can be expressed as: where is the operation of correction by the large language model, so that the large language model LLaMA selects the optimal recognition result from the Top-k words with the highest probabilities according to the context information. By fully utilizing the context comprehensive analysis ability of the large language model to self-correct and decode the speech recognition result, the speech content can be obtained more accurately to improve the speech recognition accuracy in different noise environments. For example: The following is a demo of the correction operation prompt word based on the complete sentence "We can prevent the worst case scenario", focusing on the correction operation of recognizing the word "scenario". The task is to use the context information to correct the candidate words for speech recognition and select the optimal result that best conforms to the semantics and context from the Top-k words with the highest probabilities.
[0021] Input: - List of speech recognition candidate words (Top-k): ** ["scenario", "scenery"] - Context information: ** "We can prevent the worst case ____." Output: - Corrected optimal word: "scenario" Analysis of the correction process: ** 1. Candidate word 1: "scenario" - Problem: No problem, conforms to grammar and semantics.
[0022] - Revision suggestion: Keep it.
[0023] 2. Candidate word 2: "scenery" - Problem: Semantic mismatch. "Scenery" (landscape) does not match "worst case" in the context.
[0024] - Revision suggestion: Exclude it.
[0025] Final revised result - Optimal word: "scenario". Through this revision operation, the large language model can more accurately select the optimal result from the candidate words and improve the accuracy of speech recognition.
[0026] As Figure 3 shown, when encoding the speaker's facial image frames F and the corresponding lip movement image frames L and the audio signal A in step S2 of this embodiment, encoding the speaker's facial image frames F means decoding the speaker's facial image frames F using a facial feature decoder to extract visual features Hf . It should be noted that the facial feature decoder is an existing decoder, and the type of facial feature decoder required can be selected according to actual needs. For example, as an optional implementation manner, the facial feature decoder adopted in this embodiment is composed of a 3D convolutional layer and a ResNet-18 network. The visual sequence is segmented from the original video, face detection and face alignment are performed, and then the facial region is adjusted to a size of 120*120 pixels and converted into a grayscale frame, and the key points extracted from the entire frame are smoothed to obtain the speaker's facial image frames F , and finally the speaker's facial image frames F are input into a facial feature decoder using a 3D convolutional layer followed by a ResNet-18 network for decoding and embedded into a 521-dimensional vector to obtain the extracted visual features Hf .
[0027] As Figure 3 shown, when encoding the speaker's facial image frames F and the corresponding lip movement image frames L and the audio signal A in step S2 of this embodiment, encoding the lip movement image frames L means decoding the lip movement image frames L using a lip feature decoder to extract lip featuresHl . It should be noted that the lip feature decoder is an existing decoder, and the required lip feature decoder type can be selected according to actual needs. This example is an optional implementation. The lip feature decoder used in this embodiment is composed of a 3D convolutional layer and a ResNet-18 network. A fixed bounding box of the region of interest (RoI) around the lips is derived from the aligned face frame, and the "clean" motion signal is finely distinguished within the region of interest. Then, the lip feature decoder is used to perform the same steps of facial feature extraction to obtain the final lip features. Hl .
[0028] like Figure 3 As shown, in step S2 of this embodiment, the speaker's facial image frames are respectively F and the corresponding lip motion image frames L and audio signal A When encoding, the audio signal A Encoding is the process of converting audio signals into A Use the speech feature decoder to decode and extract audio features Ha It should be noted that the speech feature decoder is an existing decoder, and the required speech feature decoder type can be selected according to actual needs. For example, as an optional implementation, the speech feature decoder used in this embodiment decodes and extracts audio features. Ha Includes: For audio signals A Perform short-time Fourier transform to transform the audio signal A Convert to Mel spectrogram to obtain spectral amplitude, then apply Mel-scale logarithmic filter to the obtained spectral amplitude, and finally process the Mel-scale amplitude feature of the filtered audio through a 2D convolutional network to obtain audio features Ha .
[0029] like Figure 3 As shown, in step S3 of this embodiment, the encoded features are cross-modally fused to obtain context features. Oa include: S3.1, lip features Hl Get lip context representation Fl , the audio features obtained using the modal interaction network Ha and visual features Hf Enhanced speech features are obtained by performing inter-modal correlation representation Fa ; S3.2, using the modal interaction network to obtain the lip context association representation Fl and enhanced speech features Fa Interactive fusion to obtain context features Oa .
[0030] likeFigure 3 As shown, in this embodiment, for lip features Hl The network module for obtaining the lip context - associated representation Fl is simply referred to as the ConMamba network. Through the ConMamba network, for lip features Hl obtain the lip context - associated representation Fl The functional expression is: , , , , where, is a feed - forward network, , and are all intermediate features, is the Mamba layer, is the convolutional layer, is the layer normalization operation; The enhanced speech feature Ha and visual feature Hf obtained by using the cross - modal interaction network for the obtained audio features Fa The functional expression for the inter - modal association representation is: , , , , , where, and are the audio feature and visual feature respectively after passing through the feed - forward network and skip connection, and are all intermediate features, is the cross - modal attention operation, and there is: , , , , , where, is the concatenation operation, ~ are the outputs of the 1st to th attention heads, is the output weight matrix, is the output of any th attention head, is the softmax activation function, , and are the query, key, and value of the cross-modal attention operation respectively, , and are the weight parameters of the query, key, and value respectively, is the dimension of the key.
[0031] In this embodiment, the modal interaction network is used to interact and fuse the obtained lip context correlation representation Fl and the enhanced speech feature Fa to obtain the context feature Oa The functional expression is: , , , , , where and are the lip context correlation representation and the enhanced speech feature after passing through the feed-forward network and the skip connection respectively, is the feed-forward network, and are both intermediate features, is the cross-modal attention operation, is the convolutional layer, is the layer normalization operation, and there is: , , , , , where is the concatenation operation, ~ are the outputs of the 1st to th attention heads, is the output weight matrix, is the output of any th attention head, is the softmax activation function, , and are the query, key, and value of the cross-modal attention operation respectively, , and are the weight parameters of the query, key, and value respectively, is the dimension of the key.
[0032] To verify the cross-modal fusion speech recognition method with image enhancement and large model empowerment in this embodiment, the model is trained and tested on the large-scale audio-visual dataset LRS2 in this embodiment to obtain the optimal speech recognition model. The LRS2 dataset collects thousands of hours of spoken sentences and phrases, as well as the corresponding faces; LRS2 consists of 143,000 utterances, which contains 2.3 million words and a vocabulary of 41,000. As a comparison of the method in this embodiment: the comparison method CM-seq2seq method (see details in P. Ma, S. Petridis, M. Pantic, End-to-end audio-visual speech recognition with conformers, in: 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2021, pp. 7613–7617.) and EG-seq2seq method (see details in B. Xu, C. Lu, Y. Guo, J. Wang, Discriminative multi-modality speech recognition, in: Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 14433–14442.). Moreover, the Word Error Rate (WER) is used to measure the accuracy of the recognition result. The lower the word error rate, the better the recognition effect. The word error rate is the ratio of the edit distance and the label length. The edit distance is a metric for measuring the similarity of two strings, generally referring to the minimum number of edit operations required to convert one string into another through three edit operations: word replacement, word insertion, and word deletion. Table 1 shows the recognition results of the method in this embodiment and the CM-seq2seq and EG-seq2seq methods on the LRS2 dataset.
[0033] Table 1 Comparison of the recognition results of the method in this embodiment and the existing methods on the LRS2 dataset
[0034] As can be seen from Table 1, compared with the CM-seq2seq and EG-seq2seq methods, the method of this embodiment obtains a lower word error rate than the CM-seq2seq and EG-seq2seq methods at the signal-to-noise ratio levels of -5, 0, 5, 10, or clean in the LRS2 dataset, that is, a better recognition effect is obtained.
[0035] In summary, to solve the problem that traditional audio-based automatic speech recognition systems are vulnerable to environmental noise interference when performing speech interaction in a noisy environment, considering that complex background noise in practical applications has a significant impact on speech quality and clarity, seriously weakening the performance of speech recognition, and the language interaction of humans is essentially multimodal. In addition to auditory information, visual information also plays an important role in language understanding. The method of this embodiment separates the target speech from the background noise through facial visual information to improve the quality of the speech signal, further combines the enhanced audio signal with lip visual information to improve the accuracy of speech recognition, and captures the temporal dynamic relationship between audio and visual features. This embodiment further self-corrects the decoding of the speech recognition result through the context comprehensive analysis ability of the large language model in the speech decoding stage, more accurately obtains the speech content, and improves the speech recognition accuracy in different noise environments.
[0036] In addition, this embodiment also provides a cross-modal fusion speech recognition system with image enhancement and large model empowerment, including a microprocessor and a memory connected to each other, and the microprocessor is programmed or configured to execute the cross-modal fusion speech recognition method with image enhancement and large model empowerment.
[0037] In addition, this embodiment also provides a computer-readable storage medium, in which a computer program or instruction is stored, and the computer program or instruction is programmed or configured to execute the cross-modal fusion speech recognition method with image enhancement and large model empowerment through a processor.
[0038] In addition, this embodiment also provides a computer program product, including a computer program or instruction, and the computer program or instruction is programmed or configured to execute the cross-modal fusion speech recognition method with image enhancement and large model empowerment through a processor.
[0039] Those skilled in the art should understand that the technical solutions provided by the embodiments of the present invention can be in the form of a method, a system, or a computer program product. Therefore, the present invention can be implemented in the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can be in the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code. The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, as well as the combination of processes and / or blocks in the flowchart and / or block diagram. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in the process Figure 1 one process or multiple processes and / or blocks Figure 1 These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in the process Figure 1 one process or multiple processes and / or blocks Figure 1 These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Therefore, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0040] The above is only the preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions within the idea of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A cross-modal fusion speech recognition method with image enhancement and large model empowerment, characterized in that: The following steps are included: S1, obtaining a facial image frame of the speaker in the audio and video F and the corresponding lip motion image frames L and audio signal A ; S2, respectively, the speaker's facial image frame F and the corresponding lip motion image frames L and audio signal A Encode; S3, cross-modal fusion of the encoded features to obtain context features Oa ; S4, the context features Oa Input the large model enhanced speech recognition decoder to obtain the final speech recognition result, including: S4.1, through the multi-layer perceptron MLP to the context features Oa Perform a linear transformation to project it into the corresponding label category vector space; S4.2, the context features after linear transformation Oa Use the Softmax activation function to get the probability distribution of the category ; S4.3, probability distribution of categories Select and retain the top-k words with high probability ; S4.4, Top-k high probability words Use the large language model to correct the final speech recognition result .
2. The image enhancement and large model-enabled cross-modal fusion speech recognition method according to claim 1 is characterized in that: In step S2, the speaker's facial image frames are respectively F and the corresponding lip motion image frames L and audio signal A When encoding, the speaker's facial image frame F Encoding refers to converting the speaker's facial image frame into F Decode using facial feature decoder to extract visual features Hf .
3. The image enhancement and large model-enabled cross-modal fusion speech recognition method according to claim 2 is characterized in that: In step S2, the speaker's facial image frames are respectively F and the corresponding lip motion image frames L and audio signal A When encoding, the lip motion image frame L Encoding means converting the lip motion image frames L Lip feature decoder is used to extract lip features Hl .
4. The image enhancement and large model-enabled cross-modal fusion speech recognition method according to claim 3 is characterized in that: In step S2, the speaker's facial image frames are respectively F and the corresponding lip motion image frames L and audio signal A When encoding, the audio signal A Encoding is the process of converting audio signals into A Use the speech feature decoder to decode and extract audio features Ha .
5. The image enhancement and large model-enabled cross-modal fusion speech recognition method according to claim 4 is characterized in that: In step S3, the encoded features are cross-modally fused to obtain context features. Oa include: S3.1, lip features Hl Get lip context representation Fl , the audio features obtained using the modal interaction network Ha and visual features Hf Enhanced speech features are obtained by performing inter-modal correlation representation Fa ; S3.2, using the modal interaction network to obtain the lip context association representation Fl and enhanced speech features Fa Interactive fusion to obtain context features Oa .
6. The image enhancement and large model-enabled cross-modal fusion speech recognition method according to claim 5, characterized in that: The lip features Hl Get lip context representation Fl The function expression is: , , , , in, is a feed-forward network, , and All are intermediate features. For the Mamba layer, is the convolutional layer, is a layer normalization operation; the audio features obtained by using the modal interaction network Ha and visual features Hf Enhanced speech features are obtained by performing inter-modal correlation representation Fa The function expression is: , , , , , in, and They are the audio features and visual features after the feedforward network and jump connection, and All are intermediate features. is a cross-modal attention operation, and has: , , , , , in, For splicing operation, ~ For the 1st~ The output of an attention head is is the output weight matrix, For any The output of an attention head is is the softmax activation function, , and are the query, key, and value of the cross-modal attention operation, respectively. , and are the weight parameters for query, key and value respectively, The dimension of the key.
7. The image enhancement and large model-enabled cross-modal fusion speech recognition method according to claim 5, characterized in that: The lip context association representation obtained by using the modal interaction network Fl and enhanced speech features Fa Interactive fusion to obtain context features Oa The function expression is: , , , , , in, and They are the lip context representation and enhanced speech features after the feedforward network and skip connection, is a feed-forward network, and All are intermediate features. is the cross-modal attention operation, is the convolutional layer, is the layer normalization operation, and: , , , , , in, For splicing operation, ~ For the 1st~ The output of an attention head is is the output weight matrix, For any The output of an attention head is is the softmax activation function, , and are the query, key, and value of the cross-modal attention operation, respectively. , and and are weight parameters for query, key, and value respectively. The dimension of the key.
8. An image-enhanced and large-model-enabled cross-modal fusion speech recognition system, comprising a microprocessor and a memory connected to each other, characterized in that: The microprocessor is programmed or configured to execute the image enhancement and large model-enabled cross-modal fusion speech recognition method as described in any one of claims 1 to 7.
9. A computer-readable storage medium having a computer program or instruction stored therein, characterized in that: The computer program or instruction is programmed or configured to execute the image enhancement and large model-enabled cross-modal fusion speech recognition method described in any one of claims 1 to 7 through a processor.
10. A computer program product comprising a computer program or instructions, characterized in that The computer program or instruction is programmed or configured to execute the image enhancement and large model-enabled cross-modal fusion speech recognition method described in any one of claims 1 to 7 through a processor.
Citation Information
Cited By
Training method of humanoid robot language recognition model and language recognition method
CN121260149A