A method, device, and medium for fusing audio and video semantic recognition
Patent Information
- Application Number
- CN202610442519.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-07
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2046-04-07
AI Technical Summary
[0004]本发明提供一种融合音频与视频的语义识别方法和装置,以解决在复杂声学环境中语音识别效果差的技术问题
[0016] The beneficial effects of this invention are as follows: By dynamically estimating the signal-to-noise ratio (SNR) and assigning weights, it effectively solves the problem of static and singular audio-video fusion strategies in noisy environments found in existing technologies. First, weights are dynamically calculated based on the quality of real-time audio information, strengthening the auxiliary role of video information at low SNR and relying more on audio information at high SNR, thus achieving adaptive adjustment of modal contribution. Second, feature alignment and concatenation ensure the spatiotemporal consistency of cross-modal information, laying the foundation for subsequent fusion. Finally, a pre-trained feature fusion model is introduced, combining weights to split and re-fuse joint features, enabling more refined integration of bimodal information, thereby significantly improving the accuracy of semantic recognition and the robustness of the overall system in complex acoustic environments.
Smart Images

Figure CN121983035B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech recognition, and more particularly to a semantic recognition method and apparatus that integrates audio and video. Background Technology
[0002] In complex acoustic environments such as vehicle cockpits and industrial control centers, the performance of standalone speech recognition systems deteriorates sharply when background noise is excessive. To improve robustness, existing technologies often incorporate video information as an aid. By extracting the speaker's lip movement features and fusing them with audio features, recognition accuracy can be improved in noisy environments. Current mainstream fusion methods typically involve simply concatenating or weighting audio and video features before inputting them into the recognition model for processing.
[0003] However, existing methods still suffer from a static and simplistic fusion strategy. Lacking dynamic awareness of audio quality, these methods cannot adaptively adjust the weights of audio and video based on changes in the environmental signal-to-noise ratio. In scenarios with severe noise interference, excessive weighting of low-quality audio can suppress the complementary effect of video information, leading to poor fusion results and consequently affecting the accuracy and stability of overall semantic recognition. Therefore, improvements are needed. Summary of the Invention
[0004] This invention provides a semantic recognition method and apparatus that integrates audio and video to solve the technical problem of poor speech recognition performance in complex acoustic environments.
[0005] This invention provides a semantic recognition method that integrates audio and video, comprising:
[0006] The audio information is subjected to Mel spectrum conversion to obtain a Mel spectrum feature map, and the feature map is extracted to obtain audio features. Feature extraction is performed on the video information to obtain lip features; the video information and the audio information are acquired simultaneously, and the video information includes lip information; The audio features and the lip features are aligned and concatenated to obtain a joint feature; The signal-to-noise ratio (SNR) of the Mel spectral feature map is estimated to obtain the SNR, and the audio weight and video weight are obtained based on the SNR and a preset weight allocation function. Using a pre-trained feature fusion model, the joint features are split based on the audio weights and the video weights to obtain audio branch features and video branch features. The audio branch features, the video branch features, and the joint features are then fused to obtain fused features. Semantic recognition is performed on the fused features to obtain the recognition results.
[0007] In one embodiment of the present invention, the step of extracting features from the Mel spectral feature map to obtain audio features includes: The frequency domain attention weights are calculated using a preset frequency domain attention function and the Mel spectral feature map. The Mel spectral feature map is element-wise weighted based on the frequency domain attention weights to obtain a weighted feature map. Feature extraction is performed on the weighted feature map to obtain audio features.
[0008] In one embodiment of the present invention, the step of extracting features from video information to obtain lip features includes: Spatiotemporal feature extraction is performed on the video information to obtain basic lip features; The lip mask is calculated using a preset lip region attention function and the basic lip features; The lip mask is element-wise multiplied with the basic lip features to obtain the weighted lip features; The weighted lip features are augmented with positional information to obtain lip features.
[0009] In one embodiment of the present invention, the pre-trained feature fusion model performs feature splitting on the joint features based on the audio weights and the video weights to obtain audio branch features and video branch features, and then performs feature fusion on the audio branch features, the video branch features, and the joint features to obtain fused features, including: By using a pre-trained feature fusion model and a multi-head attention mechanism, the joint features are subjected to multi-head attention calculation based on relative position to obtain attention output features; Based on the audio weights, the attention output features are split to obtain audio branch features; Based on the video weights, the attention output features are split to obtain video branch features; Based on the audio weights and the video weights, the modal difference perception coefficient is calculated; Based on the modal difference perception coefficient and the preset modal fusion coefficient, the audio branch features, the video branch features, and the attention output features are fused, and then the features are refined to obtain the fused features.
[0010] In one embodiment of the present invention, the step of performing feature decomposition on the attention output features based on the audio weights to obtain audio branch features includes: Based on the audio weights and the attention output features, calculate the audio modality mask; The initial audio branch features are obtained by element-wise multiplication of the audio modality mask and the attention output features; After downsampling the initial audio branch features to reduce the feature dimension, the features are processed by a nonlinear activation function and then upsampling the features to restore the feature dimension, thus obtaining the audio branch features.
[0011] In one embodiment of the present invention, the step of performing feature decomposition on the attention output features based on the video weights to obtain video branch features includes: Calculate the video modality mask based on the video weights and the attention output features; The initial video branch features are obtained by element-wise multiplication of the video modality mask and the attention output features; After downsampling the initial video branch features to reduce the feature dimension, the features are processed by a nonlinear activation function and then upsampling the features to restore the feature dimension, thus obtaining the video branch features.
[0012] In one embodiment of the present invention, the step of performing semantic recognition on the fused features to obtain a recognition result includes: The fusion features are subjected to linear transformation and normalization to generate the character probability distribution corresponding to each time step, and the character with the highest probability value is selected step by step. All selected characters are concatenated sequentially and deduplicated to obtain the predicted character sequence; The predicted character sequence is sequentially encoded and semantically optimized to obtain encoded features; The encoded features are decoded into natural language text and the format is adjusted to obtain the recognition result.
[0013] The present invention also provides a semantic recognition device that integrates audio and video, comprising: The audio feature acquisition module is used to perform Mel spectrum conversion on the audio information to obtain a Mel spectrum feature map, and to extract features from the Mel spectrum feature map to obtain audio features; The lip feature acquisition module is used to extract features from video information to obtain lip features; the video information and the audio information are acquired simultaneously, and the video information includes lip information; The feature splicing module is used to align and splice the audio features and the lip features to obtain joint features; The weight calculation module is used to estimate the signal-to-noise ratio of the Mel spectral feature map, obtain the signal-to-noise ratio, and obtain the audio weight and video weight based on the signal-to-noise ratio and a preset weight allocation function. The feature fusion module is used to perform feature splitting on the joint features based on the audio weights and the video weights using a pre-trained feature fusion model to obtain audio branch features and video branch features, and to perform feature fusion on the audio branch features, the video branch features, and the joint features to obtain fused features; The semantic recognition module is used to perform semantic recognition on the fused features to obtain the recognition result.
[0014] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the semantic recognition method for fusing audio and video.
[0015] The present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the semantic recognition method for fusing audio and video.
[0016] The beneficial effects of this invention are as follows: By dynamically estimating the signal-to-noise ratio (SNR) and assigning weights, it effectively solves the problem of static and singular audio-video fusion strategies in noisy environments found in existing technologies. First, weights are dynamically calculated based on the quality of real-time audio information, strengthening the auxiliary role of video information at low SNR and relying more on audio information at high SNR, thus achieving adaptive adjustment of modal contribution. Second, feature alignment and concatenation ensure the spatiotemporal consistency of cross-modal information, laying the foundation for subsequent fusion. Finally, a pre-trained feature fusion model is introduced, combining weights to split and re-fuse joint features, enabling more refined integration of bimodal information, thereby significantly improving the accuracy of semantic recognition and the robustness of the overall system in complex acoustic environments. Attached Figure Description
[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. It is obvious that the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.
[0018] In the attached diagram: Figure 1 A flowchart illustrating a semantic recognition method that integrates audio and video according to an embodiment of the present invention; Figure 2 This is a schematic diagram of data processing provided in an embodiment of the present invention.
[0019] Figure 3 This is a schematic diagram of a semantic recognition device that integrates audio and video according to an embodiment of the present invention; Figure 4 This is a schematic diagram of an electronic device provided in one embodiment of the present invention. Detailed Implementation
[0020] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. In the absence of conflict, the following embodiments and features in the embodiments can be combined with each other.
[0021] It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. The drawings only show the components related to the present invention and are not drawn according to the actual number, shape and size of the components in the actual implementation. In the actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.
[0022] In the following description, numerous details are explored to provide a more thorough explanation of embodiments of the invention. However, it will be apparent to those skilled in the art that embodiments of the invention may be practiced without these specific details. In other embodiments, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring embodiments of the invention.
[0023] This invention discloses a semantic recognition method that integrates audio and video, applicable to interactive scenarios with noisy environments and significant background noise, such as in-vehicle cockpits and industrial control centers. In these scenarios, single-function speech recognition systems are prone to performance degradation due to background noise interference. Existing audio-video fusion solutions often employ fixed strategies, failing to fully consider real-time changes in audio signal quality. When the environmental signal-to-noise ratio is too low, excessive involvement of low-quality audio information can interfere with or even inhibit the effective use of video information, resulting in unsatisfactory overall recognition performance. To address this issue, this invention dynamically adjusts weights based on actual audio quality, thereby achieving more robust and accurate semantic recognition under complex acoustic conditions.
[0024] Please see Figure 1 and Figure 2 The semantic recognition method includes the following steps: S10, performing Mel spectrum conversion on the audio information to obtain a Mel spectrum feature map, and extracting features from the Mel spectrum feature map to obtain audio features.
[0025] When obtaining raw audio information, it is necessary to process the input raw audio information to obtain a Mel spectrum feature map that reflects the characteristics of human hearing and is closely related to the speech content. The input raw audio information is a single-channel audio signal of fixed duration, such as a raw audio signal of 3 seconds in length with a sampling rate of 44.1kHz or 48kHz.
[0026] First, the raw audio information is preprocessed. Preprocessing operations include noise reduction to reduce the impact of ambient background noise and sampling rate normalization. For example, the sampling rate of the raw audio information is uniformly reduced to 16kHz, which is a commonly used sampling rate in speech processing. This can reduce the amount of data while retaining the main speech information.
[0027] Next, after preprocessing, a Short-Time Fourier Transform (STFT) is performed on the preprocessed raw audio information to transform it from the time domain to the frequency domain. The parameters of the STFT are set as follows: frame length 25ms, frame shift 10ms, and a Hamming window is used to reduce spectral leakage. This transformation process converts continuous raw audio information into a series of overlapping short-time spectral frames.
[0028] Finally, a Mel filter bank is applied to these short-time spectral frames. The Mel filter bank is a set of triangular bandpass filters designed based on the Mel scale, capable of simulating the nonlinear perceptual characteristics of the human ear—high resolution at low frequencies and low resolution at high frequencies. In this embodiment, the number of Mel filter banks is set to 40, meaning that each short-time spectral frame is ultimately mapped to a 40-dimensional feature vector. After performing this operation on all short-time spectral frames, a two-dimensional Mel spectral feature map is obtained. Its time dimension corresponds to the total number of frames after STFT and frame shift, and its frequency dimension is fixed at 40, representing the energy distribution across 40 Mel frequency bands. The Mel spectral feature map is... , For spectrum conversion operation, For the above preprocessing operations, This is the raw audio information before preprocessing.
[0029] After obtaining the Mel spectral feature map, S10 also includes the following step: calculating the frequency domain attention weight using a preset frequency domain attention function and the Mel spectral feature map.
[0030] After obtaining the Mel spectral feature map, in order to selectively enhance the frequency components closely related to the speech content and suppress potential noise or irrelevant information, it is necessary to calculate a set of weights that reflect the importance of different frequency bands, i.e., frequency domain attention weights. This process is accomplished through a pre-defined frequency domain attention function, FreqAtt. The frequency domain attention function is a parametric linear transformation combined with nonlinear activation operations.
[0031] Specifically, the frequency domain attention function takes the entire Mel spectrum feature map as input. First, the function internally defines attention weights and bias parameters, which are learned through a large amount of data during the model training phase. Its goal is to learn which Mel frequency bands in the Mel spectrum feature map have general or context-dependent importance for representing speech content (especially speech content related to lip movements).
[0032] Subsequently, the input Mel spectral feature map is subjected to global average pooling, followed by a linear transformation with the attention weights, and then a bias parameter is added. This linear transformation aims to recombine and evaluate the original Mel spectral feature map. The result of the linear transformation is then fed into the Sigmoid activation function. The Sigmoid function maps the evaluation value corresponding to each Mel frequency band to a scalar between 0 and 1.
[0033] Ultimately, the function outputs a frequency domain attention weight corresponding to the frequency dimension (40 dimensions) of the Mel frequency feature map. The closer the frequency domain attention weight is to 1, the more critical it is in determining the features of the corresponding Mel frequency band; the closer it is to 0, the lower its importance. In this way, the frequency domain attention weight is not fixed but dynamically generated based on the input audio content. The formula for calculating the frequency domain attention function is as follows:
[0034] in, For activation function, This is a global average pooling operation. This is a spectral feature map of Mel. For attention weights, This is the bias parameter.
[0035] After obtaining the Mel spectral feature map, S10 also includes the following steps: element-wise weighting of the Mel spectral feature map based on frequency domain attention weights to obtain a weighted feature map.
[0036] After calculating the dynamic frequency domain attention weights, the original Mel spectrum feature map needs to be finely adjusted using these weights to generate a weighted feature map that enhances key information. This adjustment is done element-wise. Specifically, the calculated frequency domain attention weights are multiplied element-wise by the Mel spectrum feature map. This means that for each time frame and each Mel frequency point in the Mel spectrum feature map, the value is multiplied by the corresponding element in the frequency domain attention weights. If a Mel frequency point is deemed important in the current context, its value will be relatively enhanced across all time frames; conversely, if it is deemed secondary or interfering, its value will be relatively weakened.
[0037] This process is equivalent to applying an adaptive filter to the Mel spectral feature map. It automatically boosts spectral regions related to articulation and phoneme characteristics (especially formant frequencies highly correlated with lip shape) based on the specific speech content, while suppressing spectral regions that may contain background noise or non-critical harmonic components. Therefore, the generated weighted feature map retains the temporal structure and basic spectral shape of the Mel spectral feature map, but the distribution of Mel frequency values within it has been recalibrated based on semantic importance. This allows subsequent models to more easily focus on the most relevant information, thus enabling the extraction of more discriminative audio features.
[0038] After obtaining the Mel spectral feature map, S10 also includes the following steps: performing feature extraction on the weighted feature map to obtain audio features.
[0039] The weighted feature map is input into a pre-trained audio feature extraction network (Light-CNN, a lightweight convolutional neural network) for deep feature extraction and encoding, ultimately yielding the audio features. This audio feature extraction network consists of two convolutional layers, one pooling layer, and one fully connected layer.
[0040] The first convolutional layer performs a sliding convolution operation on the weighted feature map to extract basic preliminary features, such as the coordinated changes of certain frequency band combinations within a short time window, corresponding to phonemes or acoustic events in speech. The second convolutional layer, based on the preliminary features extracted in the first layer, performs deeper combinations and abstractions to capture more complex time-frequency structures. Each convolutional layer is followed by a non-linear activation function and batch normalization to introduce non-linearity and stabilize the training process.
[0041] Subsequently, pooling layers are used to globally average the feature maps output by the convolutional layers across both temporal and frequency spatial dimensions, pooling all activation values from each feature channel into a single scalar. This step transforms the variable-length temporal feature map into a fixed-length feature vector, enabling the processing of input audio of arbitrary duration. Simultaneously, global average pooling integrates information from the entire temporal context and imparts a degree of translation invariance to the model.
[0042] Finally, this fixed-length feature vector is fed into a fully connected layer. The fully connected layer performs a final linear transformation and dimensionality mapping on this vector, projecting it into a predefined high-dimensional semantic space. In this embodiment, the output dimension of the fully connected layer is set to 256 dimensions. Therefore, for the entire input audio segment, the entire audio feature extraction network ultimately outputs a 256-dimensional audio feature. The entire process from weighted feature map to audio feature completes the transformation from a two-dimensional time-frequency representation to a one-dimensional high-level semantic representation. The audio feature is... Audio features The calculation formula is as follows:
[0043] in, For encoding operations, This is a spectral feature map of Mel. For element-wise multiplication, For frequency domain attention function, For the real number field, The temporal position of the audio feature (positively correlated with the duration of the original input audio information, for example, 1 second of original audio information corresponds to 100 temporal positions). The dimension of the audio features (e.g., 256 dimensions).
[0044] Please see Figure 1 and Figure 2 The semantic recognition method further includes the following steps: S20, extracting features from the video information to obtain lip features; the video information and audio information are collected simultaneously, and the video information includes lip information. Step S20 includes the following steps: extracting spatiotemporal features from the video information to obtain basic lip features.
[0045] The acquired video information is a continuous sequence of image frames, where each frame contains an image of the speaker's lips. The video information is... ,in, This refers to the number of image frames (e.g., a 3-second video with a frame rate of 30fps corresponds to 90 image frames). , These are the height and width of the image, respectively. This represents the number of channels.
[0046] To extract effective information representing lip movement and shape from these image frames, spatiotemporal feature extraction is required. For example, a pre-trained video feature extraction network (3D Convolutional Neural Network, 3DCNN) can be used for spatiotemporal feature extraction. Specifically, the 3DCNN is a lightweight C3D network. The input to this 3D convolutional neural network is a sequence of image frames. For each image frame, pre-normalization is performed, uniformly adjusting the size to 112 pixels in height and width, while retaining the RGB channels. Therefore, the video information can be represented as a four-dimensional tensor.
[0047] The advantage of 3D convolutional neural networks lies in the fact that their convolutional kernels slide and compute simultaneously in both spatial (height and width) and temporal (frame sequences) dimensions. Spatial convolutional operations can capture the static contours, shapes, and textures of the lips in a single image frame, such as the boundaries of the upper and lower lips, the curvature of the cupid's bow, and the position of the corners of the mouth. Temporal convolutional operations can capture the dynamic changes in the lip region between consecutive frames, such as the speed of lip opening and closing, the coherence of mouth shape transitions, and subtle vibration patterns during pronunciation. Through the stacking of multiple layers of 3D convolutions, activation functions, and pooling layers, the network can abstract layer by layer from the original pixels, ultimately extracting a high-level feature representation that contains both spatial structure and temporal dynamics, i.e., basic lip features.
[0048] Step S20 also includes the following steps: calculating the lip mask using a preset lip region attention function and basic lip features.
[0049] A lip mask is essentially a weighted map that spatially corresponds to the basic lip features. Its value ranges from 0 to 1, and it indicates the importance of each spatial location in the image belonging to the lip region. The calculation of the lip mask is performed by a pre-defined lip region attention function. This function relies on a pre-trained lip detection model (e.g., the MTCNN-Lip model, pre-trained on the CelebV-Text dataset). This lip detection model is independently trained; its task is to take basic lip features as input and output a lip mask for the lip region.
[0050] Step S20 also includes the following steps: performing element-wise multiplication of the lip mask and the basic lip features to obtain weighted lip features.
[0051] After calculating the lip mask, the lip mask is multiplied element-wise with the basic lip features at the corresponding temporal position, spatial height, spatial width, and feature channel. For each value in the basic lip features, it is multiplied by the weight value of the corresponding spatial position in the lip mask. At this point, in the lip region marked with a high weight (close to 1) in the lip mask, the basic lip features are almost completely preserved; while in the background region marked with a low weight (close to 0) in the lip mask, the basic lip features are significantly attenuated or suppressed.
[0052] Step S20 also includes the following steps: enhancing the positional information of the weighted lip features to obtain lip features.
[0053] For tasks requiring sequence understanding (such as lip reading), the temporal order information in weighted lip features is crucial. Since 3D convolutional neural networks and subsequent weighting operations may weaken the relative temporal position information of weighted lip features to some extent, positional information needs to be injected. This can be achieved by enhancing the positional information of the weighted lip features, for example, using sine and cosine positional encoding. Sine and cosine positional encoding generates a unique encoding vector for each temporal position in the weighted lip features (i.e., the index position corresponding to each frame's features). The dimension of this encoding vector is the same as the feature dimension of the temporal position in the weighted lip features. The generation rule for the encoding vector is: for positions with even-numbered feature dimension indices, a sine function is used to generate the encoding vector; for positions with odd-numbered feature dimension indices, a cosine function is used to generate the encoding vector, thus ensuring differentiated encodings for different positions and different dimensions. Lip Features lip features The calculation formula is as follows:
[0054]
[0055]
[0056] in, For video information, For spatiotemporal feature extraction operations, Let be the attention function for the lip region. For element-wise multiplication, For sine and cosine position encoding, For time-series location index, Indexed by feature dimensions, The number of image frames. Dimensions of lip features.
[0057] Please see Figure 1 and Figure 2 The semantic recognition method also includes the following steps: S30, aligning and concatenating the audio features and lip features to obtain joint features.
[0058] S30 includes the following steps: size alignment of audio features and lip features to obtain size-aligned audio features and lip features.
[0059] After independently extracting audio and lip features, the inconsistency in dimensionality needs to be addressed before proceeding to the cross-modal fusion stage. Audio and lip features originate from different physical signals and processing flows, so their feature dimensions—that is, the length of the feature vector corresponding to each time step and the number of time step sizes—may not be the same. Therefore, they need to be mapped to a unified feature dimension.
[0060] The alignment process can be accomplished by two linear transformation layers, processing audio features and lip features respectively. Essentially, a linear transformation layer is a fully connected layer that linearly projects input features from a high-dimensional space to a new space of a specified dimension. The weight matrices of the two linear transformation layers are initialized using the Xavier method at the start of training to ensure the stability of the training process, and the optimal parameters are learned through backpropagation during subsequent training.
[0061] After linear transformation, audio features and lip features, which may have originally had different dimensions, are mapped to the same preset common feature space, with their dimensions uniformly set to a fixed value, such as 256 dimensions. Furthermore, considering that the temporal steps of audio features and lip features may differ due to different sampling rates (e.g., the sequence length of audio features may not correspond to the frame number of lip features), this embodiment also adjusts the temporal dimension of the lip features using linear interpolation to ensure that its temporal steps are completely consistent with those of the audio features. The size-aligned audio features and lip features are represented as follows: , :
[0062]
[0063] in, , , , , , Here are the weight matrices for the two linear transformation layers. , These are audio features and lip features, respectively. The feature dimension after mapping (e.g., 256). For timing steps, For the dimensions of audio features, Dimensions of lip features.
[0064] S30 also includes the following steps: based on the multi-head attention mechanism, calculate the temporal attention weights between the size-aligned audio features and lip features, and adjust the temporal alignment of the size-aligned lip features according to the temporal attention weights so that the audio features and lip features are temporally aligned.
[0065] Due to the inherent physical delay between sound propagation and the visual representation of lip movements, as well as the slight offsets that may be introduced during feature extraction, audio features and lip features describing the same phoneme may not be precisely aligned on the timeline, leading to potential temporal asynchrony between the size-aligned audio and lip features. To address this issue, a cross-modal temporal alignment mechanism based on Q-Former technology can be used. Q-Former is a lightweight module based on the multi-head attention mechanism in the Transformer architecture. This Transformer architecture can be a Temporal Alignment Network (TAN). Specifically, the size-aligned lip features are used as the query, and the size-aligned audio features are used as both the key and value; the computation process is implemented through a multi-head attention mechanism.
[0066] First, the query, key, and value are segmented into multiple parallel attention heads, each independently calculating attention weights within a different sub-feature space. For each attention head, the dot product similarity between the lip features (query) and the audio features (key) is calculated, and after scaling and softmax normalization, a set of temporal attention weight matrices is obtained. This temporal attention weight matrix quantifies the degree of attention given to all audio features at each lip feature time step.
[0067] Then, for each frame of lip features, the most relevant audio feature is found among all audio features, and the information of that audio feature is fused in. Finally, the outputs of all attention heads are concatenated and passed through a linear projection layer, and then residually connected to the original lip feature input.
[0068] Through the multi-head attention and residual connection operations described above, the output is the temporally aligned lip feature. Simultaneously, the audio feature itself is also processed through a parallel self-attention and residual connection branch to obtain the temporally aligned audio feature. The temporally aligned audio feature and lip feature are represented as follows: , :
[0069]
[0070]
[0071] in, For multi-head attention computation operations (number of attention heads) ), For the feature dimensions of a single attention head, This indicates a residual join operation. , , , These are respectively query, key, value, and transpose. , These are the size-aligned audio features and lip features, respectively.
[0072] S30 also includes the following step: concatenating the temporally aligned audio features with the lip features to obtain joint features. Joint features The calculation formula is as follows:
[0073] in, For feature splicing operations, , These are temporally aligned audio features and lip features, and joint features, respectively. , For timing steps, This represents the feature dimension after mapping.
[0074] Please see Figure 1 and Figure 2 The semantic recognition method further includes the following steps: S40, estimating the signal-to-noise ratio (SNR) of the Mel spectral feature map to obtain the SNR, and obtaining the audio weights and video weights based on the SNR and a preset weight allocation function. The SNR estimation process is completed by a pre-trained SNR estimation model (a lightweight convolutional neural network, such as SNR-Est-LightCNN), which includes three convolutional layers and one pooling layer.
[0075] The first convolutional layer performs a convolution operation using 16 kernels to scan the input Mel-frequency feature map to extract underlying acoustic patterns, such as harmonic structures or noise textures. The output of each kernel is then passed through a ReLU activation function to introduce non-linearity. The kernel size of the first convolutional layer is... With a stride of 1 and padding of 1, this configuration can fully extract local features without reducing the size of the Mel spectral feature map, balancing feature extraction accuracy and computational efficiency. The calculation formula for the convolution operation of the first convolutional layer is as follows:
[0076] in, For activation function, For batch normalization operations, For convolution operations, This is a spectral feature map of Mel. This represents the number of output channels.
[0077] This is followed by a second convolutional layer, whose input is the output of the first convolutional layer. This second convolutional layer uses 8 kernels for deeper feature abstraction, further fusing contextual information to distinguish speech components from background noise. The kernel size of the second convolutional layer is... With a stride of 1 and padding of 1, this configuration can fully extract local features without reducing the feature size, balancing feature extraction accuracy and computational efficiency. The calculation formula for the convolution operation of the second convolutional layer is as follows:
[0078] Next is the convolution operation of the third convolutional layer. This third layer uses a single kernel, and its function is to map the rich features extracted by the first two layers into a single-channel response map. The output of the third convolutional layer is passed through a sigmoid non-linear activation function, compressing the values to a preliminary range between 0 and 1. Each value on this response map can be initially understood as a rough estimate of the signal-to-noise ratio in that local region. The kernel size of the third convolutional layer is... With a stride of 1 and padding of 0, this configuration is primarily used for channel dimension compression and feature fusion. Without altering the temporal dimension of the feature map, it maps multi-channel features to a single channel, providing suitable input for subsequent average pooling operations while reducing computational complexity. The calculation formula for the convolution operation of the third convolutional layer is as follows:
[0079] Finally, the pooling layer applies an adaptive average pooling operation to this single-channel response map. This adaptive average pooling operation globally compresses the response map into a single scalar value in the spatial dimension. This scalar value is the final estimated signal-to-noise ratio. The formula for calculating the average pooling operation of the pooling layer is as follows:
[0080] in, For average pooling operation, This refers to the output size.
[0081] The signal-to-noise ratio (SNR) estimation model is designed to be very lightweight, with a small number of parameters in its three convolutional layers, ensuring that the SNR estimation process does not impose a significant computational burden. Through this approach, the SNR estimation model can transform the Mel spectral feature map into a SNR with a clear physical meaning; for example, it may output a higher SNR, such as 8.5, in a clear speech environment, while outputting a lower SNR, such as 2.3, in a noisy environment.
[0082] After obtaining the signal-to-noise ratio (SNR), the audio and video weights need to be dynamically calculated based on the SNR. Traditional methods using a single linear function to allocate weights are often insufficiently flexible in the low to medium SNR range. Therefore, this embodiment designs a hierarchical weight allocation function. This hierarchical weight allocation function uses different calculation logics to determine the video weights based on the different SNR ranges. The calculation formula for the hierarchical weight allocation function is as follows:
[0083]
[0084] in, For video weight, For audio weights, For interval constraint functions, This refers to the signal-to-noise ratio.
[0085] Specifically, when the signal-to-noise ratio (SNR) is in the low range, such as less than or equal to 3, it indicates that the ambient noise is high and the audio information quality is poor. In this case, more reliance should be placed on relatively reliable video information (lip movements), so the video weight calculation formula is designed to quickly approach a high value, such as 0.8. When the SNR is in the middle range, it indicates that the audio information quality is average, containing both usable information and interference. In this case, the video weight decreases linearly as the SNR increases, reflecting the complementary trend of the two modes. When the SNR is in the high range, such as greater than or equal to 6, it indicates that the audio information is very clear and is the primary information source. In this case, the video weight calculation formula is designed to stabilize at a low value, for example, by limiting its lower limit to 0.2 through an interval constraint function, ensuring that the video information still provides necessary assistance and verification, but the dominant role is given to the audio information.
[0086] This design ensures that the sum of video and audio weights is 1 under all circumstances, conforming to the characteristics of probabilistic weights, and that the allocation of video and audio weights is dynamically driven by the estimated signal-to-noise ratio. Through this adaptive weighting mechanism, the information sources upon which decisions are based can be intelligently adjusted according to the quality of the current audio information, thereby achieving the goal of obtaining high accuracy using high-quality audio information in quiet environments and maintaining stable performance by relying on video information in noisy environments.
[0087] Please see Figure 1 and Figure 2 The semantic recognition method also includes the following steps: S50, using a pre-trained feature fusion model, the joint features are split based on audio weights and video weights to obtain audio branch features and video branch features, and the audio branch features, video branch features, and joint features are fused to obtain fused features.
[0088] Step S50 includes the following steps: S51, using a pre-trained feature fusion model, perform multi-head attention calculation on the joint features based on relative position to obtain attention output features. The feature fusion model (StitchFusion framework) is an architecture specifically designed for deep fusion of multimodal features. Its core is to complete cross-modal information interaction and feature stitching through a modality adapter, and to integrate a multi-head attention mechanism with a feedforward network to achieve fine-grained semantic fusion.
[0089] Joint features are the direct concatenation of audio features and lip features within a unified spatiotemporal framework. Although physical alignment is achieved, the two modalities remain relatively independent, lacking deep semantic interaction and information integration. To fully explore cross-modal associations and generate more discriminative feature representations, attention layers based on multi-head attention mechanisms in pre-trained feature fusion models can be used.
[0090] Multi-head attention mechanisms allow models to simultaneously focus on complex dependencies between cross-modal features from multiple attention heads. Each attention head independently computes dynamic association weights within and between modalities, automatically focusing on the information fragments most important to the current recognition task. The advantage of this is that it allows the model to simultaneously attend to information from different locations and different representation subspaces, thereby capturing richer and more complex dependencies.
[0091] Specifically, no. Each attention point can be based on joint features Generate corresponding query vectors respectively. Key vector Sum value vector The calculation formula is as follows:
[0092]
[0093]
[0094] in, , , For the first The weight matrix of each attention head. , , , The mapped feature dimensions, The feature dimension of a single attention head.
[0095] No. The output of each attention head is The calculation formula is as follows:
[0096] in, For activation function, , , The first The query vector, key vector, and value vector of each attention head.
[0097] Finally, the outputs of all attention heads can be concatenated to obtain the attention output features. The calculation formula is as follows:
[0098] Among them, attention output features , This is a feature splicing operation.
[0099] Step S50 further includes the following step: S52, performing feature decomposition on the attention output features based on audio weights to obtain audio branch features. Specifically, S52 includes the following steps: Calculate the audio modality mask based on audio weights and attention output features; The initial audio branch features are obtained by element-wise multiplication of the audio modality mask and the attention output features; After downsampling the initial audio branch features to reduce the feature dimension, the features are processed by a nonlinear activation function and then upsampled by a linear transformation to restore the feature dimension, thus obtaining the audio branch features.
[0100] Step S50 further includes the following step: S53, performing feature decomposition on the attention output features based on video weights to obtain video branch features. Specifically, S53 includes the following steps: Calculate the video modality mask based on video weights and attention output features; The initial video branch features are obtained by element-wise multiplication of the video modality mask and the attention output features; After downsampling the initial video branch features to reduce the feature dimension, the features are processed by a nonlinear activation function and then upsampled by a linear transformation to restore the feature dimension, thus obtaining the video branch features.
[0101] Audio weights quantify the relative reliability and importance of audio information, while video weights quantify the relative reliability and importance of video information. To apply these audio and video weights to the attention output features, a modality mask with dimensions identical to the attention output features can be set for each branch: the audio branch and the video branch. Thus, an audio modal mask is constructed. With video modal mask The calculation formula is as follows:
[0102]
[0103] in, For video weight, For audio weights, For time-series step index, Indexed by feature dimensions, This is a temperature coefficient (with a value of 0.1, used to adjust the concentration of mask weights). The function is used to normalize the mask (ensuring that the sum of the masks for each time step and each feature dimension is 1). In the attention output features, the index of the time step With feature dimension index The corresponding features are as follows. For the timing step index in the audio modality mask With feature dimension index The corresponding features are as follows. For video modal masking, the index at the timing step With feature dimension index The corresponding features are as follows.
[0104] Next, element-wise multiplication is performed on the audio modality mask and the attention output features to obtain the initial audio branch features. Element-wise multiplication of the video modality mask and attention output features yields the initial video branch features. Element-wise multiplication refers to multiplying the modality mask and the corresponding values of the attention output features one by one. The effect is that the feature vectors at all positions in the attention output features are scaled by the same mask. If the mask is high, most of the information in the attention output features is preserved because the model considers this information reliable; conversely, if the mask is low, the information in the attention output features is significantly attenuated. Initial audio branch features Initial video branch features The calculation formula is as follows:
[0105]
[0106] in, For audio modal mask, For video modal mask, This is the output feature for attention.
[0107] To further refine and enhance the effective information in the initial audio branch features and initial video branch features, and at the same time perform appropriate dimensionality reduction to reduce computational complexity, a series of nonlinear transformations are then applied.
[0108] First, a downsampling linear transformation is performed on the initial audio and video branch features. This transformation projects the high-dimensional initial audio and video branch features into a lower-dimensional subspace using a weight matrix. This downsampling process not only achieves computational lightweighting, but more importantly, it forces the features to be expressed in a compressed representation. This helps the network learn and retain the most core and essential audio-related information while filtering out some secondary or noise-related details. The calculation formula for the downsampling process is as follows:
[0109]
[0110] in, , , The mapped feature dimensions, For the feature dimensions of a single attention head, For the number of attention heads ( ), , , , These are the downsampling weight matrix and bias term for the audio branch, respectively. , These are the downsampling weight matrix and bias term for the video branch, respectively. The initial audio branch features after reducing feature dimensionality, The initial video branch features after reducing the feature dimension.
[0111] The initial audio and video branch features, after dimensionality reduction, are then processed through a non-linear activation function, such as the GELU function. The introduction of this activation function provides the model with non-linear expressive power, enabling the network to learn and construct more complex interactions between elements in the initial audio and video branch features. The relevant calculation formulas are as follows:
[0112]
[0113] in, This is a regularization operation (probability set to 0.1). The initial audio branch features after reducing feature dimensionality, To reduce the feature dimension of the initial video branch features, , These are the initial audio branch features and the initial video branch features after regularization.
[0114] Finally, to match the feature dimensions of subsequent network layers and refine the features, the initial audio and video branch features, after regularization, can be upsampled linearly to restore their feature dimensions to the same level as the attention output features. The relevant calculation formulas are as follows:
[0115]
[0116] in, , , The mapped feature dimensions, , These are the upsampling weight matrix and bias term for the audio branch, respectively. , These are the upsampling weight matrix and bias term for the video branch, respectively. , These are the initial audio branch features and the initial video branch features after regularization, respectively. To recover the audio branch features after feature dimension, To recover the video branch features after the feature dimensions are restored.
[0117] Step S50 also includes the following steps: S54, calculate the modal difference perception coefficient based on audio weights and video weights; perform feature fusion on audio branch features, video branch features, and attention output features according to the modal difference perception coefficient and the preset modal fusion coefficient, and then refine the features to obtain fused features.
[0118] First, based on the previously dynamically calculated audio and video weights, the modal difference perception coefficient is calculated. The modal difference perceptual coefficient is a metric that measures the differences between audio and video branch features. The calculation formula is as follows:
[0119] in, For video weight, For audio weights.
[0120] Subsequently, feature fusion operations can be performed to obtain initial fused features. The calculation formula is as follows:
[0121] in, For layer normalization operation, The preset modal fusion coefficient (e.g., 0.2). To recover the audio branch features after feature dimension, To recover the video branch features after defining the feature dimensions. Modal difference perceptual coefficient. The introduction of this feature can make the modal adapter biased towards higher quality audio or video branch features, enabling more accurate cross-modal stitching.
[0122] Finally, the initial fused features are refined to obtain the final fused features. The initial fused features may still contain some redundancy or noise, and need to be further integrated with their internal information to form a more compact and discriminative representation. Feature refinement is achieved through a feedforward network (FNN), which contains two linear transformation layers and a nonlinear activation function.
[0123] The feedforward network performs a linear transformation on the initial fused features, projecting them into a higher-dimensional space to facilitate the network's learning of more complex feature interactions. Next, a ReLU activation function is applied, enabling the network to capture complex dependencies between features. Then, a linear transformation reduces the feature dimensions to restore or project them onto the target output dimension. After this series of refinements, the intermediate fused features are obtained. The calculation formula is as follows:
[0124] in, , , The mapped feature dimensions, , Here are the weight matrices for the two linear transformation layers. , These are the bias terms for the two linear transformation layers. These are the initial fusion features.
[0125] After obtaining the intermediate fusion features, the intermediate fusion features can be further processed by feature splitting, feature fusion, and feature refinement to obtain the final fusion features. , , For timing steps, This represents the mapped feature dimension. By performing multiple feature splitting, fusion, and refinement processes, the effective fusion of features from the two modalities is ensured. This results in the final fused feature containing both audio speech semantics and video lip movement information, providing accurate support for subsequent semantic recognition. The process of sequentially performing feature splitting, fusion, and refinement on the intermediate fused features is similar to that of sequentially performing feature splitting, fusion, and refinement on the attention output features, and will not be elaborated further here.
[0126] Please see Figure 1 and Figure 2 The semantic recognition method also includes the following steps: S60, performing semantic recognition on the fused features to obtain the recognition result.
[0127] Step S60 includes the following steps: performing linear transformation and normalization on the fused features to generate the character probability distribution corresponding to each time step, and selecting the character with the highest probability value for each time step.
[0128] The fusion feature obtained simultaneously contains speech semantic information from audio and lip movement information from video. Then, this continuous high-dimensional feature vector sequence needs to be mapped into a discrete character sequence, that is, semantic recognition needs to be performed.
[0129] First, the fused features at each time step need to be converted into a character probability distribution. This is achieved by a linear layer, also known as a character mapping network (Char-Map-LL). The output dimension of the character mapping network is the character dictionary dimension. For example, the character dictionary includes 3000 commonly used Chinese characters, 26 English letters, and 10 punctuation marks. By multiplying the fused features by the weight matrix inside the character mapping network and adding a bias term, the 256-dimensional features of the fused features at each time step are linearly transformed into a 3036-dimensional vector. Each dimension of this 3036-dimensional vector corresponds to a specific character in the character dictionary, and its value represents the score that the fused features at the current time step tend to predict for that character.
[0130] However, the numerical ranges of these scores are not fixed, and there is a lack of comparability between different categories, so they cannot be directly used for judgment. In order to convert these raw scores into normalized and interpretable probabilities, it is necessary to perform normalization processing on the transformed 3036-dimensional vector. For example, the Softmax function is used as a normalization tool to ensure that the output value of each element is between 0 and 1. After processing by the Softmax function, the original 3036-dimensional raw score vector is converted into a probability distribution vector, that is, a character probability distribution. In the character probability distribution, each value explicitly represents the probability that the fused feature corresponds to the character at the current time step. For example, at a certain time step, the probability corresponding to the character "I" may be 0.85, the probability corresponding to the character "you" is 0.10, and the sum of the probabilities of other characters is 0.05, which indicates that the model is confident that the pronunciation or lip shape at the current time step corresponds to the character "I". Wherein, the character probability distribution the calculation formula is as follows:
[0131] wherein, , is the dimension of the character dictionary, is the mapped feature dimension, , are the weight matrix and bias term of the character mapping network respectively, is the probability of the -th time step corresponding to the -th character. The higher the probability, the higher the possibility that the time step corresponds to the character, is the final fused feature.
[0132] After the character probability distribution of each time step is generated, it is necessary to select the character with the highest probability in the character probability distribution of each time step as the prediction output of the current time step. In this way, a most probable character is determined for each time step of the fused feature. The -th time step character the calculation formula is as follows:
[0133] wherein, refers to the operation of selecting the character with the highest probability for the -th time step.
[0134] Step S60 further comprises the following step: performing temporal concatenation on all selected characters, and performing deduplication processing to obtain a predicted character sequence.
[0135] Since one pronunciation or one word usually lasts for multiple consecutive time steps, directly concatenating the predicted characters of these time steps in order will lead to the occurrence of a large number of repeated characters. For example, for the pronunciation of the Chinese character "wo (I)", its lip shape and corresponding audio features may show high consistency from the 30th frame to the 40th frame in succession, so "wo (I)" may be predicted as the character with the highest probability at these 11 consecutive time steps. If these 11 "wo (I)" characters are directly concatenated in order, an intermediate result of "wo wo wo wo wo wo wo wo wo wo wo" will be obtained, which obviously does not meet the requirement of the final text output. Therefore, when performing time-series concatenation on all selected characters, deduplication processing needs to be performed synchronously to eliminate redundancy caused by consecutive identical characters.
[0136] The specific operation is to use the lightweight BERT-Tiny model to traverse the obtained character sequence from front to back in the order of time steps. During the traversal process, a character is added to the final sequence only when it is different from its immediately preceding character; if the current character is the same as the preceding character, it is skipped. This process ensures that any two adjacent characters are different in the final output sequence. Through this deduplication operation, the above 11 consecutive "wo (I)" characters will be compressed into a single "wo (I)" character.
[0137] Finally, the predicted character sequence can be obtained after time-series concatenation and deduplication processing. This predicted character sequence already has the preliminary prototype of text, its length is far smaller than the original number of time steps, and adjacent characters are all different, which provides a basic input for subsequent further semantic optimization. For example, the processed predicted character sequence may be "wo jin tian qu gong yuan (I go to the park today)", although there may be individual character errors (for example, "gong (work)" should be "gong (public)"), it is overall a readable text sequence without consecutive repeated characters.
[0138] Step S60 further comprises the following steps: sequentially performing an encoding operation and semantic optimization on the predicted character sequence to obtain encoded features.
[0139] After obtaining the predicted character sequence (e.g., "wo jin tian qu gong yuan (I go to the park today)"), although the structure of the predicted character sequence is close to text, it may still contain problems such as character prediction errors, incoherent semantics or missing words. To solve these problems, a pre-trained large language model (BERT-Tiny) can be introduced to perform correction and optimization at the semantic level.
[0140] First, an encoding operation needs to be performed on the predicted character sequence in character form, converting it into a high-dimensional vector representation that can be processed by the model. Since large language models are trained based on a specific vocabulary and tokenization method, the first step of encoding is tokenization. For example, the Byte Pair Encoding (BPE) method is used to tokenize the predicted character sequence. For Chinese-dominated sequences, BPE can treat each independent Chinese character (such as "我", "今", "天") as a basic token , and can also effectively process English letters and punctuation marks. For example, after the sequence "wo jin tian qu gong yuan" (I went to Gongyuan today) is tokenized by BPE, a token sequence is obtained , in which "我", "今", "天", "去", "工", "园" are each assigned an independent token.
[0141] After obtaining the token sequence, it needs to be converted into a vector rich in semantic information. This process is completed through the word embedding layer and the position encoding layer. The word embedding layer is a pre-trained lookup table, which maps each token to a fixed-dimensional vector through the word embedding operation. The dimension of the word embedding vector of a large language model is 768. Therefore, each token in the sequence (such as "我") will be converted into a 768-dimensional vector, which has learned the rich semantic information of the token during pre-training. However, the word embedding vector itself does not contain the order information of the token in the sequence. In order to provide the model with the temporal structure information of the sequence, additional position encoding is required. Position encoding generates a vector also of 768 dimensions for each position in the sequence (e.g., the first position, the second position). This vector is generated by sine and cosine functions, and has uniqueness and extrapolability. Finally, by adding the word embedding vector of each position and the corresponding position encoding vector element-wise, the initial encoding feature that integrates word meaning and position information is obtained. The initial encoding feature is calculated as follows:
[0142] wherein, is the word embedding operation, is the position encoding operation, is the predicted character sequence.
[0143] Next, semantic optimization needs to be performed on the initial encoding features. The multi-layer self-attention network of a large language model, based on the self-attention mechanism, allows each semantic vector in the initial encoding features (such as "工") to attend to all other semantic vectors in the initial encoding features (including itself, "我", "今 天", "去", "园"), and calculate a weighted context representation.
[0144] Through the self-attention mechanism, the model can deeply understand the global semantics and local dependencies of the entire sequence. For example, when the model processes the semantic vector "工", it will combine the strong semantic implication of the preceding context "我今天去" (expressing an intention to go to a place) and the hint of the following context "园", so as to judge in the context that "工园" is not a reasonable collocation, while "公园" is a correct word that conforms to the daily language habit.
[0145] The multi-layer nonlinear transformation inside the model will refine and correct the input initial encoding features based on this global context understanding, and output an optimized encoding feature with higher semantic consistency. This encoding feature not only corrects possible wrong characters (such as adjusting the semantic vector of "工" to "公"), but also may complete missing semantic components, making the representation of the entire sequence more accurate and fluent.
[0146] Step S60 further comprises the following step: decoding the encoding features into a natural language text, and performing format adjustment to obtain a recognition result.
[0147] For the optimized encoding feature, the 768-dimensional vector at each position (corresponding to one token position in the original sequence) contains the corrected semantic information. The decoding operation can be implemented by a linear output layer. The output layer is a weight matrix, which maps the 768-dimensional feature vector to the entire vocabulary space. In this embodiment, the vocabulary is consistent with the vocabulary used during tokenization. Specifically, for each position vector in the optimized encoding feature sequence, calculation is performed through this linear layer to obtain a logic score vector with the same size as the vocabulary. Subsequently, the Softmax function is used to normalize this score vector into a probability distribution, which represents the probability of occurrence of each token at this position.
[0148] In order to obtain the final text sequence, it is necessary to select the most likely token from the token probability distribution at each position, for example, simply selecting the token with the highest probability at each position as the output. For example, at the position originally corresponding to "工", after semantic optimization, the probability of the token "公" may become the highest, while the probability of "工" decreases significantly. Through decoding, "公" will be selected at this position. Processing all positions in the sequence in sequence, and arranging the selected tokens in order, an optimized token sequence is obtained. Then, according to the previously used BPE tokenization vocabulary, these token identifiers are inversely converted into corresponding characters. For example, the token sequence is converted back to the character sequence "我今天去公园".
[0149] Subsequently, it is necessary to adjust the format of the character sequence to generate a standardized recognition result. Format adjustment includes removing special markers that may be introduced by word segmentation, and adding appropriate punctuation to the text according to the context to enhance readability. For example, if the model recognizes that the sequence expresses a complete declarative sentence, a period may be automatically added at the end. After these processing steps, the original predicted character sequence "wo jin tian qu gong yuan" (the character sequence with a character error: "我今天去工园") is finally optimized and output as a semantically smooth natural language text "我今天去公园。" that conforms to linguistic specifications, thereby completing the entire process from multi-modal fusion features to the final recognition result.
[0150] It can be seen that in the above solution, dynamic signal-to-noise ratio estimation and weight distribution effectively solve the problem that the static and single audio and video fusion strategy in the prior art exists in noisy environments. First, the weight is dynamically calculated according to the quality of real-time audio information: when the signal-to-noise ratio is low, the auxiliary role of video information is strengthened, and when the signal-to-noise ratio is high, audio information is more relied on to dominate, thereby realizing adaptive adjustment of modal contribution. Second, feature alignment and concatenation ensure the spatio-temporal consistency of cross-modal information, laying a foundation for subsequent fusion. Finally, a pre-trained feature fusion model is introduced, which splits and re-fuses the joint features in combination with the weights, so that the bimodal information can be integrated more finely, thereby significantly improving the accuracy of semantic recognition and the robustness of the overall system in complex acoustic environments.
[0151] Please refer to Figure 3 , the present invention also discloses a semantic recognition device fusing audio and video, and the above semantic recognition method can be applied to the semantic recognition device. The semantic recognition device may include an audio feature acquisition module 100, a lip feature acquisition module 200, a feature concatenation module 300, a weight calculation module 400, a feature fusion module 500, and a semantic recognition module 600.
[0152] The audio feature acquisition module 100 is configured to perform Mel spectrum conversion on audio information to obtain a Mel spectrum feature map, and perform feature extraction on the Mel spectrum feature map to obtain audio features.
[0153] The lip feature acquisition module 200 is configured to perform feature extraction on video information to obtain lip features; the video information and the audio information are collected simultaneously, and the video information includes lip information.
[0154] The feature concatenation module 300 is configured to align and concatenate the audio features and the lip features to obtain joint features.
[0155] The weight calculation module 400 is configured to perform signal-to-noise ratio estimation on the Mel spectrum feature map to obtain a signal-to-noise ratio, and obtain audio weights and video weights according to the signal-to-noise ratio and a preset weight distribution function.
[0156] The feature fusion module 500 is used to perform feature splitting on joint features based on audio weights and video weights using a pre-trained feature fusion model to obtain audio branch features and video branch features, and then perform feature fusion on the audio branch features, video branch features, and joint features to obtain fused features.
[0157] The semantic recognition module 600 is used to perform semantic recognition on the fused features and obtain the recognition results.
[0158] For specific limitations regarding the semantic recognition device, please refer to the limitations of the semantic recognition method above, which will not be repeated here. Each module in the aforementioned semantic recognition device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the memory in the electronic device in hardware form, or stored in the memory of the electronic device in software form, so that the memory can call and execute the operations corresponding to each module.
[0159] Please see Figure 4 The electronic device 700 may include a memory 710, a processor 720, and a bus, and may also include a computer program stored in the memory 710 and executable on the processor 720, such as a program for a semantic recognition method that integrates audio and video.
[0160] The memory 710 includes at least one type of readable storage medium, including flash memory, portable hard drive, multimedia card, card-type memory (e.g., SD or DX memory), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 710 can be an internal storage unit of the electronic device 700, such as the portable hard drive of the electronic device 700. In other embodiments, the memory 710 can also be an external storage unit of the electronic device 700, such as a plug-in portable hard drive, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the electronic device 700. Furthermore, the memory 710 can include both internal and external storage units of the electronic device 700. The memory 710 can be used not only to store application software and various types of data installed on the electronic device 700, such as code for semantic recognition methods that integrate audio and video, but also to temporarily store data that has been output or will be output.
[0161] In some embodiments, the processor 720 may be composed of integrated circuits, such as a single packaged integrated circuit or multiple integrated circuits packaged with the same or different functions, including combinations of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 720 is the control unit of the electronic device 700, connecting various components of the electronic device 700 via various interfaces and lines. It executes programs or modules stored in the memory 710 (e.g., programs for semantic recognition methods that integrate audio and video), and calls data stored in the memory 710 to perform various functions of the electronic device 700 and process data.
[0162] The processor 720 executes the operating system of the electronic device 700 and various installed applications. The processor 720 executes the applications to implement the steps in the aforementioned semantic recognition method that integrates audio and video.
[0163] A computer program can be divided into one or more modules. One or more modules are stored in memory 710 and executed by processor 720 to complete this application. The one or more modules can be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of the computer program in electronic device 700. For example, the computer program can be divided into an audio feature acquisition module 100, a lip feature acquisition module 200, a feature splicing module 300, a weight calculation module 400, a feature fusion module 500, a semantic recognition module 600, etc.
[0164] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the invention. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in the present invention should still be covered by the claims of the present invention.
Claims
1. A semantic recognition method integrating audio and video, characterized in that, include: The audio information is subjected to Mel spectrum conversion to obtain a Mel spectrum feature map, and the feature map is extracted to obtain audio features. Feature extraction is performed on the video information to obtain lip features; The video information and the audio information are collected simultaneously, and the video information includes lip information; The audio features and the lip features are aligned and concatenated to obtain a joint feature; The signal-to-noise ratio (SNR) of the Mel spectral feature map is estimated to obtain the SNR, and the audio weight and video weight are obtained based on the SNR and a preset weight allocation function. By using a pre-trained feature fusion model and a multi-head attention mechanism, the joint features are subjected to multi-head attention calculation based on relative position to obtain attention output features; Based on the audio weights, the attention output features are split to obtain audio branch features; Based on the video weights, the attention output features are split to obtain video branch features; Based on the audio weights and the video weights, the modal difference perception coefficient is calculated; Based on the modal difference perception coefficient and the preset modal fusion coefficient, the audio branch features, the video branch features, and the attention output features are fused, and then the features are refined to obtain the fused features. Semantic recognition is performed on the fused features to obtain the recognition results.
2. The semantic recognition method integrating audio and video according to claim 1, characterized in that, The step of extracting audio features from the Mel spectral feature map includes: The frequency domain attention weights are calculated using a preset frequency domain attention function and the Mel spectral feature map. The Mel spectral feature map is element-wise weighted based on the frequency domain attention weights to obtain a weighted feature map. Feature extraction is performed on the weighted feature map to obtain audio features.
3. The semantic recognition method integrating audio and video according to claim 1, characterized in that, The step of extracting features from video information to obtain lip features includes: Spatiotemporal feature extraction is performed on video information to obtain basic lip features; The lip mask is calculated using a preset lip region attention function and the basic lip features; The lip mask is element-wise multiplied with the basic lip features to obtain the weighted lip features; The weighted lip features are augmented with positional information to obtain lip features.
4. The semantic recognition method integrating audio and video according to claim 1, characterized in that, The step of splitting the attention output features based on the audio weights to obtain audio branch features includes: Based on the audio weights and the attention output features, calculate the audio modality mask; The initial audio branch features are obtained by element-wise multiplication of the audio modality mask and the attention output features; After downsampling the initial audio branch features to reduce the feature dimension, the features are processed by a nonlinear activation function and then upsampling the features to restore the feature dimension, thus obtaining the audio branch features.
5. The semantic recognition method integrating audio and video according to claim 1, characterized in that, The step of splitting the attention output features based on the video weights to obtain video branch features includes: Calculate the video modality mask based on the video weights and the attention output features; The initial video branch features are obtained by element-wise multiplication of the video modality mask and the attention output features; After downsampling the initial video branch features to reduce the feature dimension, the features are processed by a nonlinear activation function and then upsampling the features to restore the feature dimension, thus obtaining the video branch features.
6. The semantic recognition method integrating audio and video according to claim 1, characterized in that, The semantic recognition of the fused features to obtain the recognition result includes: The fusion features are subjected to linear transformation and normalization to generate the character probability distribution corresponding to each time step, and the character with the highest probability value is selected step by step. All selected characters are concatenated sequentially and deduplicated to obtain the predicted character sequence; The predicted character sequence is sequentially encoded and semantically optimized to obtain encoded features; The encoded features are decoded into natural language text and the format is adjusted to obtain the recognition result.
7. A semantic recognition device integrating audio and video, characterized in that, The semantic recognition device, which integrates audio and video as described in any one of claims 1 to 6, comprises: The audio feature acquisition module is used to perform Mel spectrum conversion on the audio information to obtain a Mel spectrum feature map, and to extract features from the Mel spectrum feature map to obtain audio features; The lip feature acquisition module is used to extract features from video information to obtain lip features; the video information and the audio information are acquired simultaneously, and the video information includes lip information; The feature splicing module is used to align and splice the audio features and the lip features to obtain joint features; The weight calculation module is used to estimate the signal-to-noise ratio of the Mel spectral feature map, obtain the signal-to-noise ratio, and obtain the audio weight and video weight based on the signal-to-noise ratio and a preset weight allocation function. The feature fusion module is used to perform feature splitting on the joint features based on the audio weights and the video weights using a pre-trained feature fusion model to obtain audio branch features and video branch features, and to perform feature fusion on the audio branch features, the video branch features, and the joint features to obtain fused features; The semantic recognition module is used to perform semantic recognition on the fused features to obtain the recognition result.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes a computer program, it implements the steps of the semantic recognition method that integrates audio and video as described in any one of claims 1 to 6.
9. A computer-readable storage medium storing a computer program, characterized in that, When a computer program is executed by a processor, it implements the steps of the semantic recognition method for fusing audio and video as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Voice interaction method and device based on lip language enhancement, equipment and storage medium
CN120600019A