Audio recognition optimization method and system based on multi-modal feature fusion
The audio recognition method based on multimodal feature fusion solves the problem of insufficient audio recognition accuracy in complex environments. Through audio-video mapping, lip movement quality evaluation and multimodal fusion, the accuracy and robustness of audio recognition are improved.
Patent Information
- Application Number
- CN202510980058.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-16
- Publication Date
- 2025-09-09
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing audio recognition technology lacks accuracy in complex environments, especially in the presence of background noise interference and unclear pronunciation, making it difficult to accurately identify effective information, resulting in insufficient recognition accuracy and robustness.
Through the multimodal feature fusion method, audio and video data are received and audio and video segmentation mapping is performed, an audio-video mapping sequence is constructed, facial confidence, lip area integrity and frame stability detection are performed, a lip movement quality index sequence is generated, an individual lip reference template is constructed and affine transformation is performed, and finally, confidence recognition fusion of audio and video information is performed in the multimodal fusion model.
The accuracy and robustness of audio recognition are improved, and audio information can be better recognized in complex environments.
Smart Images

Figure CN120612941A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of audio recognition technology, and in particular to an audio recognition optimization method and system based on multimodal feature fusion. Background Art
[0002] With the continuous development of audio recognition technology, its application scope has gradually expanded to voice assistants, voice translation, video surveillance, automatic subtitle generation and other fields. However, existing audio recognition technology still faces many challenges in practical application. In complex environments, such as those with background noise interference, changes in the speaker's speech speed, and unclear pronunciation, it is often difficult to accurately recognize effective information based solely on audio signals, resulting in a significant reduction in recognition accuracy. For example, in noisy public places, voice commands may be drowned out by the surrounding noise, making it impossible for voice assistants to correctly understand and execute user commands. In security monitoring scenarios, background noise may mask critical sound information, affecting the timely detection and handling of abnormal events. In addition, a single audio signal cannot provide sufficient contextual information, resulting in insufficient accuracy and robustness of the recognition results. Summary of the Invention
[0003] This application provides an audio recognition optimization method and system based on multimodal feature fusion, which solves the technical problem of insufficient audio recognition accuracy in the prior art.
[0004] In a first aspect of the present application, an audio recognition optimization method based on multimodal feature fusion is provided, the method comprising: Target audio and video data and audio recognition object features are received; the audio recognition object features perform audio and video segmentation mapping on the target audio and video data to construct an audio-video mapping sequence; for the video sequence in the audio-video mapping sequence, facial confidence detection, lip area integrity detection and frame stability detection are performed on each frame image in the video sequence to generate various facial confidence indicators, various lip area integrity indicators and various frame stability indicators; the various facial confidence indicators, various lip area integrity indicators and various frame stability indicators are weightedly averaged and fused to generate a lip movement quality indicator sequence; an individual lip reference template is constructed to extract and affine transform the lip area in the video sequence to generate an affine lip movement image sequence; based on the lip movement quality indicator sequence, the affine lip movement image sequence and the audio sequence in the audio-video mapping sequence are confidence-recognized and fused in a multimodal fusion model to generate a fused audio recognition result.
[0005] A second aspect of the present application provides an audio recognition optimization system based on multimodal feature fusion, the system comprising: A receiving module is used to receive target audio and video data and audio recognition object features; a mapping module is used to perform audio and video segmentation mapping on the target audio and video data using the audio recognition object features to construct an audio-video mapping sequence; a detection module is used to perform facial confidence detection, lip area integrity detection and frame stability detection on each frame image in the video sequence in the audio-video mapping sequence, and generate each facial confidence index, each lip area integrity index and each frame stability index; a weighted fusion module is used to fuse the each facial confidence index, each lip area integrity index and each frame stability index by weighted average to generate a lip movement quality index sequence; an extraction module is used to construct an individual lip reference template, extract and affine transform the lip area in the video sequence, and generate an affine lip movement image sequence; a fusion module is used to perform confidence recognition fusion on the affine lip movement image sequence and the audio sequence in the audio-video mapping sequence in a multimodal fusion model based on the lip movement quality index sequence to generate a fused audio recognition result.
[0006] One or more technical solutions provided in this application have at least the following technical effects or advantages: First, target audio and video data and audio recognition object features are received. The audio recognition object features perform audio and video segmentation mapping on the target audio and video data to construct an audio-video mapping sequence. Then, facial confidence detection, lip region integrity detection, and frame stability detection are performed on each frame image in the video sequence within the audio-video mapping sequence to generate facial confidence indices, lip region integrity indices, and frame stability indices. Furthermore, a weighted average fusion of these facial confidence indices, lip region integrity indices, and frame stability indices is performed to generate a lip movement quality indicator sequence. Next, an individual lip reference template is constructed, and the lip region in the video sequence is extracted and affine transformed to generate an affine lip movement image sequence. Finally, based on the lip movement quality indicator sequence, the affine lip movement image sequence and the audio sequence in the audio-video mapping sequence are fused using confidence recognition in a multimodal fusion model to generate a fused audio recognition result. This method solves the technical problem of insufficient audio recognition accuracy in the prior art and achieves the technical effect of improving audio recognition accuracy through multimodal feature fusion. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0008] Figure 1 A flowchart of an audio recognition optimization method based on multimodal feature fusion provided in an embodiment of the present application; Figure 2 A schematic diagram of the structure of an audio recognition optimization system based on multimodal feature fusion provided in an embodiment of the present application.
[0009] Explanation of reference numerals: receiving module 11 , mapping module 12 , detection module 13 , weighted fusion module 14 , extraction module 15 , fusion module 16 . DETAILED DESCRIPTION
[0010] This application solves the technical problem of insufficient audio recognition accuracy in the prior art by providing an audio recognition optimization method and system based on multimodal feature fusion.
[0011] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only some of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0012] It should be noted that the terms "including" and "having" are intended to cover non-exclusive inclusions. For example, a process, method, system, product or server that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or modules that are not clearly listed or are inherent to these processes, methods, products or devices.
[0013] Example 1, as Figure 1 As shown, the present application provides an audio recognition optimization method based on multimodal feature fusion, wherein the method includes: Receive target audio and video data and audio recognition object features.
[0014] In this embodiment of the present application, the system receives target audio and video data and features of an audio recognition target. Specifically, the target audio and video data includes audio and video streams captured by audio and video acquisition devices (such as cameras and microphones). These data typically contain speech content and lip images related to the audio recognition target. The audio data includes an audio signal stream representing the sound waveform produced by the audio recognition target during vocalization. The video data includes video frames, each of which displays visual information such as the audio recognition target's mouth shape and facial expression.
[0015] After receiving the audio and video data, the system also receives audio recognition object features. Audio recognition object features refer to the features of the audio recognition object (such as sound features and facial features), including but not limited to the speaker's voice spectrum, pitch, speaking speed, rhythm and other time-frequency features in the audio signal, as well as the speaker's facial features (such as lip shape, lip movement, etc.).
[0016] The audio recognition object feature performs audio and video segmentation mapping on the target audio and video data to construct an audio-video mapping sequence.
[0017] The target audio and video data are temporally and spatially aligned based on extracted audio recognition object features (such as the speaker's voice and facial features). Valid speech segments within the audio signal are identified using features within the audio signal (such as the sound spectrum and voice activity detection), specifically identifying audio segments corresponding to specific actions or expressions in the video. Simultaneously, facial features (such as facial expression recognition and lip movement tracking) are used to locate the speaker's facial region within the video frame and determine the time period corresponding to the audio signal. By precisely aligning each audio signal segment with its corresponding lip movement video segment, each audio clip is matched to its corresponding video clip (i.e., the speaker's lip movements), thereby constructing an audio-video mapping sequence. This audio-video mapping sequence accurately associates the audio signal with features such as lip movements and facial expressions in the video content, laying the foundation for subsequent lip movement quality assessment, audio recognition, and fusion processing.
[0018] For the video sequence in the audio-video mapping sequence, facial confidence detection, lip area integrity detection and frame stability detection are performed on each frame image in the video sequence to generate various facial confidence indicators, various lip area integrity indicators and various frame stability indicators; the various facial confidence indicators, various lip area integrity indicators and various frame stability indicators are weightedly averaged and fused to generate a lip movement quality indicator sequence.
[0019] Each frame in a video sequence contains information about the lip movements of the subject being recognized, and these movements match the pronunciation in the audio. These lip movements are then quality-assessed to determine their temporal and spatial consistency, as well as their correspondence with the audio signal. Based on preset evaluation metrics (facial confidence index, lip region integrity index, and frame stability index), lip movement quality is evaluated and a lip movement quality index sequence is generated. This lip movement quality index sequence reflects the overall quality of the lip movements in the video sequence and can effectively assess the degree of match between lip movements and the audio signal, as well as the clarity and stability of the video.
[0020] Specifically, facial confidence detection is first performed on each frame in the video sequence to identify the lip region. Image processing algorithms are then used to locate the face, generating a facial confidence index for each frame. This index measures the clarity and visibility of the lip region in the image. Next, lip region integrity detection is performed on each frame. By analyzing the localization of lip keypoints, any missing or occluded lip regions are detected, generating a lip region integrity index that reflects the integrity and quality of the lip region. Subsequently, frame stability detection is performed to calculate the structural similarity of the lip images between each frame and its neighboring frames to assess the stability of lip movements in the video. Specifically, the structural similarity index (SSIM) is used to calculate the stability of the lip region images between each frame and its neighboring frames, generating a frame stability index that reflects the consistency and stability of lip movements in the video sequence. Finally, the facial confidence index, lip region integrity index, and frame stability index are weighted averaged to obtain a comprehensive quality score for each frame, thereby generating a sequence of lip movement quality indicators.
[0021] For example, assuming the input video sequence contains 10 frames (denoted as F1 to F10), we need to extract the three-dimensional lip movement quality indicators for each frame and output a lip movement quality score (LQS) in the range [0, 1] for subsequent confidence fusion. For each frame, we calculate the following indicators: Clarity score: Laplace transform variance is used, normalized to [0, 1]. The formula is: S = min (1, Laplace transform variance / 500). For example, if the Laplace transform variance of F1 is 420, then S = 0.84.
[0022] Detection confidence: The object detection network outputs a confidence score C∈[0,1], for example, F1 is 0.92.
[0023] Landmark Consistency: If the average Euclidean distance offset of the lip key points in three consecutive frames is less than the threshold δ=1.5, the score is 1; otherwise, it decreases according to the offset ratio, for example, F1 is 0.95.
[0024] Calculate the weighted average of facial confidence index: FCI=0.4S+0.3C+0.3Landmark Consistency, for example, FCI of F1: F1 FCI = 0.4⋅0.84+0.3⋅0.92+0.3⋅0.95=0.336+0.276+0.285=0.897. Lip Region Integrity Index (LRIC): Construct a lip keypoint template K={k1,...,k20}, each point corresponds to an influence weight w k For example, if k5 and k9 are missing, wk5 =0.08, w k9 =0.12, the full score is 1, and the completeness is downgraded to: LRCI=1-(w k5 +w k9 )=1-(0.08+0.12)=0.8.
[0025] Frame Stability Index (FSI): Based on the similarity between each frame and its ±1 frame, the lip region image is evaluated using SSIM or structural similarity, calculating the average structural similarity. For example, if the SSIM of F1, F0, and F2 for the lip region are 0.89 and 0.92, respectively, then the FSI of F1 = (0.89 + 0.92) / 2 = 0.905.
[0026] The weighted fusion generates the lip quality score (LQS): LQS=α⋅FCI+β⋅LRCI+γ⋅FSI, where α=0.4, β=0.3, and γ=0.3, all of which are empirically determined. Substituting the F1 data into the following: LQS=0.4⋅0.897+0.3⋅0.80+0.3⋅0.905=0.3588+0.24+0.2715=0.8703.
[0027] Furthermore, facial confidence detection is performed on each frame image in the video sequence to generate various facial confidence indicators, including: The Laplace transform variance of each frame image is calculated to generate a clarity score; a preset target detection network is called to identify the confidence, detection frame stability, and key point extraction stability of each frame image, and the confidence scores are integrated with the clarity scores to generate the facial confidence indicators.
[0028] For each frame in the video sequence, facial confidence detection is performed and a corresponding facial confidence index is generated. Specifically, the system converts each frame into a grayscale image, ensuring that the processed image data only contains brightness information and no color information. Next, the system performs a Laplace transform on the grayscale image, highlighting the image's edge information by calculating the image's second-order derivative, generating a result that reflects the image's sharpness. The system calculates the variance of the Laplace transformed image. The variance value indicates the degree of fluctuation in the image's pixel values. A higher variance value generally means that the image contains more detail and edge information, resulting in a clearer image. Using the Laplace transform variance calculation method, a clarity score is obtained for each frame.
[0029] The system uses a pre-defined object detection network (such as YOLOv5-face and RetinaFace) to detect facial regions in each frame and generate facial recognition metrics, including confidence, bounding box stability, and keypoint extraction stability. Specifically, confidence indicates the reliability of the detected facial region and is typically determined by the probability value output by the network. A higher confidence indicates a more reliable facial region detection. Frame stability is measured by calculating the intersection over union (IoU) between the current and previous frames. A higher IoU value indicates greater consistency and stability between frames. Keypoint extraction stability assesses the ability to accurately locate facial key points (such as the eyes and mouth) across frames. The accuracy of detecting key areas such as the mouth and eyes, in particular, directly impacts the evaluation of lip movement quality. Finally, the system combines the facial confidence, bounding box stability, keypoint extraction stability, and clarity score to generate the final facial confidence metric through a weighted average. The facial confidence index combines the clarity of the image, the stability of the detection frame, and the accuracy of key point extraction. It can comprehensively reflect the facial recognition quality of each frame of the image and provide a solid data foundation for subsequent lip movement quality evaluation.
[0030] Furthermore, performing lip region integrity detection on each frame image in the video sequence to generate each lip region integrity index includes: A lip key point template is constructed, including the influence weight of each key point on the lip shape recognition accuracy; after extracting the lip area of each frame image, it is compared with the lip key point template to determine the missing key point set; with 1 as the origin, the lip integrity is degraded according to the influence weight of each point in the missing key point set to obtain the integrity index of each lip area.
[0031] The lip keypoint template defines the location of each keypoint in the lip shape and their weighted influence on lip shape recognition accuracy. The position of each keypoint is related to its corresponding lip morphological feature, and these keypoints contribute to lip shape recognition to varying degrees. For example, some keypoints (such as the corners of the mouth and the peak of the lip) have a greater impact on overall lip shape recognition accuracy, while other keypoints (such as the details of the lip edge) have a relatively smaller impact. Based on experimental data and previous lip shape sample sets, the system assigns an influence weight to each keypoint, as shown in Table 1. This weight reflects the contribution of the point to lip shape recognition accuracy.
[0032] Table 1 divides the lip area into multiple key points, each of which has a different influence weight on the lip shape recognition accuracy.
[0033]
[0034] After constructing the lip keypoint template, the system performs lip region extraction on each frame, extracting the lip region from the video frame using image processing algorithms (such as region growing, edge detection, and convolutional neural networks). After lip region extraction, the system compares the extracted lip region with the pre-constructed lip keypoint template. During this comparison, the system determines the integrity of the lip region by checking whether the keypoints in each frame match the pre-defined keypoints in the template. If certain keypoints are missing or incorrectly located in the image, the system generates a missing keypoint set, which contains all keypoints that were not correctly identified or located.
[0035] After identifying the missing key point set, the system downgrades the lip integrity according to the influence weights of these missing key points, and obtains the integrity index of each lip area. Specifically, the system takes 1 as the origin, calculates the influence weight of each point in the missing key point set, and downgrades the integrity of the lip area according to the weight. The missing key points with a larger influence weight will lead to a significant decrease in lip integrity, while the missing key points with a smaller influence weight will lead to a smaller decrease in integrity. The degradation process is weighted according to the degree of influence of the missing key points. The higher the weight of the missing key points, the greater its impact on the integrity. Finally, the lip area integrity index of each frame image is obtained. This index reflects the integrity of the lip area in each frame image. The higher the value, the better the integrity of the lip area, the fewer missing key points, and the higher the image quality.
[0036] For example, in a frame of image, the missing key points detected are the left corner of the mouth (K1) and the right lip peak (K4), and the missing key point set is {K1, K4}.
[0037] According to the weight data, the influence weight of the left corner of the mouth is 0.15, and the influence weight of the right lip peak is 0.12. Therefore, the lip integrity degradation value is: Degradation value = 0.15 (left corner of the mouth weight) + 0.12 (right lip peak weight) = 0.27.
[0038] Assuming that the original lip integrity index is 1 (completely complete), the integrity index of the lip area after degradation is: integrity index = 1-degradation value = 1-0.27 = 0.73, that is, the lip area integrity index of this frame image is 0.73, indicating that there is a certain degree of missing or incompleteness in the lip area.
[0039] Furthermore, the lip key point template is constructed, including: Collect multiple groups of lip shape missing samples, each group of lip shape missing samples includes multiple lip morphology images that are missing the same key point; collect standard complete lip samples, including multiple complete lip morphology images corresponding to the multiple groups of lip shape missing samples; call a pre-trained lip morphology recognition network to recognize the multiple groups of lip shape missing samples and the standard complete lip samples respectively, count the lip shape recognition accuracy reduction index of each group of lip shape missing samples relative to the standard complete lip samples, and generate the influence weight of each key point on the lip shape recognition accuracy; construct the lip key point template based on the influence weight of each key point on the lip shape recognition accuracy.
[0040] First, multiple sets of missing lip shape samples are collected. Each set includes multiple lip morphology images with the same key points missing. For example, images are collected in different lip states (such as closed, open, and smiling) with a specific key point missing (such as the corner of the mouth key point or the upper lip outline point), ensuring that these images cover a variety of lip shape states. Next, standard complete lip samples are collected. This sample set includes complete lip morphology images corresponding to the multiple sets of missing lip shape samples, meaning that each sample image contains all the lip key points, serving as a standard reference. A pre-trained lip shape recognition network is then used to identify the multiple sets of missing lip shape samples and the standard complete lip samples. The pre-trained lip recognition network can identify lip morphology in images and locate key points based on deep learning algorithms such as convolutional neural networks (CNNs). By identifying each group of missing lip samples and standard complete lip samples, the lip recognition accuracy degradation index between each group of missing lip samples and the standard complete samples is calculated. This index reflects the impact of the missing key points on the overall lip recognition accuracy. The greater the accuracy degradation, the greater the impact of the missing key points on lip recognition. Based on the impact of each key point on lip recognition accuracy, the influence weight of each key point on lip recognition accuracy is generated. Finally, based on the influence weight of each key point, a lip key point template is constructed. This template will be used for subsequent lip area integrity detection and lip movement quality assessment.
[0041] Furthermore, the frame stability detection is completed by calculating the lip image structure similarity index between each frame image and a set of neighboring frames that meet a preset number.
[0042] The stability of lip movements can be assessed by analyzing the structural similarity of lip images between adjacent frames. Specifically, the current frame and several neighboring frames (e.g., the five frames preceding and following) are selected for analysis. These neighboring frames must meet a preset number requirement to ensure the representativeness and stability of the selected frame set. Next, a lip image structural similarity index, such as the Structural Similarity Index (SSIM) or other similarity metrics, is used to calculate the structural similarity between the current frame and each neighboring frame. The SSIM compares brightness, contrast, and structural information and quantifies the overall structural similarity between two frames. A higher similarity value indicates that the lip regions between adjacent frames are stable and consistent, and the lip movements are relatively natural; a lower similarity value indicates significant inter-frame discrepancies, suggesting possible instability or abnormality in the lip movements. By calculating the similarity index between the current frame and its neighboring frames, a frame stability index is generated for each frame, reflecting the stability and continuity of the lip movements in the video sequence.
[0043] An individual lip reference template is constructed, and the lip region in the video sequence is extracted and affine transformed to generate an affine lip motion image sequence.
[0044] By collecting a large number of lip image samples and extracting key points (such as lip contours, mouth corners, etc.) from the lip area in each image, an individual lip reference template is established.
[0045] The system extracts the lip region from each frame in the video sequence and then uses individual lip reference templates to align the lip regions in the image, thereby standardizing them and ensuring consistency between lip images across different frames. Next, affine transformation technology is used to transform the lip region in each frame, standardizing the lip shape in the image and generating an affine lip motion image sequence.
[0046] Furthermore, constructing an individual lip reference template, extracting and affine transforming the lip region in the video sequence, and generating an affine lip motion image sequence includes: After extracting the lip area of each frame image in the video sequence, the lip area is aligned with the individual lip reference template to calculate the lip morphology affine transformation matrix; the lip morphology affine transformation matrix is applied to the lip area of each frame image to unify the lip basic structure and generate the affine lip movement image sequence.
[0047] First, the lip region is extracted for each frame in the video sequence. A facial detection algorithm (such as one based on a convolutional neural network or Haar features) is used to locate the lip region in the image and extract image information from this region. After extraction, these lip images are aligned with an individual lip reference template. The individual lip reference template is a standard lip shape template constructed from multiple lip image samples. By extracting and normalizing lip key points, each lip image has a consistent reference frame. Lip key points include the lip contour, mouth corners, upper lip, and lower lip. Next, a lip morphological affine transformation matrix is calculated to normalize the lip region in each frame. By comparing the positions of the lip key points in each frame with those in the individual lip reference template, an affine transformation algorithm (such as a least-squares registration method) is used to calculate the lip morphological affine transformation matrix. The lip morphological affine transformation matrix eliminates positional, rotational, and scaling differences in the image, ensuring that the lip region maintains consistent morphology across frames. The calculated affine transformation matrix for the lip morphology is then applied to the lip region of each frame. This matrix translates, rotates, and scales the lip region, thereby unifying the lip infrastructure. Ultimately, the resulting affine lip movement image sequence provides stable and standardized lip image data, providing reliable data input for subsequent lip movement quality evaluation and audio recognition optimization.
[0048] Furthermore, the steps of constructing the individual lip reference template include: A lip image sample set is collected, and lip key points are extracted for each lip image sample; the lip key points of each lip image sample are normalized, the point-by-point average position of the lip key points is calculated, and the lip individual reference template is constructed.
[0049] First, the system collects a sample set of lip images, consisting of multiple standard lip images. Each image depicts the lip morphology of a different individual, covering a wide range of lip shape variations (e.g., lip morphology across age, gender, and expression). Next, the system extracts lip keypoints from each lip image sample. Using a facial landmark detection algorithm, multiple keypoints of the lip region are extracted from each image, including the lip corners, peaks, upper and lower lip edges, and other detailed locations. After keypoint extraction, the system normalizes the lip keypoints of each lip image sample so that all keypoints in the sample are mapped to a unified reference frame. Typically, normalization is performed based on fixed lip points (such as the corners or center of the lip), using translation, rotation, and scaling operations to ensure consistent, standard positions of keypoints across all image samples. Finally, the system calculates the point-by-point average position of the lip keypoints. This average position is then averaged across all samples to determine the average position of each keypoint. Based on these average positions, an individual lip reference template is constructed.
[0050] Based on the lip movement quality indicator sequence, the affine lip movement image sequence and the audio sequence in the audio-video mapping sequence are confidently recognized and fused in a multimodal fusion model to generate a fused audio recognition result.
[0051] In the multimodal fusion model, the fusion weight is determined based on the lip movement quality indicator sequence, and the affine lip movement image sequence and the audio sequence in the audio-video mapping sequence are fused using weighted averaging to generate the fused audio recognition result.
[0052] Furthermore, based on the lip movement quality indicator sequence, the affine lip movement image sequence and the audio sequence in the audio-video mapping sequence are subjected to confidence recognition fusion in a multimodal fusion model to generate a fused audio recognition result, including: The multimodal fusion model includes a visual channel and an audio channel, the visual channel is constructed based on a 3D convolution kernel, and the audio channel is constructed through a 1D convolution kernel; the audio content of the audio sequence is recognized through the audio channel to generate a first recognition result; the lip shape of the affine lip movement image sequence is recognized through the visual channel to generate a second recognition result; the visual fusion weight is determined using the lip movement quality indicator sequence, and the second recognition result is fused with the first recognition result to generate the fused audio recognition result.
[0053] The multimodal fusion model consists of two channels: a visual channel and an audio channel. The audio channel uses a 1D convolution kernel to process the audio sequence; the audio sequence is extracted through a convolution layer, and then gradually downsampled and feature extracted to obtain a feature vector of the audio content. This feature vector is used to generate the first recognition result, that is, the audio recognition output. The visual channel uses a 3D convolution kernel to process the affine lip movement image sequence. 3D convolution can extract information in both the temporal and spatial dimensions, making it suitable for processing dynamic changes in video data (such as lip movements). The feature vector extracted by 3D convolution is used to generate the second recognition result, that is, the lip shape recognition output based on the lip movement image sequence.
[0054] The audio channel processes the input audio sequence using a 1D convolution kernel, extracting the temporal features of the audio signal and generating an audio recognition result. This result represents the confidence level of the audio content. For example, when recognizing "hello," the audio channel's recognition result might be 0.9 (indicating the confidence level of "hello"). The visual channel processes the affine lip movement image sequence using a 3D convolution kernel, extracting the spatiotemporal features of the lips and generating a lip shape recognition result. For example, when recognizing the lip shape of "hello," the visual channel's recognition result might be 0.85 (indicating the confidence level of "hello"). Next, the fusion weights of the visual and audio channels are determined based on a sequence of lip movement quality indicators. Higher lip movement quality indicates a greater contribution from visual information, while lower lip movement quality indicates a greater reliance on audio information. Finally, the audio and lip shape recognition results are weighted and fused to produce a comprehensive recognition result. The formula is: Fusion Result = Visual Weight × Lip Shape Recognition Result + Audio Weight × Audio Recognition Result. For example, if the audio weight is 0.65 and the visual weight is 0.35, the final fused recognition result is 0.8825, which indicates the confidence level of recognition after multimodal information fusion.
[0055] In an embodiment of the present application, the audio channel uses a 1D convolution kernel to process the input audio sequence to extract the time domain features of the audio signal. Specifically, the input audio signal is preprocessed and converted into frequency domain features, such as Mel-frequency cepstral coefficients (MFCC) or spectrograms. These feature data constitute the input of the audio sequence and are processed by a 1D convolutional neural network. The 1D convolution kernel slides on the time dimension of the audio sequence, gradually performing convolution operations with the audio data of each time step, thereby extracting local features of the audio signal, such as frequency changes, pitch fluctuations, etc. After each convolution operation, the output generated generates a new feature vector, which represents the local features of the current time step, and is nonlinearly transformed through an activation function (such as ReLU) to enhance the learning and processing capabilities of the network. Next, the convolution output is downsampled through a pooling layer to reduce the dimension of the features and retain key information in the audio signal. After convolution, pooling, and activation operations, the resulting feature map is flattened and input into the fully connected layer for further processing. Through the linear transformation and activation function of the fully connected layer, the audio recognition result is finally output. This result is usually expressed as a confidence value, indicating the accuracy and reliability of the model's recognition of the audio content.
[0056] The visual channel uses 3D convolution kernels to process affine lip motion image sequences to extract the spatiotemporal features of the lips. Specifically, the input lip motion image sequence consists of each frame processed with an affine transformation, representing the movement of the lips at different time steps. The 3D convolution kernel simultaneously convolves the video sequence in both the temporal and spatial dimensions, capturing the temporal changes and spatial features of the lip region, thereby identifying the movement patterns and morphological changes of the lips. After each convolution operation, the generated feature map contains the spatiotemporal information of the lip movement, reflecting the morphology and movement trajectory of the lips at different time steps. After processing with 3D convolution, pooling layers (such as max pooling), and activation functions (such as ReLU), the model gradually extracts the deep features of the lip shape and ultimately generates a lip shape recognition result. This recognition result is typically expressed as a confidence value, reflecting the accuracy and reliability of the model's lip shape recognition.
[0057] The quality of lip movements is assessed using the Lip Quality Score (LQS) sequence. These metrics characterize the quality of the video based on information such as facial confidence and lip region integrity for each frame. The higher the lip movement quality, the greater the contribution of visual information to the final fusion result.
[0058] The visual fusion weight is determined based on the lip movement quality index sequence. For example, if the lip movement quality is high (LQS close to 1), the visual channel weight is high; if the lip movement quality is low, the visual channel weight is low.
[0059] For example, the lip quality index sequence (LQS) is [0.9, 0.8, 0.85], and the visual weight = The weights of the visual channels are W1=0.9 / (0.9+0.8+0.85)=0.35, W2=0.8 / (0.9+0.8+0.85)=0.31, and W3=0.85 / (0.9+0.8+0.85)=0.33. These weights reflect the contribution of each frame image to the final fusion audio recognition result.
[0060] In summary, the embodiments of the present application have at least the following technical effects: First, target audio and video data and audio recognition object features are received. The audio recognition object features perform audio and video segmentation mapping on the target audio and video data to construct an audio-video mapping sequence. Then, facial confidence detection, lip region integrity detection, and frame stability detection are performed on each frame image in the video sequence within the audio-video mapping sequence to generate facial confidence indices, lip region integrity indices, and frame stability indices. Furthermore, a weighted average fusion of these facial confidence indices, lip region integrity indices, and frame stability indices is performed to generate a lip movement quality indicator sequence. Next, an individual lip reference template is constructed, and the lip region in the video sequence is extracted and affine transformed to generate an affine lip movement image sequence. Finally, based on the lip movement quality indicator sequence, the affine lip movement image sequence and the audio sequence in the audio-video mapping sequence are fused using confidence recognition in a multimodal fusion model to generate a fused audio recognition result. This method solves the technical problem of insufficient audio recognition accuracy in the prior art and achieves the technical effect of improving audio recognition accuracy through multimodal feature fusion.
[0061] Embodiment 2 is based on the same inventive concept as the audio recognition optimization method based on multimodal feature fusion in the above embodiment. Figure 2 As shown, the present application provides an audio recognition optimization system based on multimodal feature fusion, wherein the system includes: A receiving module 11 is used to receive target audio and video data and audio recognition object features; a mapping module 12 is used to perform audio and video segmentation mapping on the target audio and video data using the audio recognition object features to construct an audio-video mapping sequence; a detection module 13 is used to perform facial confidence detection, lip area integrity detection and frame stability detection on each frame image in the video sequence in the audio-video mapping sequence, and generate each facial confidence index, each lip area integrity index and each frame stability index; a weighted fusion module 14 is used to fuse the each facial confidence index, each lip area integrity index and each frame stability index by weighted average to generate the lip movement quality index sequence; an extraction module 15 is used to construct an individual lip reference template, extract and affine transform the lip area in the video sequence, and generate an affine lip movement image sequence; a fusion module 16 is used to perform confidence recognition fusion on the affine lip movement image sequence and the audio sequence in the audio-video mapping sequence in a multimodal fusion model based on the lip movement quality index sequence to generate a fused audio recognition result.
[0062] Furthermore, the detection module 13 is configured to perform the following method: The Laplace transform variance of each frame image is calculated to generate a clarity score; a preset target detection network is called to identify the confidence, detection frame stability, and key point extraction stability of each frame image, and the confidence scores are integrated with the clarity scores to generate the facial confidence indicators.
[0063] Furthermore, the detection module 13 is configured to perform the following method: A lip key point template is constructed, including the influence weight of each key point on the lip shape recognition accuracy; after extracting the lip area of each frame image, it is compared with the lip key point template to determine the missing key point set; with 1 as the origin, the lip integrity is degraded according to the influence weight of each point in the missing key point set to obtain the integrity index of each lip area.
[0064] Furthermore, the extraction module 15 is configured to perform the following method: Collect multiple groups of lip shape missing samples, each group of lip shape missing samples includes multiple lip morphology images that are missing the same key point; collect standard complete lip samples, including multiple complete lip morphology images corresponding to the multiple groups of lip shape missing samples; call a pre-trained lip morphology recognition network to recognize the multiple groups of lip shape missing samples and the standard complete lip samples respectively, count the lip shape recognition accuracy reduction index of each group of lip shape missing samples relative to the standard complete lip samples, and generate the influence weight of each key point on the lip shape recognition accuracy; construct the lip key point template based on the influence weight of each key point on the lip shape recognition accuracy.
[0065] Furthermore, the detection module 13 is configured to perform the following method: Frame stability detection is completed by calculating the lip image structure similarity index between each frame image and a set of neighboring frames that meet a preset number.
[0066] Furthermore, the extraction module 15 is configured to perform the following method: After extracting the lip area of each frame image in the video sequence, the lip area is aligned with the individual lip reference template to calculate the lip morphology affine transformation matrix; the lip morphology affine transformation matrix is applied to the lip area of each frame image to unify the lip basic structure and generate the affine lip movement image sequence.
[0067] Furthermore, the extraction module 15 is configured to perform the following method: A lip image sample set is collected, and lip key points are extracted for each lip image sample; the lip key points of each lip image sample are normalized, the point-by-point average position of the lip key points is calculated, and the lip individual reference template is constructed.
[0068] Furthermore, the fusion module 16 is configured to perform the following method: The multimodal fusion model includes a visual channel and an audio channel, the visual channel is constructed based on a 3D convolution kernel, and the audio channel is constructed through a 1D convolution kernel; the audio content of the audio sequence is recognized through the audio channel to generate a first recognition result; the lip shape of the affine lip movement image sequence is recognized through the visual channel to generate a second recognition result; the visual fusion weight is determined using the lip movement quality indicator sequence, and the second recognition result is fused with the first recognition result to generate the fused audio recognition result.
[0069] It should be noted that the order in which the embodiments of the present application are presented is for illustrative purposes only and does not necessarily represent the superiority or inferiority of the embodiments. Furthermore, the foregoing descriptions of specific embodiments of this specification are provided. The processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0070] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
[0071] This specification and drawings are merely illustrative of the present application and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Obviously, those skilled in the art may make various modifications and variations to this application without departing from the scope of this application. Thus, this application is intended to include such modifications and variations as fall within the scope of this application and its equivalents.
Claims
1. An audio recognition optimization method based on multimodal feature fusion, characterized in that: Methods include: receiving target audio and video data and audio recognition object features; The audio recognition object feature performs audio and video segmentation mapping on the target audio and video data to construct an audio-video mapping sequence; For the video sequence in the audio-video mapping sequence, performing facial confidence detection, lip region integrity detection, and frame stability detection on each frame image in the video sequence, and generating each facial confidence index, each lip region integrity index, and each frame stability index; A weighted average fusion of the facial confidence indices, the lip region integrity indices, and the frame stability indices is performed to generate a lip movement quality indicator sequence; Constructing an individual lip reference template, extracting and affine transforming the lip region in the video sequence, and generating an affine lip motion image sequence; Based on the lip movement quality indicator sequence, the affine lip movement image sequence and the audio sequence in the audio-video mapping sequence are confidently recognized and fused in a multimodal fusion model to generate a fused audio recognition result.
2. The audio recognition optimization method based on multimodal feature fusion according to claim 1, characterized in that: Performing facial confidence detection on each frame image in the video sequence to generate various facial confidence indicators, including: Calculating the Laplace transform variance of each frame of image to generate a clarity score; A preset target detection network is called to identify the confidence, detection frame stability, and key point extraction stability of each frame image, and is integrated with the clarity score to generate the various facial confidence indicators.
3. The audio recognition optimization method based on multimodal feature fusion according to claim 1, characterized in that: Performing lip region integrity detection on each frame image in the video sequence to generate each lip region integrity index includes: Construct a lip key point template, including the influence weight of each key point on the lip shape recognition accuracy; After extracting the lip region of each frame image, the lip region is compared with the lip key point template to determine the missing key point set; Taking 1 as the origin, the lip integrity is degraded according to the influence weight of each point in the missing key point set to obtain the integrity index of each lip area.
4. The audio recognition optimization method based on multimodal feature fusion according to claim 3, characterized in that: Construct a lip key point template, including: Collect multiple groups of lip shape missing samples, each group of lip shape missing samples includes multiple lip morphology images with the same key point missing; Collecting standard complete lip samples, including complete multiple lip morphology images corresponding to the multiple groups of lip shape missing samples; Invoking a pre-trained lip shape recognition network to recognize the multiple groups of lip shape missing samples and the standard complete lip samples respectively, calculating a lip shape recognition accuracy degradation index of each group of lip shape missing samples relative to the standard complete lip samples, and generating an influence weight of each key point on the lip shape recognition accuracy; The lip key point template is constructed according to the influence weight of each key point on the lip shape recognition accuracy.
5. The audio recognition optimization method based on multimodal feature fusion according to claim 1, characterized in that: Frame stability detection is completed by calculating the lip image structure similarity index between each frame image and a set of neighboring frames that meet a preset number.
6. The audio recognition optimization method based on multimodal feature fusion according to claim 1, characterized in that: Constructing an individual lip reference template, extracting and affine transforming the lip region in the video sequence, and generating an affine lip motion image sequence, including: Extracting the lip region of each frame image in the video sequence and aligning the extracted lip region with the lip individual reference template to calculate the affine transformation matrix of the lip morphology; The lip morphology affine transformation matrix is applied to the lip region of each frame image to unify the lip basic structure and generate the affine lip movement image sequence.
7. The audio recognition optimization method based on multimodal feature fusion according to claim 6, characterized in that: The steps of constructing the individual lip reference template include: Collect lip image sample sets and extract lip key points for each lip image sample; The lip key points of each lip image sample are normalized, the point-by-point average position of the lip key points is calculated, and the lip individual reference template is constructed.
8. The audio recognition optimization method based on multimodal feature fusion according to claim 1, characterized in that: Based on the lip movement quality indicator sequence, the affine lip movement image sequence and the audio sequence in the audio-video mapping sequence are subjected to confidence recognition fusion in a multimodal fusion model to generate a fused audio recognition result, including: The multimodal fusion model includes a visual channel and an audio channel, wherein the visual channel is constructed based on a 3D convolution kernel, and the audio channel is constructed through a 1D convolution kernel; performing audio content recognition on the audio sequence through the audio channel to generate a first recognition result; performing lip shape recognition on the affine lip movement image sequence through the visual channel to generate a second recognition result; The visual fusion weight is determined using the lip movement quality indicator sequence, and the second recognition result is fused with the first recognition result to generate the fused audio recognition result.
9. An audio recognition optimization system based on multimodal feature fusion, characterized in that: A system for implementing the audio recognition optimization method based on multimodal feature fusion according to any one of claims 1 to 8, comprising: A receiving module, configured to receive target audio and video data and audio recognition object features; A mapping module, configured to perform audio-video segmentation mapping on the target audio and video data using the audio recognition object feature, and construct an audio-video mapping sequence; a detection module, configured to perform facial confidence detection, lip region integrity detection, and frame stability detection on each frame image in the video sequence in the audio-video mapping sequence, and generate each facial confidence index, each lip region integrity index, and each frame stability index; A weighted fusion module, configured to fuse the facial confidence indices, the lip region integrity indices, and the frame stability indices by weighted average to generate the lip movement quality indicator sequence; An extraction module is used to construct an individual lip reference template, extract and perform affine transformation on the lip area in the video sequence, and generate an affine lip motion image sequence; A fusion module is used to perform confidence recognition fusion on the affine lip movement image sequence and the audio sequence in the audio-video mapping sequence in a multimodal fusion model based on the lip movement quality indicator sequence to generate a fused audio recognition result.