An audio-video semantic analysis method and system based on multi-modal contrastive learning

By using a multimodal contrastive learning method, audio and video features are extracted, normalized, and weighted to generate a unified semantic representation. This solves the problem of inconsistent emotional expression in AI video generation systems and improves the consistency and quality of generated content.

CN121528248BActive Publication Date: 2026-05-08BEIJING LIUJINSUIYUE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING LIUJINSUIYUE TECH CO LTD
Filing Date
2025-11-17
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing AI video generation systems lack mechanisms for detecting and optimizing the emotional matching between audio and video, resulting in inconsistent emotional expression.

Method used

By employing a multimodal contrastive learning approach, emotional and content features of audio and video are extracted separately. The emotional intensity difference coefficient is calculated, normalized, and then weighted, fused, and jointly encoded to generate a unified semantic representation of audio and video.

Benefits of technology

It enables effective detection and optimization of the emotional matching degree of audio and video, improves the coordination and quality of generated content, and is suitable for post-production quality inspection of digital character generation, AI short video production, virtual anchor system and AI film and television production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121528248B_ABST
    Figure CN121528248B_ABST
Patent Text Reader

Abstract

The application discloses an audio-video semantic analysis method and system based on multi-modal contrast learning, and relates to the technical field of audio-video processing. The method comprises the following steps: acquiring audio-video data and extracting audio signals and video frame sequences; acquiring audio emotional features and audio content features, video emotional features and video content features; calculating a difference coefficient of emotional intensity between the audio emotional features and the video emotional features, and performing intensity normalization processing; weighting and fusing the normalized audio emotional features and the normalized video emotional features to obtain cross-modal emotional features; and jointly encoding the cross-modal emotional features, the audio content features and the video content features to generate unified audio-video semantic representation. The method solves the problem that existing AI video generation systems often directly output audio-video content, lack detection and optimization mechanisms for the emotional matching degree in the generated results, and thus cause the AI video generation system to have the problem of uncoordinated emotional expression.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio and video processing technology, specifically to a multimodal contrastive learning method and system for audio and video semantic analysis. Background Technology

[0002] With the rapid development of artificial intelligence technology, significant progress has been made in the field of audio and video processing. However, traditional post-processing methods mainly focus on technical indicators such as clarity and fluency, lacking the ability to specifically detect and optimize consistency at the emotional semantic level. Furthermore, they mostly focus on extracting content features, while the processing of emotional features is not detailed and precise enough, resulting in an imbalance in the intensity of emotional features across different modalities. Summary of the Invention

[0003] This application provides a multimodal contrastive learning-based audio and video semantic analysis method and system, which solves the technical problem that existing AI video generation systems often directly output audio and video content and lack a mechanism for detecting and optimizing the emotional matching degree of the generated audio and video, resulting in inconsistent emotional expression in AI video generation systems.

[0004] The technical solution to the above-mentioned technical problems in this application is as follows:

[0005] In a first aspect, this application provides a multimodal contrastive learning-based audio-visual semantic analysis method, the method comprising:

[0006] Acquire audio and video data, and extract audio signals and video frame sequences respectively;

[0007] Audio emotion features and audio content features are obtained based on the audio signal, and video emotion features and video content features are obtained based on the video frame sequence;

[0008] Calculate the emotional intensity difference coefficient between the audio emotional features and the video emotional features, and perform intensity normalization processing on the audio emotional features and the video emotional features based on the emotional intensity difference coefficient to obtain normalized audio emotional features and normalized video emotional features;

[0009] The normalized audio sentiment features and the normalized video sentiment features are weighted and fused to obtain cross-modal sentiment features;

[0010] The cross-modal emotion features are jointly encoded with the audio content features and the video content features to generate a unified audio-visual semantic representation.

[0011] Secondly, this application provides a multimodal contrastive learning-based audio and video semantic analysis system, comprising:

[0012] The information acquisition module is used to acquire audio and video data and extract audio signals and video frame sequences respectively;

[0013] The feature acquisition module is used to acquire audio emotion features and audio content features based on the audio signal, and to acquire video emotion features and video content features based on the video frame sequence.

[0014] The coefficient calculation module is used to calculate the emotional intensity difference coefficient between the audio emotional features and the video emotional features, and to perform intensity normalization processing on the audio emotional features and the video emotional features based on the emotional intensity difference coefficient to obtain normalized audio emotional features and normalized video emotional features.

[0015] The feature processing module is used to perform weighted fusion of the normalized audio sentiment features and the normalized video sentiment features to obtain cross-modal sentiment features;

[0016] The semantic generation module is used to jointly encode the cross-modal sentiment features with the audio content features and the video content features to generate a unified audio and video semantic representation.

[0017] This application provides one or more technical solutions, which have at least the following technical effects or advantages:

[0018] This application provides a multimodal contrastive learning-based audio and video semantic analysis method and system. By acquiring audio and video data and extracting audio and video features respectively, the system performs intensity normalization and weighted fusion of emotional features, and finally performs adaptive fusion based on scene type to generate an optimized unified semantic representation. This method can effectively detect and optimize the emotional matching degree of audio and video, and solves the problem of inconsistent emotional expression in existing AI video generation systems.

[0019] The aforementioned technical solutions not only enable precise processing and analysis of audio and video data but also provide innovative solutions for multiple fields. In digital character generation systems, they can detect whether the generated character's facial expressions match the emotional nuances of their voice, providing optimization suggestions for expression adjustments. In AI short video production platforms, they can identify inconsistencies between visual content and voice-over emotion, assisting creators in making fine-tuning adjustments. In virtual anchor systems, they can assess the emotional consistency of generated content in real time, ensuring that the output audio and video content is natural and harmonious. Furthermore, they can be applied to post-production quality control in AI film and television production, automatically identifying and marking emotionally mismatched segments requiring manual intervention, improving overall production efficiency and content quality. Attached Figure Description

[0020] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a flowchart illustrating a multimodal contrastive learning-based audio and video semantic analysis method provided in an embodiment of this application.

[0022] Figure 2 This is a schematic diagram of the structure of a multimodal contrastive learning audio and video semantic analysis system provided in an embodiment of this application.

[0023] The components represented by each number in the attached diagram are explained below:

[0024] Information acquisition module 11, feature acquisition module 12, coefficient calculation module 13, feature processing module 14, semantic generation module 15. Detailed Implementation

[0025] This application provides a multimodal contrastive learning-based audio and video semantic analysis method and system to address the technical problem that existing AI video generation systems often directly output audio and video content, lacking a detection and optimization mechanism for the emotional matching degree of the generated audio and video, resulting in inconsistent emotional expression in AI video generation systems.

[0026] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0027] In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the stated features. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.

[0028] In the description of this application, the term "for example" is used to mean "used as an example, illustration, or description." Any embodiment described as "for example" in this application is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use this application. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that this application can be made without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid unnecessarily obscuring the description of this application. Therefore, this application is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed in this application.

[0029] Example 1, as Figure 1 As shown in the embodiments of this application, a multimodal contrastive learning-based audio and video semantic analysis method is provided, including:

[0030] S10: Acquire audio and video data, and extract audio signals and video frame sequences respectively;

[0031] In this embodiment, audio and video data are acquired through audio and video acquisition devices, such as high-definition cameras and professional recording equipment. After obtaining the audio and video data, existing audio and video separation techniques are used to extract the audio signals and video frame sequences separately. For example, FFmpeg, an open-source audio and video processing tool, can efficiently and accurately separate audio and video. The separated audio signals and video frame sequences will serve as the basis for subsequent feature extraction and analysis.

[0032] Audio signals include information such as the frequency, amplitude, and timbre of sound. A video frame sequence consists of a series of consecutive video frames, each containing visual information such as color, texture, and object shape.

[0033] S20: Obtain audio emotional features and audio content features based on the audio signal, and obtain video emotional features and video content features based on the video frame sequence;

[0034] In this embodiment, audio emotion features and audio content features, as well as video emotion features and video content features, are obtained based on the audio signal and video frame sequence, respectively. For the audio signal, a combination of convolutional neural networks and long short-term memory networks from deep learning is used for feature extraction. Convolutional neural networks can effectively capture local features in the audio signal, while long short-term memory networks can process the temporal information of the audio signal. By combining the two networks, audio emotion features and audio content features are accurately extracted. Audio emotion features reflect the emotional information conveyed by the audio, such as joy, sadness, and anger; audio content features contain the semantic content in the audio, such as the textual information of the speech.

[0035] For video frame sequences, pre-trained visual models, such as ResNet and VGG, are used to extract features from the video frames. By inputting the video frames into a pre-trained model trained on large-scale image data, video sentiment features and video content features are extracted. Video sentiment features reflect the emotional atmosphere conveyed by the video images, while video content features describe the specific content in the video images, such as character actions and scenes.

[0036] Specifically, step S20 in the method includes:

[0037] The audio signal is processed by an audio decoupling network to obtain the audio emotional features and the audio content features, wherein the audio decoupling network includes an audio emotional feature extraction branch and an audio content feature extraction branch;

[0038] The video frame sequence is processed by a video decoupling network to obtain the video emotion features and the video content features. The video decoupling network includes a video emotion feature extraction branch and a video content feature extraction branch.

[0039] In this embodiment, a decoupled network is used to process audio signals and video frame sequences, separating audio emotional features from audio content features. For the audio decoupled network, the audio emotional feature extraction branch focuses on capturing emotion-related information in the audio, such as changes in tone and speed of speech. The audio content feature extraction branch, on the other hand, focuses on extracting the semantic content of the audio, converting it into understandable textual information through speech recognition and analysis.

[0040] Similarly, a video decoupling network is used to process video frame sequences to obtain video sentiment features and video content features. The video decoupling network includes a video sentiment feature extraction branch and a video content feature extraction branch. The video sentiment feature extraction branch analyzes elements such as color distribution, lighting effects, and facial expressions in the video frames to determine the emotional atmosphere conveyed by the video. For example, bright colors and cheerful facial expressions may indicate joy, while dark colors and sad facial expressions may convey sadness. The video content feature extraction branch identifies specific objects, human actions, and scenes in the video frame, providing a foundation for subsequent semantic analysis.

[0041] The construction steps of the audio decoupling network include:

[0042] A sample audio signal set is constructed by retrieving multiple sample audio signals using big data.

[0043] Audio emotion feature labels are annotated for each of the sample audio signals in the sample audio signal set to construct a sample audio emotion feature label set;

[0044] Audio content feature labels are annotated for each of the sample audio signals in the sample audio signal set to construct a sample audio content feature label set;

[0045] Using the sample audio signal set as input and the sample audio emotion feature label set as supervision label, the audio emotion feature extraction branch is constructed.

[0046] Using the sample audio signal set as input and the sample audio content feature label set as supervision label, construct the audio content feature extraction branch;

[0047] The audio decoupling network is obtained by integrating the inputs of the audio emotion feature extraction branch and the audio content feature extraction branch.

[0048] In this embodiment, firstly, a large number of sample audio signals are collected through big data retrieval to form a sample audio signal set. The sample audio signal set covers audio from different scenarios, with different emotional expressions and different semantic content, in order to ensure the generalization ability of the network.

[0049] Next, each sample audio signal in the sample audio signal set is labeled. Audio emotional features and audio content features are labeled to form a sample audio emotional feature label set and a sample audio content feature label set. For example, the sample audio emotional feature label set is labeled with emotion categories such as joy, sadness, and anger; the sample audio content feature label set is labeled with specific text information, etc.

[0050] Secondly, when constructing the audio sentiment feature extraction branch, the sample audio signal set is used as input, and the sample audio sentiment feature label set is used as supervision labels. By combining convolutional neural networks and long short-term memory networks, the network parameters are continuously adjusted to enable the network to accurately extract sentiment features from the audio signal. Similarly, when constructing the audio content feature extraction branch, the sample audio signal set is used as input, and the sample audio content feature label set is used as supervision labels to train the network to extract semantic content from the audio.

[0051] Finally, the inputs of the audio sentiment feature extraction branch and the audio content feature extraction branch are integrated, allowing the two branches to share the input audio signal, thus obtaining a complete audio decoupling network. The audio decoupling network can effectively separate the sentiment and content features from the audio signal, providing accurate feature data for subsequent audio-visual semantic analysis.

[0052] For example, an audio decoupling network is constructed and trained based on a convolutional neural network and a long short-term memory network. The specific steps are as follows:

[0053] First, data preparation involves collecting sample audio signals and dividing them into training, validation, and test sets. Audio sentiment feature labels and audio content feature labels are then applied to the sample audio signals in the training and validation sets to form corresponding label sets. These labels are then divided into training, validation, and test sets in a 7:1.5:1.5 ratio.

[0054] Next, the model is constructed, including an audio sentiment feature extraction branch and an audio content feature extraction branch. The audio sentiment feature extraction branch adopts an architecture combining convolutional neural networks and long short-term memory networks. Convolutional layers are used to capture local features of the audio signal, while long short-term memory layers are used to process the temporal information of the audio signal. The audio content feature extraction branch also adopts a similar architecture to extract semantic content from the audio. The number of nodes in the input layer is equal to the dimension of the input features. For example, if the sample audio signal set has 5 features, then the input layer contains 5 nodes. 1-3 hidden layers are set, and the number of nodes in each layer is adjusted experimentally, such as 64, 32, etc. The activation function is ReLU. The number of nodes in the output layer is equal to the predicted audio sentiment features and audio content features. For example, the prediction takes 1 node. The output layer generally does not use an activation function and directly outputs continuous values.

[0055] Next, model training involves inputting the sample audio signal set into the constructed branch network, using the sample audio sentiment feature label set and the sample audio content feature label set as supervision labels for training. The training framework is constructed using the Adam optimizer and the mean squared error (MSE) loss function, with a batch size of 32 and a total of 50 training epochs. An early stopping mechanism (patience=5) is introduced: if the validation set loss does not decrease for 5 consecutive epochs, the training process is automatically terminated, resulting in the trained audio sentiment feature extraction branch and audio content feature extraction branch. This effectively avoids overfitting while ensuring the model reaches convergence.

[0056] Finally, the branches are integrated, combining the pre-trained audio emotion feature extraction branch and the audio content feature extraction branch to obtain the audio decoupled network. Simultaneously, the model is validated using a validation set to prevent overfitting.

[0057] Furthermore, the construction steps of the video decoupling network include:

[0058] A sample video frame sequence set is constructed by retrieving multiple sample video frame sequences using big data.

[0059] Video sentiment feature labels are annotated for each of the sample video frame sequences in the sample video frame sequence set to construct a sample video sentiment feature label set;

[0060] Video content feature labels are annotated for each of the sample video frame sequences in the sample video frame sequence set to construct a sample video content feature label set;

[0061] Using the sample video frame sequence set as input and the sample video sentiment feature label set as supervision label, the video sentiment feature extraction branch is constructed.

[0062] Using the sample video frame sequence set as input and the sample video content feature tag set as supervision tags, the video content feature extraction branch is constructed.

[0063] The inputs of the video emotion feature extraction branch and the video content feature extraction branch are integrated to obtain the video decoupling network.

[0064] In this embodiment, the construction of the video decoupling network also employs a method similar to that used in training the audio decoupling network. First, multiple sample video frame sequences are collected through big data retrieval to construct a sample video frame sequence set. Then, video sentiment feature labels are applied to each sample video frame sequence in the sample video frame sequence set to construct a sample video sentiment feature label set, such as labels indicating cheerful, melancholic, or tense emotional atmospheres. Simultaneously, video content feature labels are applied to construct a sample video content feature label set, including specific character actions, scenes, and other content.

[0065] Secondly, using a set of sample video frame sequences as input and a set of sample video sentiment feature labels as supervision labels, a video sentiment feature extraction branch is constructed. Utilizing a pre-trained visual model, such as ResNet, combined with deep learning algorithms, the network is enabled to extract sentiment features from the video frame sequences. Then, using a set of sample video frame sequences as input and a set of sample video content feature labels as supervision labels, a video content feature extraction branch is constructed to train the network to recognize specific content within the video frames.

[0066] Finally, the inputs of the video sentiment feature extraction branch and the video content feature extraction branch are integrated to form a video decoupling network. This video decoupling network can effectively separate sentiment features and content features in a video frame sequence, providing an important feature foundation for audio and video semantic analysis.

[0067] By constructing audio decoupling networks and video decoupling networks, we can more accurately obtain the emotional and content features of audio and video, thereby improving the accuracy and reliability of audio and video semantic analysis.

[0068] For example, a video decoupling network is constructed and trained based on a pre-trained visual model ResNet, and the specific steps are as follows:

[0069] First, data acquisition and preprocessing are performed. Multiple sample video frame sequences are collected and processed into a format suitable for model input, such as standardizing video frame size and resolution. The sample video frame sequence set is divided into a training set, a validation set, and a test set in a 6:2:2 ratio. Video sentiment feature labels and video content feature labels are then applied to the sample video frame sequences in the training and validation sets to form corresponding label sets.

[0070] Next, the model is built, constructing a video sentiment feature extraction branch and a video content feature extraction branch. The video sentiment feature extraction branch is based on a pre-trained ResNet model, with fully connected layers and a Softmax activation function added to output the sentiment feature classification results. The video content feature extraction branch also utilizes a ResNet model, with convolutional layers, pooling layers, and fully connected layers added to extract key features of the video content. The number of nodes in the input layer is determined based on the feature dimension of the input video frame. For example, if the input video frame is a 3-channel image, the number of nodes in the input layer is related to the 3-channel feature dimension. Two to four hidden layers are set, and the number of nodes in each layer is determined by cross-validation, such as 128 or 64. The LeakyReLU activation function is used to avoid the gradient vanishing problem. The number of nodes in the output layer is determined based on the number of categories of the predicted video sentiment features and video content features. The output layer uses the Sigmoid activation function to output probability values.

[0071] Next, the model is trained by inputting the sample video frame sequence set into the constructed branch network, using the sample video sentiment feature label set and the sample video content feature label set as supervision labels, respectively. The training framework is constructed using the Adagrad optimizer and cross-entropy loss function, with a batch size of 16 and a total of 80 training epochs. A learning rate decay mechanism is introduced, reducing the learning rate to 0.8 times its original value every 10 epochs. A Dropout mechanism with a Dropout rate of 0.3 is also introduced to prevent overfitting. Training is terminated early when the accuracy on the validation set fails to improve for 10 consecutive epochs, resulting in the trained video sentiment feature extraction branch and video content feature extraction branch.

[0072] Finally, branch fusion is performed, integrating the inputs of the trained video sentiment feature extraction branch and video content feature extraction branch to obtain the integrated video decoupling network. The video decoupling network is then comprehensively tested using a test set to evaluate its performance in different scenarios. Based on the test results, the model is fine-tuned to further improve its accuracy and generalization ability.

[0073] S30: Calculate the emotional intensity difference coefficient between the audio emotional features and the video emotional features, and perform intensity normalization processing on the audio emotional features and the video emotional features according to the emotional intensity difference coefficient to obtain normalized audio emotional features and normalized video emotional features;

[0074] In this embodiment, to ensure effective comparison and analysis of audio and video emotional features on the same scale, an emotional intensity difference coefficient is calculated between the two. The emotional intensity difference coefficient reflects the degree of difference in emotional expression intensity between audio and video.

[0075] First, the difference between audio and video emotional features across various dimensions is calculated and then weighted and summed to obtain the emotional intensity difference coefficient. For example, if a certain dimension of the audio emotional features represents the intensity of anger, and the video emotional features also have a corresponding intensity value in that dimension, the absolute value of the difference between the two is calculated, and then appropriate weights are assigned according to the importance of that dimension. The results of all dimensions are then summed to obtain the emotional intensity difference coefficient.

[0076] Secondly, based on the calculated emotional intensity difference coefficient, the emotional features of both audio and video are normalized. This adjusts the emotional intensity of both to the same range, avoiding the impact of excessive intensity differences on subsequent analysis results. For audio emotional features, the value of each dimension is divided by the emotional intensity difference coefficient to obtain normalized audio emotional features; a similar operation is performed on video emotional features to obtain normalized video emotional features.

[0077] Specifically, step S30 in the method includes:

[0078] Calculate the feature intensity values ​​of the audio emotion features and the video emotion features respectively to obtain the audio emotion intensity value and the video emotion intensity value;

[0079] The emotional intensity difference coefficient is calculated based on the audio emotional intensity value and the video emotional intensity value;

[0080] The intensity of the audio emotional features is adjusted according to the emotional intensity difference coefficient to obtain the normalized audio emotional features, and the intensity of the video emotional features is adjusted according to the emotional intensity difference coefficient to obtain the normalized video emotional features.

[0081] In this embodiment, the feature intensity values ​​of audio emotion features and video emotion features are first calculated separately. Vector norm is an indicator of vector strength. For audio emotion features, vector norm is calculated to obtain the audio emotion intensity value; for video emotion features, a similar calculation method is used to obtain the video emotion intensity value.

[0082] For example, the audio emotion feature is a 4-dimensional vector [0.8, 0.6, 0.3, 0.9]. Calculate the L2 norm: 1.3 represents the audio emotional intensity value. The larger the vector value and the larger the norm, the stronger the emotional intensity.

[0083] Next, based on the audio and video emotional intensity values, an emotional intensity difference coefficient is calculated. For example, the relative ratio of the two is used as the emotional intensity difference coefficient.

[0084] Secondly, after obtaining the emotional intensity difference coefficient, the audio and video emotional features are adjusted for intensity. For audio emotional features, the feature value of each dimension is multiplied by an adjustment coefficient calculated based on the emotional intensity difference coefficient. For example, the ratio of the audio emotional intensity value to the emotional intensity difference coefficient is used as the adjustment coefficient, and the values ​​of each dimension of the audio emotional feature are multiplied to obtain the normalized audio emotional feature. For video emotional features, a similar method is used, determining the adjustment coefficient based on the relationship between the video emotional intensity value and the emotional intensity difference coefficient, and adjusting the values ​​of each dimension of the video emotional feature to obtain the normalized video emotional feature.

[0085] By normalizing the data, the emotional features of audio and video can be compared and analyzed on a unified scale, providing a data foundation for audio and video semantic analysis.

[0086] The emotional intensity difference coefficient is calculated based on the audio emotional intensity value and the video emotional intensity value, including:

[0087] Based on the audio emotional intensity value, calculate the relative ratio of the video emotional intensity value to the audio emotional intensity value;

[0088] The relative ratio is used as the emotional intensity difference coefficient.

[0089] In this embodiment, the relative ratio is calculated based on the audio emotional intensity value. Assuming the audio emotional intensity value is A and the video emotional intensity value is V, then the relative ratio, i.e., the emotional intensity difference coefficient K = V / A, is calculated.

[0090] For example, the audio sentiment features are [0.8, 0.6, 0.3, 0.9], and the L2 norm is calculated to be 1.3, so the audio sentiment intensity value is 1.3; the video sentiment features are [1.2, 0.8, 0.5, 1.0], and the L2 norm is calculated to be approximately 1.82, so the video sentiment intensity value is 1.82.

[0091] Coefficient of difference This study uses a relative ratio as the emotional intensity difference coefficient to avoid the limitations of simple difference calculations. A difference coefficient greater than 1 indicates that the video's emotional intensity is stronger than the audio's; a difference coefficient less than 1 indicates that the video's emotional intensity is weaker than the audio's; and a difference coefficient equal to 1 means that both have the same emotional intensity. Since difference calculations may not accurately reflect the relative magnitude of the emotional intensity, the relative ratio can uniformly measure the intensity difference under different combinations of audio and video emotional intensities.

[0092] Further, the intensity of the audio emotional features is adjusted according to the emotional intensity difference coefficient to obtain the normalized audio emotional features, and the intensity of the video emotional features is adjusted according to the emotional intensity difference coefficient to obtain the normalized video emotional features, including:

[0093] Based on the square root value of the emotional intensity difference coefficient, the audio emotional features are multiplied to obtain the normalized audio emotional features.

[0094] Based on the square root value of the emotional intensity difference coefficient, the video emotional features are divided to obtain the normalized video emotional features.

[0095] In this embodiment, the square root value of the emotional intensity difference coefficient is used to adjust the emotional features of audio and video, avoiding over- or under-adjustment due to excessively large or small difference coefficients. For audio emotional features, the value of each dimension is multiplied by the square root value of the emotional intensity difference coefficient.

[0096] For example, if the emotional intensity difference coefficient K is 1.4, its square root is approximately 1.183, and the audio emotional features are [0.8, 0.6, 0.3, 0.9], then the normalized audio emotional features are [0.8×1.183, 0.6×1.183, 0.3×1.183, 0.9×1.183], that is, [0.947, 0.710, 0.355, 1.065], and the new intensity value is approximately 1.54.

[0097] For video sentiment features, the value of each dimension is divided by the square root of the sentiment intensity difference coefficient. Taking a sentiment intensity difference coefficient K of 1.4 (square root approximately 1.183) and video sentiment features of [1.2, 0.8, 0.5, 1.0] as an example, the normalized video sentiment features are [1.2÷1.183, 0.8÷1.183, 0.5÷1.183, 1.0÷1.183], approximately [1.014, 0.676, 0.423, 0.845], with a new intensity value ≈ 1.54. By adjusting the intensity of both modalities to 1.54, intensity balance is achieved.

[0098] After intensity adjustment of the audio and video sentiment features, the resulting normalized audio sentiment features and normalized video sentiment features can be compared and analyzed more accurately and effectively at the same scale.

[0099] S40: The normalized audio sentiment features and the normalized video sentiment features are weighted and fused to obtain cross-modal sentiment features;

[0100] In this embodiment, emotional information from both audio and video modalities is comprehensively utilized to perform weighted fusion of normalized audio and video emotional features. First, the weights of audio and video are determined based on their respective contributions to emotional expression in different scenarios. For example, in music videos, audio may be more crucial for emotional expression, thus the weight of audio is set higher; while in silent theatrical performance videos, the visuals and actors' body language are more important for conveying emotion, so the weight of video can be increased accordingly.

[0101] After determining the weights, the value of each dimension of the normalized audio sentiment feature is multiplied by the audio weight, and the value of each dimension of the normalized video sentiment feature is multiplied by the video weight. Then, the values ​​of the corresponding dimensions are added together to obtain the cross-modal sentiment features.

[0102] Specifically, step S40 in the method includes:

[0103] Identify the current scene type based on the audio and video data;

[0104] Obtain the corresponding preset audio fusion weights and preset video fusion weights based on the current scene type;

[0105] Based on the preset audio fusion weights and the preset video fusion weights, the normalized audio sentiment features and the normalized video sentiment features are weighted and combined to obtain the cross-modal sentiment features.

[0106] In this embodiment, the current scene type is first identified based on audio and video data. A scene classification model can be used, which is trained on audio and video data from different scenes and can accurately identify scene types such as music performances, theatrical performances, news broadcasts, and sports events. Although the intensity is the same, the direction and numerical distribution of the feature vectors are completely different. For example, audio features may express sadness in a trembling voice, while video features may express sadness in facial expressions, so a weighted fusion approach is used.

[0107] For example, the weight format is described as [audio weight, video weight], and in a dialogue scenario it is represented as [0.5, 0.5], since voice information is just as important as facial expressions and lip movements; while in a silent scenario, the weight is [0.2, 0.8], because there is only background noise or no sound, so the audio weight is 0.2, and the video weight is 0.8 because in a silent scenario, the main way to judge emotions is by visual information.

[0108] Next, based on the identified current scene type, the corresponding preset audio fusion weights and preset video fusion weights are obtained. The importance of audio and video in emotional expression varies depending on the scene type.

[0109] Finally, based on preset audio fusion weights and preset video fusion weights, the normalized audio sentiment features and normalized video sentiment features are weighted and combined. This process is repeated to obtain complete cross-modal sentiment features. By using weighted fusion, the sentiment information from both audio and video modalities is fully utilized, laying the foundation for more accurate audio-visual semantic analysis in the future.

[0110] For example, if the weight of audio is 0.6 and the weight of video is 0.4, the normalized audio sentiment feature is [0.947, 0.710, 0.355, 1.065], and the normalized video sentiment feature is [1.014, 0.676, 0.423, 0.845], then the value of the first dimension of the cross-modal sentiment feature is 0.947 × 0.6 + 1.014×0.4=0.974; the value of the second dimension is 6×0.710+0.4×0.676=0.426+0.270=0.696; the value of the third dimension is 0.6×0.355+0.4×0.423=0.213+0.169=0.382; the value of the fourth dimension is 0.6×1.065+0.4×0.845=0.639+0.338=0.977, and the final cross-modal sentiment feature = [0.974,0.696,0.382,0.977].

[0111] S50: Jointly encode the cross-modal emotion features with the audio content features and the video content features to generate a unified audio-visual semantic representation.

[0112] In this embodiment, cross-modal sentiment features are effectively fused with audio content features and video content features to generate a unified audio-visual semantic representation, employing a joint coding approach. Joint coding can fully explore the intrinsic connections between the three types of features, integrating information from different modalities into a unified semantic space.

[0113] First, feature preprocessing is performed on cross-modal sentiment features, audio content features, and video content features, including normalization to ensure consistent numerical ranges and prevent differences in feature scale from affecting coding performance. Simultaneously, dimensionality reduction is performed to remove redundant information from the features, reducing computational load and improving coding efficiency.

[0114] Next, a suitable encoder is selected for joint encoding. For example, the Transformer architecture in deep learning can be used. The Transformer architecture has parallel computing capabilities and long sequence processing capabilities, and can capture complex dependencies between features. The preprocessed cross-modal sentiment features, audio content features, and video content features are input into the Transformer encoder. The encoder uses a multi-head self-attention mechanism to perform weighted summation of the input features and learn the correlation between features.

[0115] Furthermore, positional encoding information is introduced during the encoding process. Since the Transformer model itself does not possess positional information, adding positional encoding enables the model to distinguish features at different locations and better understand the order and context of features. Positional encoding can be calculated using sine and cosine functions, incorporating positional information into the feature vector.

[0116] Finally, after multi-layer stacked computation by the Transformer encoder, a unified semantic representation of audio and video is output. This semantic representation includes emotional and content information from both audio and video, comprehensively and accurately describing the semantic content of the audio and video.

[0117] In summary, compared to existing technologies, this application overcomes the inaccuracy issues caused by modal differences and inconsistencies in intensity in traditional audio-visual semantic analysis methods by performing a series of operations, including intensity normalization, weighted fusion, and joint encoding, on audio and video sentiment features. In the intensity normalization stage, by calculating and adjusting the sentiment intensity difference coefficient, audio and video sentiment features can be compared and analyzed on the same scale, avoiding analysis errors caused by excessive intensity differences. The weighted fusion stage determines the weights based on the contribution of audio and video to sentiment expression in different scenarios, effectively fusing them into cross-modal sentiment features and fully utilizing multimodal information. The joint encoding stage further integrates the cross-modal sentiment features with audio and video content features to generate a unified audio-visual semantic representation, comprehensively and accurately describing the semantic content of audio and video.

[0118] In summary, the embodiments of this application have at least the following technical effects:

[0119] This application provides a multimodal contrastive learning-based audio-visual semantic analysis method. By acquiring audio and video data and extracting audio and video features separately, it normalizes and weights the emotional features, and finally performs adaptive fusion based on scene type to generate an optimized unified semantic representation. This method effectively detects and optimizes the emotional matching degree of audio and video, solving the problem of inconsistent emotional expression in existing AI video generation systems. This technical solution not only accurately processes and analyzes audio and video data but also provides innovative solutions for multiple fields. In digital character generation systems, it can detect whether the generated character's facial expressions match the emotional tone of their voice and provide optimization suggestions for expression adjustments. In AI short video production platforms, it can identify inconsistencies between the visual content and the emotional tone of the voiceover, assisting creators in making fine-tuning adjustments. In virtual anchor systems, it can evaluate the emotional consistency of the generated content in real time, ensuring that the output audio and video content is natural and harmonious. Furthermore, it can be applied to post-production quality inspection in AI film and television production, automatically identifying and marking emotionally mismatched segments that require manual intervention, improving overall production efficiency and content quality.

[0120] Example 2, as Figure 2 As shown, based on the same inventive concept as the multimodal contrastive learning audio-visual semantic analysis method provided in Embodiment 1, this application also provides a multimodal contrastive learning audio-visual semantic analysis system, including:

[0121] The information acquisition module 11 is used to acquire audio and video data and extract audio signals and video frame sequences respectively;

[0122] Feature acquisition module 12 is used to acquire audio emotion features and audio content features based on the audio signal, and to acquire video emotion features and video content features based on the video frame sequence;

[0123] The coefficient calculation module 13 is used to calculate the emotional intensity difference coefficient between the audio emotional features and the video emotional features, and to perform intensity normalization processing on the audio emotional features and the video emotional features based on the emotional intensity difference coefficient to obtain normalized audio emotional features and normalized video emotional features.

[0124] Feature processing module 14 is used to perform weighted fusion of the normalized audio sentiment features and the normalized video sentiment features to obtain cross-modal sentiment features;

[0125] The semantic generation module 15 is used to jointly encode the cross-modal sentiment features with the audio content features and the video content features to generate a unified audio and video semantic representation.

[0126] In one embodiment, the feature acquisition module 12 is specifically used for:

[0127] The audio signal is processed by an audio decoupling network to obtain the audio emotional features and the audio content features, wherein the audio decoupling network includes an audio emotional feature extraction branch and an audio content feature extraction branch;

[0128] The video frame sequence is processed by a video decoupling network to obtain the video emotion features and the video content features. The video decoupling network includes a video emotion feature extraction branch and a video content feature extraction branch.

[0129] Furthermore, in one embodiment of the application, the construction steps of the audio decoupling network include:

[0130] A sample audio signal set is constructed by retrieving multiple sample audio signals using big data.

[0131] Audio emotion feature labels are annotated for each of the sample audio signals in the sample audio signal set to construct a sample audio emotion feature label set;

[0132] Audio content feature labels are annotated for each of the sample audio signals in the sample audio signal set to construct a sample audio content feature label set;

[0133] Using the sample audio signal set as input and the sample audio emotion feature label set as supervision label, the audio emotion feature extraction branch is constructed.

[0134] Using the sample audio signal set as input and the sample audio content feature label set as supervision label, construct the audio content feature extraction branch;

[0135] The audio decoupling network is obtained by integrating the inputs of the audio emotion feature extraction branch and the audio content feature extraction branch.

[0136] Furthermore, in one embodiment of the application, the construction steps of the video decoupling network include:

[0137] A sample video frame sequence set is constructed by retrieving multiple sample video frame sequences using big data.

[0138] Video sentiment feature labels are annotated for each of the sample video frame sequences in the sample video frame sequence set to construct a sample video sentiment feature label set;

[0139] Video content feature labels are annotated for each of the sample video frame sequences in the sample video frame sequence set to construct a sample video content feature label set;

[0140] Using the sample video frame sequence set as input and the sample video sentiment feature label set as supervision label, the video sentiment feature extraction branch is constructed.

[0141] Using the sample video frame sequence set as input and the sample video content feature tag set as supervision tags, the video content feature extraction branch is constructed.

[0142] The inputs of the video emotion feature extraction branch and the video content feature extraction branch are integrated to obtain the video decoupling network.

[0143] In one embodiment, the coefficient calculation module 13 is specifically used for:

[0144] Calculate the feature intensity values ​​of the audio emotion features and the video emotion features respectively to obtain the audio emotion intensity value and the video emotion intensity value;

[0145] The emotional intensity difference coefficient is calculated based on the audio emotional intensity value and the video emotional intensity value;

[0146] The intensity of the audio emotional features is adjusted according to the emotional intensity difference coefficient to obtain the normalized audio emotional features, and the intensity of the video emotional features is adjusted according to the emotional intensity difference coefficient to obtain the normalized video emotional features.

[0147] Further, in one embodiment, the emotional intensity difference coefficient is calculated based on the audio emotional intensity value and the video emotional intensity value, including:

[0148] Based on the audio emotional intensity value, calculate the relative ratio of the video emotional intensity value to the audio emotional intensity value;

[0149] The relative ratio is used as the emotional intensity difference coefficient.

[0150] Further, the intensity of the audio emotional features is adjusted according to the emotional intensity difference coefficient to obtain the normalized audio emotional features, and the intensity of the video emotional features is adjusted according to the emotional intensity difference coefficient to obtain the normalized video emotional features, including:

[0151] Based on the square root value of the emotional intensity difference coefficient, the audio emotional features are multiplied to obtain the normalized audio emotional features.

[0152] Based on the square root value of the emotional intensity difference coefficient, the video emotional features are divided to obtain the normalized video emotional features.

[0153] In one embodiment, the feature processing module 14 specifically includes:

[0154] Identify the current scene type based on the audio and video data;

[0155] Obtain the corresponding preset audio fusion weights and preset video fusion weights based on the current scene type;

[0156] Based on the preset audio fusion weights and the preset video fusion weights, the normalized audio sentiment features and the normalized video sentiment features are weighted and combined to obtain the cross-modal sentiment features.

[0157] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this specification. Additionally, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are possible or may be advantageous.

[0158] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

[0159] This specification and accompanying drawings are merely illustrative examples of this application and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from its scope. Therefore, if such modifications and modifications fall within the scope of this application and its equivalents, this application intends to include such modifications and modifications.

Claims

1. A multimodal contrastive learning-based audio-visual semantic analysis method, characterized in that, The method includes: Acquire audio and video data, and extract audio signals and video frame sequences respectively; Audio emotion features and audio content features are obtained based on the audio signal, and video emotion features and video content features are obtained based on the video frame sequence; Calculate the emotional intensity difference coefficient between the audio emotional features and the video emotional features, and perform intensity normalization processing on the audio emotional features and the video emotional features based on the emotional intensity difference coefficient to obtain normalized audio emotional features and normalized video emotional features; The normalized audio sentiment features and the normalized video sentiment features are weighted and fused to obtain cross-modal sentiment features; The cross-modal emotion features are jointly encoded with the audio content features and the video content features to generate a unified audio-visual semantic representation.

2. The method according to claim 1, characterized in that, Based on the audio signal, audio emotional features and audio content features are obtained; based on the video frame sequence, video emotional features and video content features are obtained, including: The audio signal is processed by an audio decoupling network to obtain the audio emotional features and the audio content features, wherein the audio decoupling network includes an audio emotional feature extraction branch and an audio content feature extraction branch; The video frame sequence is processed by a video decoupling network to obtain the video emotion features and the video content features. The video decoupling network includes a video emotion feature extraction branch and a video content feature extraction branch.

3. The method according to claim 2, characterized in that, The construction steps of the audio decoupling network include: A sample audio signal set is constructed by retrieving multiple sample audio signals using big data. Audio emotion feature labels are annotated for each of the sample audio signals in the sample audio signal set to construct a sample audio emotion feature label set; Audio content feature labels are annotated for each of the sample audio signals in the sample audio signal set to construct a sample audio content feature label set; Using the sample audio signal set as input and the sample audio emotion feature label set as supervision label, the audio emotion feature extraction branch is constructed. Using the sample audio signal set as input and the sample audio content feature label set as supervision label, construct the audio content feature extraction branch; The audio decoupling network is obtained by integrating the inputs of the audio emotion feature extraction branch and the audio content feature extraction branch.

4. The method according to claim 2, characterized in that, The construction steps of the video decoupling network include: A sample video frame sequence set is constructed by retrieving multiple sample video frame sequences using big data. Video sentiment feature labels are annotated for each of the sample video frame sequences in the sample video frame sequence set to construct a sample video sentiment feature label set; Video content feature labels are annotated for each of the sample video frame sequences in the sample video frame sequence set to construct a sample video content feature label set; Using the sample video frame sequence set as input and the sample video sentiment feature label set as supervision label, the video sentiment feature extraction branch is constructed. Using the sample video frame sequence set as input and the sample video content feature tag set as supervision tags, the video content feature extraction branch is constructed. The inputs of the video emotion feature extraction branch and the video content feature extraction branch are integrated to obtain the video decoupling network.

5. The method according to claim 1, characterized in that, Calculate the emotional intensity difference coefficient between the audio emotional features and the video emotional features, and perform intensity normalization processing on the audio emotional features and the video emotional features based on the emotional intensity difference coefficient to obtain normalized audio emotional features and normalized video emotional features, including: Calculate the feature intensity values ​​of the audio emotion features and the video emotion features respectively to obtain the audio emotion intensity value and the video emotion intensity value; The emotional intensity difference coefficient is calculated based on the audio emotional intensity value and the video emotional intensity value; The intensity of the audio emotional features is adjusted according to the emotional intensity difference coefficient to obtain the normalized audio emotional features, and the intensity of the video emotional features is adjusted according to the emotional intensity difference coefficient to obtain the normalized video emotional features.

6. The method according to claim 5, characterized in that, Based on the audio emotional intensity value and the video emotional intensity value, the emotional intensity difference coefficient is calculated, including: Based on the audio emotional intensity value, calculate the relative ratio of the video emotional intensity value to the audio emotional intensity value; The relative ratio is used as the emotional intensity difference coefficient.

7. The method according to claim 5, characterized in that, The normalized audio emotional features are obtained by adjusting the intensity of the audio emotional features according to the emotional intensity difference coefficient, and the normalized video emotional features are obtained by adjusting the intensity of the video emotional features according to the emotional intensity difference coefficient, including: Based on the square root value of the emotional intensity difference coefficient, the audio emotional features are multiplied to obtain the normalized audio emotional features. Based on the square root value of the emotional intensity difference coefficient, the video emotional features are divided to obtain the normalized video emotional features.

8. The method according to claim 1, characterized in that, The normalized audio sentiment features and the normalized video sentiment features are weighted and fused to obtain cross-modal sentiment features, including: Identify the current scene type based on the audio and video data; Obtain the corresponding preset audio fusion weights and preset video fusion weights based on the current scene type; Based on the preset audio fusion weights and the preset video fusion weights, the normalized audio sentiment features and the normalized video sentiment features are weighted and combined to obtain the cross-modal sentiment features.

9. A multimodal contrastive learning-based audio-visual semantic analysis system, characterized in that, For performing the method according to any one of claims 1-8, comprising: The information acquisition module is used to acquire audio and video data and extract audio signals and video frame sequences respectively; The feature acquisition module is used to acquire audio emotion features and audio content features based on the audio signal, and to acquire video emotion features and video content features based on the video frame sequence. The coefficient calculation module is used to calculate the emotional intensity difference coefficient between the audio emotional features and the video emotional features, and to perform intensity normalization processing on the audio emotional features and the video emotional features based on the emotional intensity difference coefficient to obtain normalized audio emotional features and normalized video emotional features. The feature processing module is used to perform weighted fusion of the normalized audio sentiment features and the normalized video sentiment features to obtain cross-modal sentiment features; The semantic generation module is used to jointly encode the cross-modal sentiment features with the audio content features and the video content features to generate a unified audio and video semantic representation.

Citation Information

Patent Citations

  • Cross-modal semantic analysis method

    CN120408537A

  • Audio and video dual-mode emotion recognition method and system based on adapter fusion

    CN120411863A