Multimodal audio-video quality of experience evaluation system based on multimodal learning
The audio and video experience quality evaluation system based on multimodal learning solves the problem that traditional audio and video quality evaluation models cannot accurately capture user experience through the collaborative work of audio and video material processing, subjective quality score annotation and model training modules, and realizes efficient and automated audio and video quality evaluation.
Patent Information
- Application Number
- CN202511007572.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-22
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-07-22
AI Technical Summary
Existing audio and video quality evaluation models cannot accurately capture the real quality experience that audio and video bring to users. Traditional single-modal evaluation models cannot capture the synergy between audio and video. The subjective scoring cost is high, and the objective evaluation is disconnected from user perception.
A multimodal audio and video experience quality evaluation system based on multimodal learning is adopted. The audio and video material processing module is used to perform distortion processing. The subjective quality score annotation module is combined to obtain the real subjective score. The model training module is used to capture the bidirectional dependency relationship of audio and video features to construct a multimodal audio and video experience quality evaluation model.
It improves the accuracy, comprehensiveness and automation level of audio and video quality evaluation, provides an efficient and reliable decision-making basis for communication system coding optimization and bandwidth resource allocation, and reduces bit rate and bandwidth waste.
Smart Images

Figure CN120508944B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of audio and video quality evaluation, and in particular to a multimodal audio and video experience quality evaluation system based on multimodal learning. Background Art
[0002] With the widespread adoption of multimedia communications, the proportion of video in global mobile data continues to increase annually. In this context, compressing and optimizing video data can reduce data transmission volume, alleviate bandwidth pressure, improve transmission efficiency, and alleviate the pressure of multimedia communications. As users' demands for high-definition video quality and user experience increase, selecting appropriate audio and video experience quality evaluation models can accurately assess the video experience and guide communication systems to optimize encoding and resource allocation, reducing bitrates and conserving bandwidth.
[0003] Existing audio and video quality of experience evaluation models typically use video data as training material, leveraging neural networks to learn video features and then predict audio and video quality of experience scores. However, audio and video quality evaluation models trained solely on video cannot accurately capture the true quality of audio and video experience for users. Summary of the Invention
[0004] This application provides a multimodal audio and video experience quality evaluation system based on multimodal learning to solve the problem that existing video quality evaluation models cannot accurately capture the real quality experience brought to users by audio and video.
[0005] The multimodal audio and video experience quality evaluation system based on multimodal learning includes the following steps:
[0006] an audio and video material processing module, configured to perform distortion processing on audio and video materials of multiple audio and video types to obtain audio and video materials of multiple quality level combinations of the multiple audio and video types, wherein a quality level combination includes an audio distortion level and a video distortion level;
[0007] A subjective quality score labeling module is used to play the audio and video materials of each quality level combination of each audio and video type, obtain the user's subjective quality score, and label the audio and video materials of each quality level combination of each audio and video type with the real subjective quality score to obtain multiple labeled audio and video materials;
[0008] The model training module is used to capture the bidirectional dependency between the video and audio features of each annotated audio and video material, and use the multiple annotated audio and video materials to train a multimodal audio and video experience quality evaluation model. The multimodal audio and video experience quality evaluation model is used to comprehensively evaluate the quality score of audio and video from both audio and video perspectives.
[0009] Compared with the prior art, this application has the following advantages:
[0010] By separating audio tracks and performing joint distortion processing on multi-scene audio and video materials, combined with parameterized control of constant quality factors and audio bitrates, subjective quality score samples covering multiple quality level combinations are generated, ensuring sample diversity and scene representativeness. By employing multiple rounds of playback, two-dimensional average scoring, and standardized experimental process design, single-shot scoring errors and individual biases are effectively eliminated, resulting in highly reliable subjective scoring labels. By constructing a multimodal audio and video experience quality evaluation model to capture the bidirectional dependencies between audio and video features, a comprehensive assessment of overall audio and video quality is achieved from both audio and video dimensions. Through the collaborative work of three modules: audio and video material processing, subjective quality score annotation, and model training, this approach addresses technical issues such as the inability of traditional single-modal evaluation models to capture audio and video synergy, the high cost of subjective scoring, and the disconnect between objective evaluation and user perception. This approach improves the accuracy, comprehensiveness, and automation of audio and video quality evaluation, providing an efficient and reliable decision-making basis for coding optimization and bandwidth resource allocation in communication systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments of the present application. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0012] Figure 1 A schematic diagram of the structure of a multimodal audio and video quality of experience evaluation system based on multimodal learning provided in an embodiment of the present application is shown;
[0013] Figure 2 The following is a flowchart of an experiment for subjective quality scoring provided by an embodiment of the present application;
[0014] Figure 3 A schematic diagram of the structure of a model training module provided in one embodiment of the present application is shown;
[0015] Figure 4 A schematic diagram of the structure of a video encoder provided by an embodiment of the present application is shown;
[0016] Figure 5 A schematic diagram of the structure of a cross attention unit provided in one embodiment of the present application is shown;
[0017] Figure 6 The figure shows the impact of the audio and video quality provided by an embodiment of the present application on the user's subjective score. DETAILED DESCRIPTION
[0018] In order to make the above objectives, characteristics and advantages of the present application more obvious and easy to understand, the present application will be further described in detail below with reference to the drawings and specific embodiments.
[0019] In recent years, with the rapid development of communication technology and artificial intelligence technology, audio and video data have gradually become important resources for transmission and utilization in modern information systems. In the field of communication transmission, the explosive growth of multimedia communication services such as video conferencing, real-time live broadcast, short video platforms, etc. has made the continuous transmission and interaction of massive data in the communication system put forward very high requirements for the bandwidth of communication transmission. At the same time, with the popularization of various multimedia communication services, users' requirements for the quality and experience of high-definition audio and video content are also increasing. Under this background, how to realize the compression and optimization of audio and video data under the premise of ensuring user experience has become a problem to be solved in the field of multimedia communication. Selecting a suitable audio and video quality of experience evaluation model can effectively guide the encoding optimization and resource allocation of the communication system, and then reduce the code rate and save bandwidth.
[0020] At present, many studies have carried out in-depth research on the evaluation of audio and video quality of experience, especially in the field of communication, researchers often use quality of service (QoS) as an evaluation index to compress data in the communication system under the condition of ensuring a certain service quality. However, the quality of service index is calculated objectively by each pixel of the image block and other components of the data, and cannot well reflect the real experience of users, which leads to a certain degree of waste of code rate and bandwidth. In addition to the quality of service method, there are some common problems in the commonly used objective evaluation method of data: first, they mostly need reference videos and cannot independently evaluate the quality; second, they also cannot reflect the real feelings of users. Traditional audio and video quality of experience evaluation methods also include quality of experience (QoE) based on user experience, which evaluates data through users' subjective cognition, such as the commonly used mean opinion score (MOS). However, the mean opinion score method also has problems such as high cost, labor-intensive, poor scalability, and insufficient real-time performance. Therefore, it is urgent to establish a deep learning model to evaluate the quality of experience of video data efficiently and automatically.
[0021] Most existing deep learning-based video quality of experience evaluation methods only use video data as training materials, and the trained neural network may rely on low-level visual information such as color and brightness to give a prediction of the quality evaluation score. Most deep learning-based video quality evaluation still focuses on a single modality, such as only evaluating a single video quality.
[0022] This application takes into account that from the user's perspective, what they receive is often multimodal data that contains both video and audio. This type of model uses a single-modal processing method, only independently analyzing the video, and does not fully consider the interaction between the two in user perception. Specifically, by separately controlling the video compression parameters to generate test samples, an evaluation system for joint distortion scenarios has not been established, making it difficult to capture the complex relationship between audio and video modalities. As a result, the existing audio and video quality evaluation model does not adequately consider the synergy between audio and video, and cannot accurately capture the real experience of users in audio and video fusion scenarios.
[0023] Based on the above technical problems, this application proposes a multimodal audio and video experience quality evaluation system based on multimodal learning. Figure 1 The multimodal audio and video experience quality evaluation system based on multimodal learning includes:
[0024] The audio and video material processing module 100 is used to perform distortion processing on audio and video materials of multiple audio and video types to obtain audio and video materials of multiple quality level combinations of the multiple audio and video types, where a quality level combination includes an audio distortion level and a video distortion level.
[0025] When performing distortion processing on audio and video materials of multiple audio and video types, first extract the audio tracks and video tracks of each audio and video type of audio and video materials, perform multiple video distortion level processing on the video track of each audio and video type of audio and video materials, and perform multiple audio distortion level processing on the audio track of each audio and video type of audio and video materials, to obtain audio and video materials of multiple quality level combinations of multiple audio and video types, where one quality level combination includes an audio distortion level and a video distortion level.
[0026] Specifically, the multiple audio and video types include daytime scenes, nighttime scenes, scenes with animals and humans, and scenes without specific target objects. When processing video distortion levels, the constant rate factor (CRF) is varied to achieve different levels of video quality compression, resulting in multiple video distortion levels. When processing audio distortion levels, the audio bitrate is varied to introduce multiple audio distortion levels. By pairing audio tracks with different audio distortion levels with video tracks with different video distortion levels, multiple quality level combinations for multiple audio and video types are achieved.
[0027] Exemplarily, the audio-video material processing module 100 performs multiple video distortion level processing on the video track of the original audio-video material of each audio-video type, including: for the video track of the original audio-video material of each audio-video type, the audio-video material processing module 100 performs different levels of video quality compression by changing the constant quality factor; the audio-video material processing module 100 performs multiple audio distortion level processing on the video track of the original audio-video material of each audio-video type, including: for the audio track of the original audio-video material of each audio-video type, the audio-video material processing module 100 performs different levels of audio quality compression by changing the audio bit rate. For example, four video distortion levels are introduced by setting the constant quality factor of the video track of each audio-video material of multiple audio-video types to 1, and setting other constant quality factors to 37, 47, and 51, respectively; when performing audio distortion level processing, the audio track is processed first to ensure consistency between the processed audio track and the video track, so as to ensure audio-visual synchronization; then, three audio distortion levels are obtained by setting the bit rate of the audio track of each audio-video material of multiple audio-video types to 64 kbps, and setting other bit rates to 32 kbps and 16 kbps, respectively. For a single audio-video material of a single audio-video type, 12 quality level combinations of audio-video materials are obtained by combining the audio track of different audio distortion levels and the video track of different video distortion levels two by two. Under 20 audio-video types of audio-video materials, 240 quality level combinations are obtained.
[0028] In this embodiment, by extracting the audio track and the video track of each audio-video material of multiple scene audio-video types respectively, and performing different distortion level processing combined with specific parameter settings while maintaining audio-visual synchronization to reduce the impact on user perception experience, subjective quality score samples covering a wide range of audio-video quality are generated by two-by-two combination, ensuring the diversity, representativeness and accuracy of the score data, and ensuring that the trained multi-modal audio-video experience quality evaluation model can comprehensively and accurately evaluate the audio-video experience quality.
[0029] The subjective quality score labeling module 200 is configured to play each quality level combination of audio-video material of each audio-video type, obtain the subjective quality score of the user, and label the true subjective quality score for each quality level combination of audio-video material of each audio-video type, to obtain multiple labeled audio-video materials.
[0030] Specifically, the user's subjective quality rating refers to collecting users' subjective ratings of audio and video materials of each quality level combination of each audio and video type in multiple rounds under a standard experimental environment; using the subjective ratings of the audio and video materials of each quality level combination of each audio and video type to label the audio and video materials of each quality level combination of each audio and video type, and obtaining multiple labeled audio and video materials.
[0031] In some embodiments, the subjective quality score annotation module 200 is further configured to:
[0032] For the audio and video material of the i-th audio and video type with the j-th video distortion level and the k-th audio distortion level, play it N times.
[0033] Obtain the subjective quality score of the audio and video material of the i-th audio and video type with the j-th video distortion level and the k-th audio distortion level by the m-th user in the n-th round.
[0034] The average is taken from the round dimension to obtain the average subjective quality score of the mth user for the audio and video material of the i-th audio and video type with the j-th video distortion level and the k-th audio distortion level.
[0035] Averaging from the user dimension, we obtain the average subjective quality score of the audio and video material of the i-th audio and video type with the j-th video distortion level and the k-th audio distortion level.
[0036] The value of n ranges from 1 to N, the value of m ranges from 1 to M, the value of i ranges from 1 to I, the value of j ranges from 1 to J, and the value of k ranges from 1 to K.
[0037] Specifically, the specific values of N, M, I, J, and K are set according to actual needs. By obtaining the subjective quality score of the audio and video material with the jth video distortion level and the kth audio distortion level of the i-th audio and video type for each user in each round, the average subjective quality score of the audio and video material with the jth video distortion level and the kth audio distortion level of the i-th audio and video type, the average subjective quality score of the audio and video material with the jth video distortion level and the kth audio distortion level of the m-th user, and the average subjective quality score of the audio and video material with the jth video distortion level and the kth audio distortion level of the i-th audio and video type are obtained as the true value for training the multimodal audio and video experience quality evaluation model.
[0038] For example, refer to Figure 2, with 10 users, 2 rounds, 20 audio and video types, 4 video distortion levels, and 3 audio distortion levels. None of the 10 users participating in the subjective quality rating had ever participated in a similar experiment. All users were informed of the specific experimental procedures and details before the experiment began. In the same subjective quality rating environment, each user participated in 480 subjective quality ratings, with a short break between each subjective quality rating to maintain the reliability and validity of the ratings. Each subjective quality rating consisted of a first gray screen phase, an audio and video material playback phase, and a second gray screen phase. The first gray screen phase lasted for milliseconds to stabilize the subject's state. After the first gray screen phase, the audio and video material playback phase played 3 seconds of audio and video material to stimulate the user's sensory perception. Afterwards, subtitles appeared in the second gray screen phase to prompt the user to rate the overall quality of the audio and video material on a scale of 1-5. The higher the score, the higher the subjective quality.
[0039] For example, the subjective quality score is recorded as , determine the average subjective quality score of the mth user for the audio and video material of the i-th audio and video type with the j-th video distortion level and the k-th audio distortion level according to the following formula: :
[0040]
[0041] Among them, i represents the audio and video type, j represents the video distortion level, k represents the audio distortion level, and n represents the number of audio and video material playback rounds. Indicates averaging from the round dimension, It represents the average subjective quality score of the mth user for the audio and video material of the i-th audio and video type with the j-th video distortion level and the k-th audio distortion level.
[0042] The average subjective quality score of the audio and video material of the i-th audio and video type with the j-th video distortion level and the k-th audio distortion level is determined according to the following formula:
[0043]
[0044] Among them, i represents the audio and video type, j represents the video distortion level, k represents the audio distortion level, and n represents the number of audio and video material playback rounds. Indicates the average from the user dimension, The average subjective quality score of the audio and video materials with the jth video distortion level and the kth audio distortion level of the i-th audio and video type.
[0045] In this embodiment, by playing the same audio and video material multiple times and collecting subjective quality ratings, the randomness and error of individual ratings can be reduced, improving the stability and reliability of subjective quality ratings. Averaging across rounds and users eliminates individual rating differences and fluctuations in individual ratings, resulting in a more objective and representative average subjective quality rating. By flexibly configuring the number of users (M), rounds (N), audio and video type (I), video distortion level (J), and audio distortion level (K), the system can adapt to the needs of experiments of varying scales, ensuring the comprehensiveness and scalability of rating data. Controlling subjective quality ratings reduces external interference and improves rating consistency and comparability. By stabilizing user status through gray screen periods, playing audio and video material for a fixed duration, and maintaining rating reliability through short breaks, users can be guided to focus on audio and video quality assessment, improving rating accuracy and effectiveness. Furthermore, subjective quality ratings quantify user perceptions and provide structured ground truth labels for training multimodal audio and video quality of experience evaluation models.
[0046] By playing the same audio and video material multiple times and collecting subjective quality ratings, combined with dual-dimensional averaging of rounds and users, we can effectively reduce the error in single ratings and obtain more representative subjective quality rating results. Flexible configuration of the number of users, rounds, audio and video types, and distortion level parameters can adapt to the subjective quality rating needs of different scales, ensuring data comprehensiveness and scalability. Through standardized subjective quality rating environment control and scientific scoring process design, users can be guided to focus on evaluation, external interference can be reduced, and the accuracy of subjective quality ratings can be improved. Structured subjective quality rating labels provide a reliable quantitative perceptual benchmark for the training of multimodal audio and video experience quality evaluation models, forming a complete closed loop from data collection to model training.
[0047] The model training module 300 is used to capture the bidirectional dependency between the video features and audio features of each annotated audio and video material, and to use multiple annotated audio and video materials to train a multimodal audio and video experience quality evaluation model. The multimodal audio and video experience quality evaluation model is used to comprehensively evaluate the quality scores of audio and video from both audio and video perspectives.
[0048] Specifically, the bidirectional dependency relationship characterizes the mutual influence and constraint between the audio features in the audio track of each annotated audio and video material and the video features in the audio track of each annotated audio and video material. When using multiple annotated audio and video materials for model training to obtain a multimodal audio and video experience quality evaluation model, by capturing the bidirectional dependency relationship between the video and audio features of each annotated audio and video material for training, the multimodal audio and video experience quality evaluation model can more accurately assess the overall quality of the audio and video.
[0049] The embodiment of the application proposes a multi-modal learning-based audio-video experience quality evaluation system. Through audio track separation and joint distortion processing of multi-scene audio-video materials, combined with parameterized control of constant quality factors and audio bit rates, subjective quality score samples covering various quality level combinations are generated to ensure sample diversity and scene representativeness. Multi-round playback, double-dimension average scoring, and standardized experimental process design are adopted to effectively eliminate single scoring errors and individual biases, and obtain high-credibility subjective scoring labels. A multi-modal audio-video experience quality evaluation model is constructed to capture the bidirectional dependency relationship of audio-video features, and realize comprehensive evaluation of the overall quality of audio-video from the audio and video double dimensions. Through the collaborative work of the three modules of audio-video material processing, subjective quality scoring labeling, and model training, the technical problems of the traditional single-modal evaluation model that cannot capture the synergistic effect of audio-video, high subjective scoring cost, and the disconnection between objective evaluation and user perception are solved. The accuracy, comprehensiveness, and automation level of audio-video quality evaluation are improved, and efficient and reliable decision-making basis is provided for communication system coding optimization and bandwidth resource allocation.
[0050] In some embodiments, the audio-video material processing module 100 is further configured to extract the audio track and the video track of each labeled audio-video material.
[0051] The model training module 300 is configured to:
[0052] The video encoder in the to-be-trained model is used to perform video feature extraction on the video track of each labeled audio-video material, to obtain the video feature of each labeled audio-video material.
[0053] The audio encoder in the to-be-trained model is used to perform audio feature extraction on the audio track of each labeled audio-video material, to obtain the audio feature of each labeled audio-video material.
[0054] The multi-modal fusion module in the to-be-trained model is used to capture the bidirectional dependency relationship between the video feature and the audio feature of each labeled audio-video material, to fuse the video feature and the audio feature of each labeled audio-video material, to obtain the audio fusion feature and the video fusion feature of each labeled audio-video material. Based on the audio fusion feature and the video fusion feature of each labeled audio-video material, the predicted quality score of each labeled audio-video material is obtained.
[0055] Based on the predicted quality score of each labeled audio-video material and the label carried by the labeled audio-video material, the model parameters of the to-be-trained model are updated, to obtain a multi-modal audio-video experience quality evaluation model.
[0056] Specifically, the model training module 300 comprises a video encoder, an audio encoder, and a multi-modal fusion module. The video encoder is used to extract the video features of the video track of each annotated audio-video material, and the audio encoder is used to extract the audio features of the audio track of each annotated audio-video material. The extracted audio features and video features are simultaneously sent to the multi-modal fusion module. The multi-modal fusion module captures the bidirectional dependency between the video features and the audio features, and performs deep interaction and fusion on the two modal features to generate audio fusion features and video fusion features containing both video information and audio information. The predicted quality score of each annotated audio-video material is obtained based on the audio fusion features and the video fusion features. The model parameters of the training model are updated using the predicted quality score of the annotated audio-video material and the subjective quality score carried by the annotated audio-video material, and a multi-modal audio-video quality of experience evaluation model is obtained.
[0057] For example, referring to Figure 3 and Figure 4 , the video encoder is composed of a convolutional neural network backbone, a Transformer module, and a fully connected layer. The convolutional neural network backbone is used to extract the video features of the video track of each annotated audio-video material, and is composed of five convolutional layers with ReLU activation function. Each convolutional layer uses a convolution kernel with a size of 3*3 and a step of 2, and adopts maximum pooling and batch normalization operation after convolution. The Transformer module is used to extract the bidirectional dependency between the video features and the audio features, and is composed of two Transformer encoder layers. Each Transformer encoder layer contains 8 attention heads, and the hidden dimension is 512. The position encoding in the Transformer encoder layer is used to identify the correlation and mutual influence between the video features in the time dimension. The fully connected layer is used to adjust the dimension of the feature vector to adapt to multi-modal fusion and improve the expression ability of the video features. The video features are efficiently extracted by the video encoder, the temporal dependency is preserved, and the compatibility of the video features in multi-modal fusion is improved.
[0058] The audio encoder is composed of a Mel feature extraction layer, a convolutional neural network layer, and a fully connected layer. The Mel feature extraction layer is used to convert the audio track of each annotated audio-video material into a Mel spectrogram to provide audio representation consistent with human auditory perception. The convolutional neural network layer includes four convolution kernels with a size of 3*3 and ReLU activation function, and enters the fully connected layer after batch normalization. The audio features are mapped to the same dimension as the video features, and the video and audio features are aligned in dimension to provide a comparable feature space for multi-modal fusion.
[0059] Table 1 shows the performance comparison results of the method provided in one embodiment of the present application and other multimodal methods on the test set. In the comparison method, this experiment adopted the ConvLSTM+SVM encoder method and the C3D+CNN-Audio encoder method, wherein the ConvLSTM+SVM encoder method refers to the use of the ConvLSTM network and the support vector machine method to extract video features and audio features respectively, and the C3D+CNN-Audio encoder method refers to the use of the C3D method to extract video features, while using CNN to extract features for the audio Mel; In addition, this experiment also adopted four multimodal fusion methods MCB, MRRF, CBAM and AVFF that are different from cross-attention for performance comparison. It can be seen that the video feature and audio feature extraction method provided by this application is superior to other methods in performance.
[0060] Table 1 Performance comparison of this method and other multimodal methods on the test set
[0061]
[0062] The multimodal fusion module includes a cross-attention unit and a gating unit. The audio and video features of each audio track of annotated audio and video material are simultaneously fed into the cross-attention unit, which then outputs a feature vector and feeds it into the gating unit. After normalization, the predicted quality score for each annotated audio and video material is obtained. By calculating bidirectional attention between audio and video features, the enhancement effect of audio features on video features and the constraint effect of video features on audio features are dynamically modeled, achieving deep modal interaction.
[0063] The embodiments of the present application achieve a leap from single-modality to multi-modal collaboration in audio and video quality evaluation through the joint extraction of spatiotemporal features of the video encoder, the joint modeling of time-frequency-depth features of the audio encoder, and the bidirectional dependency capture and gated fusion of the multimodal fusion module, significantly improving the model's perception of joint distortion scenarios and providing an efficient and automated decision-making basis for communication system bandwidth optimization.
[0064] Table 2 shows the results of the audio and video ablation experiment provided by an embodiment of the present application. Referring to Table 2, it can be seen that the performance changes after removing the key components in the model, where "w / o" means that the component is removed in the method. The results show that the audio encoder, cross-attention unit, and dynamic gating unit all play a significant role in performance improvement. This also further illustrates the importance of audio in video quality evaluation and the superiority of the cross-attention fusion method.
[0065] Table 2 Audio and video ablation experiment results
[0066]
[0067] In some embodiments, the multimodal fusion module in the model to be trained is used to fuse the video features and audio features of each annotated audio and video material with the goal of capturing the bidirectional dependency between the video features and audio features of each annotated audio and video material, thereby obtaining audio fusion features and video fusion features of each annotated audio and video material, including:
[0068] For each annotated audio and video material, the following steps are performed through the cross-attention unit in the multimodal fusion module:
[0069] The Q transformation matrix, the K transformation matrix, and the V transformation matrix are combined to determine the Q, K, and V values of the video features and audio features of the annotated audio and video material.
[0070] Specifically, the Q-transform matrix is a weight matrix that maps input features to query vectors. It is used to represent the attention that audio features require for video features, or vice versa. The Q-value is the vector obtained by linearly transforming audio or video features using the Q-transform matrix. It represents the information that the current feature hopes to obtain from other features.
[0071] The K transformation matrix is a weight matrix that maps input vectors to key vectors, representing the key information of audio and video features for query matching. The K value is the vector obtained by linearly transforming the audio or video features using the K transformation matrix. It represents the key information of the current feature and is used to match query requirements.
[0072] The V transformation matrix is a weight matrix that maps input vectors to value vectors. It is used to represent the actual content of audio and video features as the attention-weighted output. The V value is the vector obtained by linearly transforming the audio or video features using the V transformation matrix. It represents the actual content of the audio and video features as the attention-weighted output.
[0073] For example, Figure 5 The structure diagram of the cross attention unit provided by an embodiment of the present application is shown. Figure 5 , determine the Q, K, and V values of the video and audio features of the annotated audio and video materials according to the following formula:
[0074]
[0075]
[0076]
[0077]
[0078]
[0079]
[0080] in, Represents video features, Represents audio features, 、 and They are Q transformation matrix, K transformation matrix and V transformation matrix respectively, and Represent the query matrices calculated from video features and audio features respectively, and Represent the key matrices calculated from video features and audio features respectively, and Represent the value matrices calculated from video features and audio features respectively.
[0081] According to the Q value of the video feature of the annotated audio and video material and the K value of the audio feature of the annotated audio and video material, the attention score of the video feature to the audio feature is calculated, and, according to the Q value of the audio feature of the annotated audio and video material and the K value of the video feature of the annotated audio and video material, the attention score of the audio feature to the video feature is calculated.
[0082] Specifically, the dot product of the Q value and the K value is used to calculate the attention score to quantify the correlation between audio features and video features, and the V value is used as a weighted target to map the correlation to the actual content to achieve feature fusion.
[0083] For example, the attention score of the video feature to the audio feature is calculated according to the following formula:
[0084]
[0085] in, represents the attention score of video features to audio features, d is the dimension of the feature, Represents the dot product similarity between the query and the key.
[0086] The attention score of audio features to video features is calculated according to the following formula:
[0087]
[0088] in, represents the attention score of audio features to video features, d is the dimension of the feature, Represents the dot product similarity between the query and the key.
[0089] According to the attention score of the video feature on the audio feature, an attention distribution of the video feature on the audio feature is determined, and according to the attention score of the audio feature on the video feature, an attention distribution of the audio feature on the video feature is determined.
[0090] Exemplarily, the attention distribution of the video feature on the audio feature is determined according to the following formula:
[0091]
[0092] The attention distribution of the audio feature on the video feature is determined according to the following formula:
[0093]
[0094] wherein, represents the attention distribution of the video feature on the audio feature, represents the attention distribution of the audio feature on the video feature, and the softmax operation is used to normalize the attention distribution into a probability distribution, so that the sum of all attention weights is 1.
[0095] According to the attention distribution of the video feature on the audio feature and the V value of the audio feature of the labeled audio-video material, the audio feature adjusted based on the video feature is obtained, and according to the attention distribution of the audio feature on the video feature and the V value of the video feature of the labeled audio-video material, the video feature adjusted based on the audio feature is obtained.
[0096] Exemplarily, the audio feature adjusted based on the video feature and the video feature adjusted based on the audio feature are determined according to the following formula:
[0097]
[0098]
[0099] wherein, represents the video feature adjusted based on the audio feature, represents the audio feature adjusted based on the video feature.
[0100] In combination with the weight matrix used for feature transformation, the audio feature adjusted based on the video feature and the video feature adjusted based on the audio feature are respectively subjected to linear transformation.
[0101] The linearly transformed audio feature is determined as the audio fusion feature of each labeled audio-video material, and the linearly transformed video feature is determined as the video fusion feature of each labeled audio-video material.
[0102] Exemplarily, the linear transformation is performed according to the following formula:
[0103]
[0104]
[0105] Exemplarily, the fused video features and the fused audio features are expressed as follows:
[0106]
[0107] in, represents the fused video features, represents the fused audio features, Represents video features, Represents audio features, Represents the weight matrix used for feature transformation, and C represents the fused video features And the fused audio features The relationship function between video features and audio features.
[0108] The embodiments of the present application capture the bidirectional dependency between audio features and video features so that the adjusted fusion features simultaneously contain interactive information from both modalities. This allows the trained multimodal audio and video quality of experience evaluation model to not only perceive the impact of video quality on audio perception when performing audio and video quality assessment, but also identify the feedback effect of audio quality on video perception, thereby improving the accuracy of the judgment of the comprehensive audio and video experience quality.
[0109] Furthermore, based on the predicted quality score of each annotated audio and video material and the annotations carried by the annotated audio and video, the model parameters of the training model are updated to obtain a multimodal audio and video experience quality evaluation model, including:
[0110] Through the gating unit in the multimodal fusion module, the following steps are performed:
[0111] Based on the audio fusion features and video fusion features of each annotated audio and video material, the gating weights of the audio modality and the gating weights of the video modality are generated. The gating weights of the audio modality are used to: characterize the contribution of the audio track of each annotated audio and video material to the predicted quality score of the annotated audio and video material; the gating weights of the video modality are used to: characterize the contribution of the video track of each annotated audio and video material to the predicted quality score of the annotated audio and video material.
[0112] Based on the audio fusion features and video fusion features of each annotated audio and video material, the gating weight of the audio modality and the gating weight of the video modality, a predicted quality score of each annotated audio and video material is obtained.
[0113] Specifically, the gating unit consists of a fully connected layer and a sigmoid function. The gating unit adaptively adjusts the contribution of each modality by generating gating weights. The gating weight of the audio modality is used to characterize the contribution of the audio track of each annotated audio and video material to the predicted quality score of the annotated audio and video material. The gating weight of the video modality is used to characterize the contribution of the video track of each annotated audio and video material to the predicted quality score of the annotated audio and video material. Based on the audio fusion features and video fusion features of each annotated audio and video material, the gating weight of the audio modality and the gating weight of the video modality, the predicted quality score of each annotated audio and video material is obtained. Based on the predicted quality score of each annotated audio and video material and the annotations carried by the annotated audio and video, the model parameters of the model to be trained are updated to obtain a multimodal audio and video experience quality evaluation model.
[0114] For example, the fused video features And the fused audio features The gating unit in the multimodal fusion module adjusts the fused video features. The gate weights and fused audio features The gating weights are used to make more accurate quality score predictions.
[0115] After the fusion of video features And the fused audio features After average pooling, a two-layer fully connected network with ReLU activation function is input, and then normalized by sigmoid activation function to obtain the fused video features. The gate weights and fused audio features Specifically, it is calculated according to the following formula:
[0116]
[0117]
[0118] in, Represents the fused video features The gate weight, Represents the fused audio features , sigmoid and ReLU represent different activation functions, FC represents the fully connected layer, and mean represents the calculated average.
[0119] Next, use the fused video features The gate weight And the fused audio features The gate weight , for the fused video features And the fused audio features Perform weighted summation to obtain the final fusion feature vector. The final fusion feature vector is calculated according to the following formula:
[0120]
[0121] in, Represents the final fused feature vector.
[0122] The final fused feature vector is input into a fully connected header to obtain the predicted quality score, and the MES loss function is used to update the model parameters of the training model to obtain a multimodal audio and video experience quality evaluation model. The specific formula is as follows:
[0123]
[0124]
[0125] Among them, FC represents a fully connected header, represents the prediction quality score, represents the MES loss function for updating model parameters, Indicates the annotation score carried by each annotated audio or video material.
[0126] The embodiment of the present application implements dynamic modal contribution adjustment through a gating unit, thereby improving the accuracy, robustness, and generalization capability of the multimodal audio and video experience quality evaluation model, and optimizing training efficiency and model interpretability.
[0127] Furthermore, it also includes:
[0128] The actual subjective quality scores of audio and video materials at all audio distortion levels under the same video distortion level are compared to determine the degree of influence of audio modality on subjective quality scores.
[0129] The actual subjective quality scores of audio and video materials at all video distortion levels under the same audio distortion level are compared to determine the impact of video modality on subjective quality scores.
[0130] According to the influence of the video modality and the influence of the audio modality on the subjective quality score, the initialization weights of the audio modality and the video modality are determined.
[0131] The gating unit is initialized according to the initialization weights of the audio modality and the video modality respectively.
[0132] Specifically, the impact refers to the magnitude of score fluctuations caused by the same modality at the same distortion level. This can be achieved by calculating the standard deviation or range of the scores, quantifying the independent contribution of each modality to the subjective quality score. Initialization weights refer to the initial parameters used in the gating unit to adjust the contribution of the audio and video modalities. This can be achieved by normalizing the ratio of the impact of the two modalities, ensuring that the weight distribution matches the actual effect of the modalities.
[0133] For example, Figure 6 The effect of the audio and video quality provided by an embodiment of the present application on the user's subjective rating is shown in FIG. 1 , where the user's subjective quality rating is compared under the levels of no distortion and severe distortion. Figure 6 As shown. By comparing the real subjective quality score differences of audio and video materials at all audio distortion levels under the same video distortion level, the impact of audio quality changes on user perception can be quantified. When the video quality maintains CRF=37, the greater the drop in score when the audio bitrate drops from 64kbps to 16kbps, the greater the impact of the audio modality on the subjective quality score. Similarly, when the audio bitrate is fixed at 32kbps, the impact of the video modality can be calculated by comparing the score changes corresponding to different video CRF values. The bimodal influence is normalized, for example, the audio influence is divided by the sum of the bimodal influence to obtain the initialization weight of the audio modality, and the remainder is used as the initialization weight of the video modality. Based on this weight, the fully connected layer parameters of the gated unit are initialized, so that the model has a reasonable modal fusion tendency at the beginning of training, avoiding the convergence direction uncertainty problem caused by random initialization.
[0134] For example, for a distortion-free video, adding the original audio increases the subjective score from 0.987 to 0.993, indicating that the original audio slightly improves the user's video quality experience. When audio quality degrades, it negatively impacts user perception, and higher audio distortion levels increase this negative impact. For example, when the audio bitrate is 32kbps, or an audio distortion level of 1, the average subjective score drops to 0.901; when the audio bitrate is 16kbps, or an audio distortion level of 2, the subjective score drops to 0.784.
[0135] For severely distorted videos, adding the original audio increased the average subjective rating from 0.248 to 0.316, demonstrating that it also improves the user experience. However, when audio quality degrades, the improvement is minimal. For example, at an audio distortion level of 1, the subjective rating increased from 0.248 to 0.251. It can even further degrade the experience. For example, at an audio distortion level of 2, the subjective rating decreased from 0.248 to 0.222.
[0136] Therefore, during audio and video quality evaluation, poor audio quality can lead to a decline in the overall quality of experience, while high-quality audio can complement the video information and improve the overall quality of experience. By constructing a modal influence assessment mechanism based on real subjective quality scores and using objective experimental data to derive initialization weights, the initial state of the gating unit is made closer to the actual modal contribution distribution, effectively shortening the model convergence cycle, improving model training efficiency, reducing the number of parameter adjustments during training, and enhancing the consistency between model predictions and user subjective perception.
[0137] In some embodiments, the invention further includes an EEG signal acquisition module, which is used to:
[0138] Obtain an EEG signal of the mth user during the nth round when watching an audio / video material of the ith audio / video type with the jth video distortion level and the kth audio distortion level.
[0139] Averaging is performed from the round dimension to obtain the average EEG signal of the mth user when watching the audio and video material of the ith audio and video type with the jth video distortion level and the kth audio distortion level.
[0140] Averaging is performed from the user dimension to obtain the average EEG signal of the audio and video material of the i-th audio and video type with the j-th video distortion level and the k-th audio distortion level.
[0141] The model training module 300 is also used to:
[0142] The average EEG signal of the audio and video material of the i-th audio and video type with the j-th video distortion level and the k-th audio distortion level is used to train an audio and video quality rating model, which is used to evaluate the audio quality level and video quality level of the audio and video.
[0143] Specifically, averaging from the round dimension refers to averaging the EEG signals generated by the same user watching the same audio and video material multiple times. Specifically, it can be implemented by arithmetic averaging or weighted averaging methods to eliminate noise interference caused by random factors during individual viewing. Averaging from the user dimension refers to averaging the EEG responses generated by different users watching the same material. Specifically, it can be implemented by group mean calculation methods to eliminate the impact of individual physiological differences on EEG characteristics. The audio and video quality rating model refers to a cross-modal mapping model constructed based on a neural network. Specifically, it can be implemented by a structure combining a convolutional neural network and a fully connected layer to extract characteristic patterns associated with audio distortion levels and video distortion levels from the average EEG signal.
[0144] For example, when a user watches audio and video materials of a specific quality combination, the EEG signal acquisition module synchronously records its original physiological response data. In the case where each user watches the same audio and video material multiple times, the multiple recorded data of the same user are first averaged to effectively eliminate signal fluctuations caused by attention fluctuations or environmental interference. The average EEG data of different users are further averaged twice to eliminate data deviations caused by differences in individual EEG characteristics. The EEG data set after double averaging can reflect the stable physiological response characteristics of a group of users under a specific quality level combination. The model training module 300 associates the processed EEG data with known audio distortion levels and video distortion levels for training, so that the model can identify the EEG feature change rules corresponding to different quality levels, and establish an objective evaluation standard from EEG signals to audio and video quality levels.
[0145] The embodiment of the present application uses a double averaging mechanism to effectively suppress random noise and individual differences while retaining quality-related features, making the mapping relationship between EEG signals and quality levels more reliable and generalizable. It achieves an objective assessment of audio and video quality levels based on group EEG characteristics, solving the problem of poor stability of traditional subjective scoring methods. By establishing a correlation model between EEG signals and audio distortion levels and video distortion levels, it is possible to accurately identify the specific modes that cause quality degradation, providing a quantifiable evaluation basis for audio and video quality optimization.
[0146] In some embodiments, it also includes: an audio and video evaluation module and a low quality cause analysis module.
[0147] The audio and video material processing module 100 is also used to extract the audio track and video track of the target audio and video.
[0148] The EEG signal acquisition module is also used to collect EEG signals of users while they are watching target audio and video.
[0149] The audio and video evaluation module is used to input the audio track and video track of the target audio and video into the multimodal audio and video experience quality evaluation model to obtain the quality score of the target audio and video, and to input the EEG signals of the user while watching the target audio and video into the audio and video quality rating model to obtain the audio quality level and video quality level of the target audio and video.
[0150] The low quality cause analysis module is used to determine the low quality cause of the target audio and video according to the audio quality level and video quality level of the target audio and video and the quality score of the target audio and video.
[0151] Specifically, the audio and video quality rating model is a machine learning model that analyzes the objective quality of audio and video based on EEG signals. Specifically, a convolutional neural network is used to extract the spatiotemporal features of EEG signals and, through a regression layer, predict the independent quality levels of audio and video. This model captures users' subconscious perception of audio and video through neurophysiological signals, compensating for individual differences in subjective ratings.
[0152] The low-quality cause analysis module establishes a mapping rule system between quality levels and comprehensive scores. Specifically, it uses a threshold comparison method. When the audio quality level falls below a preset threshold, audio distortion is determined to be the primary cause. When the video quality level does not meet the standard, video distortion is determined to be the primary cause. When both are abnormal, a composite distortion is determined. This quantitative analysis enables precise tracing of quality defects.
[0153] The audio and video material processing module 100 separates the target audio and video into independent audio and video tracks, and the EEG signal acquisition module synchronously records the neural response signals of the user when watching. In the audio and video evaluation module, the multimodal model performs feature fusion on the two tracks to generate an overall quality score, and the quality rating model parses the alpha wave energy characteristics that represent audio clarity and the theta wave energy characteristics that reflect video smoothness from the EEG signals, and outputs independent audio and video quality levels respectively. The low-quality cause analysis module establishes a correspondence between the quality level and the comprehensive score. When the audio quality level is lower than the set threshold, it triggers the audio defect judgment, and when the video quality level does not meet the standard, it triggers the video defect judgment. When both are triggered at the same time, a composite defect report is generated.
[0154] The embodiment of the present application constructs a two-dimensional evaluation system for audio and video quality by integrating EEG signal analysis and multimodal feature analysis. It achieves accurate positioning of audio and video quality problems and can clearly distinguish quality defects caused by audio distortion, video distortion or combined distortion. In a video conferencing scenario, when it is detected that the audio quality level is continuously lower than the threshold, the audio encoding parameters can be optimized first; in a streaming media transmission scenario, when the video quality level is abnormal, the video compression strategy can be adjusted in a targeted manner. This solution effectively solves the problem of difficulty in tracing the source of quality defects in multimodal scenarios, and provides a clear direction for the optimization and improvement of audio and video systems.
[0155] In some embodiments, an audio and video compression module is further included, and the audio and video compression module is used to:
[0156] Determine the target bit rate of the target audio and video based on the quality score of the target audio and video.
[0157] Compress the target audio and video according to the target bit rate and send the compressed audio and video.
[0158] Among them, the lower the quality score of audio and video, the higher the corresponding target bit rate.
[0159] Specifically, the quality score refers to the quantitative assessment of the perceived quality of audio and video material using a multimodal audio and video experience quality evaluation model. This can be achieved using a normalized numerical range or a graded scoring system, reflecting the user's perception of the overall quality of the audio and video. The target bitrate refers to a compression parameter that is dynamically adjusted based on the quality score. This can be achieved using a preset linear mapping rule or a nonlinear mapping table. For example, the quality score is divided into multiple intervals and a corresponding bitrate threshold is assigned to each interval, so that the compression process matches the corresponding bitrate level based on the score.
[0160] The multimodal audio and video quality of experience evaluation model calculates the quality score of the target audio and video, reflecting the impact of audio and video distortion levels on the user experience. Based on the inverse correlation between quality score and bitrate, audio and video with low quality scores are assigned a higher target bitrate to reduce information loss during compression and avoid further deterioration of the user experience due to secondary distortion. Audio and video with high quality scores are assigned a lower target bitrate to reduce data redundancy and optimize transmission efficiency while ensuring basic quality. The audio and video compression module performs encoding operations based on the target bitrate, generating compressed audio and video data and sending it to the target device.
[0161] For example, the bitrate mapping rule may be in the form of a piecewise function, for example, when the quality score is lower than 0.3, the highest bitrate level is adopted, when the score is between 0.3 and 0.7, the middle bitrate level is adopted, and when the score is higher than 0.7, the lowest bitrate level is adopted.
[0162] By introducing a dynamic correlation mechanism between quality score and bitrate, the system proactively identifies low-quality audio and video during the compression phase and implements protective compression strategies, effectively suppressing the quality avalanche effect caused by over-compression while achieving more efficient bandwidth utilization for high-quality audio and video. This solves the problem of further degradation in the quality of experience caused by compression during the transmission of low-quality audio and video, optimizes the transmission efficiency of high-quality audio and video, and achieves a dynamic balance between quality preservation and bandwidth utilization.
[0163] In some embodiments, a video synthesis module and a video synthesis model fine-tuning module are also included.
[0164] The video synthesis module is used to:
[0165] Input images and audio into the video synthesis model to generate target audio and video.
[0166] The video synthesis model fine-tuning module is used to:
[0167] According to the audio quality level and the video quality level of the target audio video and the quality score of the target audio video, the model parameters of the video synthesis model are fine-tuned multiple times until the audio quality level and the video quality level of the target audio video and the quality score of the target audio video obtained by using the video synthesis model all reach the target audio quality level, the target video quality level and the target quality score.
[0168] Specifically, the video synthesis model refers to an audio video generation algorithm model constructed based on a deep learning framework, and can be specifically implemented by using a generative adversarial network or an autoregressive model, and functions to convert discrete images and audio inputs into time-synchronized audio video streams. The video synthesis model fine-tuning module refers to a functional unit containing a parameter optimization algorithm, and can be specifically implemented by using a back propagation combined with a gradient descent method, and generates a parameter adjustment signal by analyzing the deviation amount of the quality evaluation result and the target value. The audio quality level and the video quality level refer to quantified indexes output by an independent modal evaluation network, and can be specifically implemented by using a classification network output level probability distribution, and are used to respectively reflect the distortion degrees of an audio track and a video track. The quality score refers to a comprehensive evaluation value output by a multi-modal audio video experience quality evaluation model, and can be specifically implemented by using a regression network output scalar value, and is used to represent the overall perception quality under the collaborative action of audio and video.
[0169] The quality of audio video synthesis is controllable by constructing a closed-loop optimization system. After the original synthesis model generates an audio video, the audio quality evaluation network and the video quality evaluation network output the quality levels of independent modalities respectively, and the multi-modal audio video experience quality evaluation model outputs a comprehensive quality score. The three evaluation indexes jointly constitute a quality feedback signal to drive the video synthesis model fine-tuning module to calculate the adjustment direction and amplitude of the model parameters. In each iteration process, the video synthesis model fine-tuning module compares the current quality indicators with the preset target values, and if any indicator does not meet the requirements, the convolution kernel weights and the attention mechanism parameters of the synthesis model are updated according to the error gradient. For example, when the audio quality level is lower than the target value, the video synthesis model fine-tuning module will focus on adjusting the neural network layer parameters of the audio generation path; when the quality score does not meet the requirements but the single-modal quality meets the requirements, the parameters of the cross-modal feature fusion module are adjusted to optimize the audio video synchronization effect. Through multiple iteration updates, the video synthesis model gradually adapts to the multi-dimensional quality constraint conditions, and finally generates an audio video content that meets the requirements of audio quality, video quality and comprehensive experience quality.
[0170] This embodiment of the application introduces a multi-dimensional quality feedback mechanism that simultaneously considers the synergistic impact of single-modality quality levels and cross-modality quality scores during model fine-tuning. This allows the optimization process to simultaneously correct for defects in independent modalities and improve audio and video perceptual consistency. In video conferencing scenarios, the quality score dynamically balances the optimization strength of lip synchronization parameters and noise reduction modules, achieving quality improvements that are more consistent with human perception. This effectively addresses the issue of poor matching between audio and video synthesis quality and multimodal evaluation criteria.
[0171] The embodiment of the present application dynamically adjusts the synthesis model parameters according to the real-time quality evaluation results, optimizing the audio and video collaborative performance while ensuring that the quality of a single modality meets the standard. For example, in the automatic generation of short videos, it can ensure that the generated video maintains high-definition image quality while maintaining precise synchronization between the background music and the rhythm of the picture, and that the overall viewing experience reaches the preset quality threshold. In addition, by establishing a multi-dimensional quality constraint mechanism, the problem of quality imbalance between modalities that may be caused by traditional single-objective optimization is avoided, providing a reliable quality control method for multimodal content generation.
[0172] The above is a detailed introduction to the multimodal audio and video experience quality evaluation system based on multimodal learning provided by this application. This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core idea of this application; at the same time, for those skilled in the art, according to the ideas of this application, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as limiting this application.
Claims
1. A multimodal audio and video experience quality evaluation system based on multimodal learning, characterized by: include: an audio and video material processing module, configured to perform distortion processing on audio and video materials of multiple audio and video types to obtain audio and video materials of multiple quality level combinations of the multiple audio and video types, wherein a quality level combination includes an audio distortion level and a video distortion level; A subjective quality score labeling module is used to play the audio and video materials of each quality level combination of each audio and video type, obtain the user's subjective quality score, and label the audio and video materials of each quality level combination of each audio and video type with the real subjective quality score to obtain multiple labeled audio and video materials; A model training module is configured to capture the bidirectional dependency between the video and audio features of each annotated audio and video material, and to use the annotated audio and video materials to train a multimodal audio and video quality of experience evaluation model. The multimodal audio and video quality of experience evaluation model is configured to comprehensively evaluate the audio and video quality scores from both audio and video perspectives. The model training module is also used to: Using the average EEG signal of the audio and video material of the i-th audio and video type with the j-th video distortion level and the k-th audio distortion level, training an audio and video quality rating model, wherein the audio and video quality rating model is used to evaluate the audio quality level and video quality level of the audio and video; The audio and video material processing module is further used to extract the audio track and video track of the target audio and video; The EEG signal acquisition module is further used to collect EEG signals of the user while watching the target audio and video; an audio and video evaluation module, configured to input the audio track and video track of the target audio and video into the multimodal audio and video experience quality evaluation model to obtain a quality score of the target audio and video, and input the EEG signals of the user during viewing of the target audio and video into the audio and video quality rating model to obtain an audio quality grade and a video quality grade of the target audio and video; The low-quality cause analysis module refers to a mapping rule system that establishes audio quality level, video quality level and quality score, and is used to determine the low-quality cause of the target audio and video based on the audio quality level and video quality level of the target audio and video and the quality score of the target audio and video.
2. The system according to claim 1, wherein The audio and video material processing module is further used to extract the audio track and video track of each annotated audio and video material; The model training module is used to: Performing video feature extraction on the video track of each of the annotated audio and video materials using a video encoder in the model to be trained to obtain video features of each of the annotated audio and video materials; Extracting audio features from the audio track of each annotated audio or video material using the audio encoder in the to-be-trained model to obtain audio features of each annotated audio or video material; By means of the multimodal fusion module in the to-be-trained model, the video features and audio features of each annotated audio and video material are fused with the goal of capturing the bidirectional dependency between the video features and the audio features of each annotated audio and video material, thereby obtaining audio fusion features and video fusion features of each annotated audio and video material; Obtaining a predicted quality score for each of the annotated audio and video materials based on the audio fusion features and the video fusion features of each of the annotated audio and video materials; Based on the predicted quality score of each annotated audio and video material and the annotations carried by the annotated audio and video, the model parameters of the model to be trained are updated to obtain the multimodal audio and video experience quality evaluation model.
3. The system according to claim 2, wherein: Through the multimodal fusion module in the to-be-trained model, with the goal of capturing the bidirectional dependency between the video features and the audio features of each annotated audio and video material, the video features and the audio features of each annotated audio and video material are fused to obtain the audio fusion features and the video fusion features of each annotated audio and video material, including: For each of the video and audio features of the annotated audio and video material, the following steps are performed by the cross attention unit in the multimodal fusion module: Determine the Q, K, and V values of the video features and audio features of the annotated audio and video material by combining the Q transformation matrix, the K transformation matrix, and the V transformation matrix; Calculating an attention score of the video feature to the audio feature based on the Q value of the video feature of the annotated audio and video material and the K value of the audio feature of the annotated audio and video material, and calculating an attention score of the audio feature to the video feature based on the Q value of the audio feature of the annotated audio and video material and the K value of the video feature of the annotated audio and video material; Determining an attention distribution of the video feature to the audio feature based on the attention score of the video feature to the audio feature, and determining an attention distribution of the audio feature to the video feature based on the attention score of the audio feature to the video feature; Obtaining, based on the attention distribution of the video feature to the audio feature and the V value of the audio feature of the annotated audio and video material, an audio feature adjusted based on the video feature; and, based on the attention distribution of the audio feature to the video feature and the V value of the video feature of the annotated audio and video material, obtaining a video feature adjusted based on the audio feature; In combination with a weight matrix for feature transformation, linearly transform the audio features adjusted based on the video features and the video features adjusted based on the audio features respectively; The audio features after linear transformation are determined as the audio fusion features of each annotated audio and video material, and the video features after linear transformation are determined as the video fusion features of each annotated audio and video material.
4. The system according to claim 2, wherein: Based on the predicted quality score of each annotated audio and video material and the annotations carried by the annotated audio and video, the model parameters of the to-be-trained model are updated to obtain the multimodal audio and video experience quality evaluation model, including: The following steps are performed by the gating unit in the multimodal fusion module: Based on the audio fusion features and video fusion features of each annotated audio and video material, generating a gating weight of the audio modality and a gating weight of the video modality, the gating weight of the audio modality is used to: characterize the contribution of the audio track of each annotated audio and video material to the predicted quality score of the annotated audio and video material, and the gating weight of the video modality is used to: characterize the contribution of the video track of each annotated audio and video material to the predicted quality score of the annotated audio and video material; Based on the audio fusion features and video fusion features of each annotated audio and video material, the gating weight of the audio modality and the gating weight of the video modality, a predicted quality score of each annotated audio and video material is obtained.
5. The system according to claim 4, wherein: Also includes: Comparing the actual subjective quality scores of audio and video materials at all audio distortion levels under the same video distortion level to determine the degree of influence of the audio modality on the subjective quality scores; Comparing the actual subjective quality scores of audio and video materials at all video distortion levels under the same audio distortion level to determine the degree of influence of the video modality on the subjective quality scores; Determining initialization weights of the audio modality and the video modality, respectively, based on the degree of influence of the video modality on the subjective quality score and the degree of influence of the audio modality on the subjective quality score; The gating unit is initialized according to respective initialization weights of the audio modality and the video modality.
6. The system according to claim 1, wherein: The subjective quality score marking module is also used to: Play the audio and video material of the i-th audio and video type with the j-th video distortion level and the k-th audio distortion level for N rounds; Obtaining, in the nth round, the subjective quality score of the audio and video material of the ith audio and video type with the jth video distortion level and the kth audio distortion level by the mth user; Averaging across rounds to obtain the average subjective quality score of the mth user for the audio and video material of the ith audio and video type with the jth video distortion level and the kth audio distortion level; Averaging from the user dimension, obtaining the average subjective quality score of the audio and video material of the i-th audio and video type with the j-th video distortion level and the k-th audio distortion level; The value of n ranges from 1 to N, the value of m ranges from 1 to M, the value of i ranges from 1 to I, the value of j ranges from 1 to J, and the value of k ranges from 1 to K.
7. The system according to claim 6, wherein: It also includes an EEG signal acquisition module, which is used to: Obtaining an electroencephalogram signal of the mth user during the nth round when watching the audio and video material of the ith audio and video type with the jth video distortion level and the kth audio distortion level; Averaging the round dimensions to obtain an average EEG signal of the mth user during the process of watching the audio and video material of the ith audio and video type with the jth video distortion level and the kth audio distortion level; Averaging is performed from the user dimension to obtain an average EEG signal of the audio and video material of the i-th audio and video type with the j-th video distortion level and the k-th audio distortion level.
8. The system according to claim 1, wherein: It also includes an audio and video compression module, which is used to: Determining a target bit rate of the target audio and video according to the quality score of the target audio and video; Compressing the target audio and video according to the target bit rate, and sending the compressed audio and video; Among them, the lower the quality score of audio and video, the higher the corresponding target bit rate.
9. The system according to claim 1, wherein: It also includes a video synthesis module and a video synthesis model fine-tuning module; The video synthesis module is used for: Inputting the image and audio into the video synthesis model to generate the target audio and video; The video synthesis model fine-tuning module is used to: According to the audio quality level and video quality level of the target audio and video and the quality score of the target audio and video, the model parameters of the video synthesis model are fine-tuned multiple times until the audio quality level and video quality level of the target audio and video and the quality score of the target audio and video obtained using the video synthesis model all reach the target audio quality level, target video quality level and target quality score.
Citation Information
Patent Citations
Video transmission method, system and device based on visual quality of images
CN101895752A
Method, device and system for monitoring audio and video quality and electronic equipment
CN113382232A
Video playing quality evaluation method based on electroencephalogram characteristics and device thereof
CN113662565A
Audio and video multi-task learning and evaluation method, computer equipment and medium
CN115905613A
Multi-modal sentiment analysis method and system based on multi-dimensional perception
CN119918010A