Multi-modal audio and video experience quality evaluation system based on multi-modal learning
Through the multi-modal audio and video experience quality evaluation system, audio and video materials are processed and subjective score marks are performed, and the multi-modal audio and video experience quality evaluation model is trained, which solves the problem that the existing models cannot capture the synergy of audio and video, and achieves efficient and accurate audio and video quality evaluation.
Patent Information
- Application Number
- CN202511007572.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-22
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2045-07-22
AI Technical Summary
The existing audio and video quality evaluation model cannot accurately capture the real quality experience brought by audio and video to users, and the traditional methods are expensive, labor-intensive, and poor scalability, which cannot reflect the synergy between audio and video.
A multimodal audio and video experience quality evaluation system based on multimodal learning is adopted, and a variety of quality level combinations are generated through the audio and video material processing module, combined with subjective quality score annotation and model training, the two-way dependence relationship of audio and video features is captured, and a multimodal audio and video experience quality evaluation model is constructed.
It improves the accuracy and automation level of audio and video quality evaluation, provides efficient and reliable decision-making basis for coding optimization and bandwidth resource allocation of communication systems, and reduces code rate and bandwidth consumption.
Smart Images

Figure CN120508944A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of audio and video quality evaluation, and in particular to a multimodal audio and video experience quality evaluation system based on multimodal learning. Background Art
[0002] With the widespread adoption of multimedia communications, the proportion of video in global mobile data continues to increase annually. In this context, compressing and optimizing video data can reduce data transmission volume, alleviate bandwidth pressure, improve transmission efficiency, and alleviate the pressure of multimedia communications. As users' demands for high-definition video quality and user experience increase, selecting appropriate audio and video experience quality evaluation models can accurately assess the video experience and guide communication systems to optimize encoding and resource allocation, reducing bitrates and conserving bandwidth.
[0003] Existing audio and video quality of experience evaluation models typically use video data as training material, leveraging neural networks to learn video features and then predict audio and video quality of experience scores. However, audio and video quality evaluation models trained solely on video cannot accurately capture the true quality of audio and video experience for users. Summary of the Invention
[0004] This application provides a multimodal audio and video experience quality evaluation system based on multimodal learning to solve the problem that existing video quality evaluation models cannot accurately capture the real quality experience brought to users by audio and video.
[0005] The multimodal audio and video experience quality evaluation system based on multimodal learning includes the following steps: an audio and video material processing module, configured to perform distortion processing on audio and video materials of multiple audio and video types to obtain audio and video materials of multiple quality level combinations of the multiple audio and video types, wherein a quality level combination includes an audio distortion level and a video distortion level; A subjective quality score labeling module is used to play the audio and video materials of each quality level combination of each audio and video type, obtain the user's subjective quality score, and label the audio and video materials of each quality level combination of each audio and video type with the real subjective quality score to obtain multiple labeled audio and video materials; The model training module is used to capture the bidirectional dependency between the video and audio features of each annotated audio and video material, and use the multiple annotated audio and video materials to train a multimodal audio and video experience quality evaluation model. The multimodal audio and video experience quality evaluation model is used to comprehensively evaluate the quality score of audio and video from both audio and video perspectives.
[0006] Compared with the prior art, this application has the following advantages: By separating audio tracks and performing joint distortion processing on multi-scene audio and video materials, combined with parameterized control of constant quality factors and audio bitrates, subjective quality score samples covering multiple quality level combinations are generated, ensuring sample diversity and scene representativeness. By employing multiple rounds of playback, two-dimensional average scoring, and standardized experimental process design, single-shot scoring errors and individual biases are effectively eliminated, resulting in highly reliable subjective scoring labels. By constructing a multimodal audio and video experience quality evaluation model to capture the bidirectional dependencies between audio and video features, a comprehensive assessment of overall audio and video quality is achieved from both audio and video dimensions. Through the collaborative work of three modules: audio and video material processing, subjective quality score annotation, and model training, this approach addresses technical issues such as the inability of traditional single-modal evaluation models to capture audio and video synergy, the high cost of subjective scoring, and the disconnect between objective evaluation and user perception. This approach improves the accuracy, comprehensiveness, and automation of audio and video quality evaluation, providing an efficient and reliable decision-making basis for coding optimization and bandwidth resource allocation in communication systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments of the present application. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0008] Figure 1 A schematic diagram of the structure of a multimodal audio and video quality of experience evaluation system based on multimodal learning provided in an embodiment of the present application is shown; Figure 2 The following is a flowchart of an experiment for subjective quality scoring provided by an embodiment of the present application; Figure 3 A schematic diagram of the structure of a model training module provided in one embodiment of the present application is shown; Figure 4 A schematic diagram of the structure of a video encoder provided by an embodiment of the present application is shown; Figure 5 A schematic diagram of the structure of a cross attention unit provided in one embodiment of the present application is shown; Figure 6 The figure shows the impact of the audio and video quality provided by an embodiment of the present application on the user's subjective score. DETAILED DESCRIPTION
[0009] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0010] In recent years, with the rapid development of communications and artificial intelligence technologies, audio and video data has gradually become a vital resource for transmission and utilization in modern information systems. In the field of communications transmission, the explosive growth of multimedia communication services, such as video conferencing, live streaming, and short video platforms, has led to the continuous transmission and interaction of massive amounts of data in communication systems, placing extremely high demands on communication transmission bandwidth. At the same time, with the popularization of various multimedia communication services, users' demands for the quality and experience of high-definition audio and video content are also increasing. Against this backdrop, how to compress and optimize audio and video data while ensuring a good user experience has become a pressing issue in the multimedia communications field. Selecting an appropriate audio and video quality of experience evaluation model can effectively guide communication systems in coding optimization and resource allocation, thereby reducing bitrates and conserving bandwidth.
[0011] Numerous studies have explored the issue of audio and video quality of experience (QoE) evaluation, particularly in the communications field. Quality of Service (QoS) is often used as an evaluation metric to compress data in communication systems while ensuring a certain level of service quality. However, QoS metrics are calculated using objective formulas based on individual data components, such as each pixel in an image block, and do not accurately reflect the user's actual experience, leading to a certain degree of waste in bitrate and bandwidth. In addition to QoS methods, commonly used objective data evaluation methods share some common issues: first, they often require reference video and cannot independently evaluate quality; second, they also fail to reflect the user's true perception. Traditional audio and video QoE evaluation methods also include user-based QoE (Quality of Experience) evaluation, which uses subjective user perception to evaluate data, such as the commonly used Mean Opinion Score (MOS). However, the MOS method also suffers from high cost, labor-intensiveness, poor scalability, and limited real-time performance. Therefore, there is an urgent need to develop deep learning models for efficient and automated video data QoE evaluation.
[0012] Most existing deep learning-based video experience quality evaluation methods only use video data as training material. The trained neural networks may rely on low-level visual information such as color and brightness to predict quality evaluation scores. Most deep learning-based video quality evaluations still focus on a single modality, for example, only evaluating the quality of a single video.
[0013] This application takes into account that from the user's perspective, what they receive is often multimodal data that contains both video and audio. This type of model uses a single-modal processing method, only independently analyzing the video, and does not fully consider the interaction between the two in user perception. Specifically, by separately controlling the video compression parameters to generate test samples, an evaluation system for joint distortion scenarios has not been established, making it difficult to capture the complex relationship between audio and video modalities. As a result, the existing audio and video quality evaluation model does not adequately consider the synergy between audio and video, and cannot accurately capture the real experience of users in audio and video fusion scenarios.
[0014] Based on the above technical problems, this application proposes a multimodal audio and video experience quality evaluation system based on multimodal learning. Figure 1 The multimodal audio and video experience quality evaluation system based on multimodal learning includes: The audio and video material processing module 100 is used to perform distortion processing on audio and video materials of multiple audio and video types to obtain audio and video materials of multiple quality level combinations of the multiple audio and video types, where a quality level combination includes an audio distortion level and a video distortion level.
[0015] When performing distortion processing on audio and video materials of multiple audio and video types, first extract the audio tracks and video tracks of each audio and video type of audio and video materials, perform multiple video distortion level processing on the video track of each audio and video type of audio and video materials, and perform multiple audio distortion level processing on the audio track of each audio and video type of audio and video materials, to obtain audio and video materials of multiple quality level combinations of multiple audio and video types, where one quality level combination includes an audio distortion level and a video distortion level.
[0016] Specifically, the multiple audio and video types include daytime scenes, nighttime scenes, scenes with animals and humans, and scenes without specific target objects. When processing video distortion levels, the constant rate factor (CRF) is varied to achieve different levels of video quality compression, resulting in multiple video distortion levels. When processing audio distortion levels, the audio bitrate is varied to introduce multiple audio distortion levels. By pairing audio tracks with different audio distortion levels with video tracks with different video distortion levels, multiple quality level combinations for multiple audio and video types are achieved.
[0017] Exemplarily, the audio and video material processing module 100 performs multiple video distortion levels processing on the video track of the original audio and video material of each audio and video type, including: for the video track of the original audio and video material of each audio and video type, the audio and video material processing module 100 performs different levels of video quality compression by changing the constant quality factor; the audio and video material processing module 100 performs multiple audio distortion levels processing on the video track of the original audio and video material of each audio and video type, including: for the audio track of the original audio and video material of each audio and video type, the audio and video material processing module 100 performs different levels of audio quality compression by changing the audio bit rate. For example, by setting the constant quality factor of the video track of each audio and video material of multiple audio and video types to 1, and setting other constant quality factors to 37, 47, and 51 respectively, four video distortion levels are introduced; when performing audio distortion level processing, the audio track is first processed so that the processed audio track and video track remain consistent to ensure audio and video synchronization; then, by setting the bit rate of the audio track of each audio and video material of multiple audio and video types to 64kbps, and setting other bit rates to 32kbps and 16kbps respectively, three audio distortion levels are obtained. For a single audio and video material of a single audio and video type, by combining audio tracks with different audio distortion levels and video tracks with different video distortion levels in pairs, 12 quality level combinations of audio and video materials are obtained. With audio and video materials of 20 audio and video types, 240 quality level combinations are obtained.
[0018] In this embodiment, by extracting the audio and video tracks of audio and video materials of various scene audio and video types respectively, and performing different distortion level processing in combination with specific parameter settings, the synchronization of audio and video is maintained to reduce the impact on the user's perceived experience. By combining the two, subjective quality rating samples covering a wide range of audio and video quality are generated, ensuring the diversity, representativeness and accuracy of the rating data, and ensuring that the trained multimodal audio and video experience quality evaluation model can comprehensively and accurately evaluate the audio and video experience quality.
[0019] The subjective quality score labeling module 200 is used to play the audio and video materials of each quality level combination of each audio and video type, obtain the user's subjective quality score, and label the real subjective quality score for the audio and video materials of each quality level combination of each audio and video type, thereby obtaining multiple labeled audio and video materials.
[0020] Specifically, the user's subjective quality rating refers to collecting users' subjective ratings of audio and video materials of each quality level combination of each audio and video type in multiple rounds under a standard experimental environment; using the subjective ratings of the audio and video materials of each quality level combination of each audio and video type to label the audio and video materials of each quality level combination of each audio and video type, and obtaining multiple labeled audio and video materials.
[0021] In some embodiments, the subjective quality score annotation module 200 is further configured to: For the audio and video material of the i-th audio and video type with the j-th video distortion level and the k-th audio distortion level, play it N times.
[0022] Obtain the subjective quality score of the audio and video material of the i-th audio and video type with the j-th video distortion level and the k-th audio distortion level by the m-th user in the n-th round.
[0023] The average is taken from the round dimension to obtain the average subjective quality score of the mth user for the audio and video material of the i-th audio and video type with the j-th video distortion level and the k-th audio distortion level.
[0024] Averaging from the user dimension, we obtain the average subjective quality score of the audio and video material of the i-th audio and video type with the j-th video distortion level and the k-th audio distortion level.
[0025] The value of n ranges from 1 to N, the value of m ranges from 1 to M, the value of i ranges from 1 to I, the value of j ranges from 1 to J, and the value of k ranges from 1 to K.
[0026] Specifically, the specific values of N, M, I, J, and K are set according to actual needs. By obtaining the subjective quality score of the audio and video material with the jth video distortion level and the kth audio distortion level of the i-th audio and video type for each user in each round, the average subjective quality score of the audio and video material with the jth video distortion level and the kth audio distortion level of the i-th audio and video type, the average subjective quality score of the audio and video material with the jth video distortion level and the kth audio distortion level of the m-th user, and the average subjective quality score of the audio and video material with the jth video distortion level and the kth audio distortion level of the i-th audio and video type are obtained as the true value for training the multimodal audio and video experience quality evaluation model.
[0027] For example, refer to Figure 2, with 10 users, 2 rounds, 20 audio and video types, 4 video distortion levels, and 3 audio distortion levels. None of the 10 users participating in the subjective quality rating had ever participated in a similar experiment. All users were informed of the specific experimental procedures and details before the experiment began. In the same subjective quality rating environment, each user participated in 480 subjective quality ratings, with a short break between each subjective quality rating to maintain the reliability and validity of the ratings. Each subjective quality rating consisted of a first gray screen phase, an audio and video material playback phase, and a second gray screen phase. The first gray screen phase lasted for milliseconds to stabilize the subject's state. After the first gray screen phase, the audio and video material playback phase played 3 seconds of audio and video material to stimulate the user's sensory perception. Afterwards, subtitles appeared in the second gray screen phase to prompt the user to rate the overall quality of the audio and video material on a scale of 1-5. The higher the score, the higher the subjective quality.
[0028] For example, the subjective quality score is recorded as , determine the average subjective quality score of the mth user for the audio and video material of the i-th audio and video type with the j-th video distortion level and the k-th audio distortion level according to the following formula: :
[0029] Among them, i represents the audio and video type, j represents the video distortion level, k represents the audio distortion level, and n represents the number of audio and video material playback rounds. Indicates averaging from the round dimension, It represents the average subjective quality score of the mth user for the audio and video material of the i-th audio and video type with the j-th video distortion level and the k-th audio distortion level.
[0030] The average subjective quality score of the audio and video material of the i-th audio and video type with the j-th video distortion level and the k-th audio distortion level is determined according to the following formula:
[0031] Among them, i represents the audio and video type, j represents the video distortion level, k represents the audio distortion level, and n represents the number of audio and video material playback rounds. Indicates the average from the user dimension, The average subjective quality score of the audio and video materials with the jth video distortion level and the kth audio distortion level of the i-th audio and video type.
[0032] In this embodiment, by playing the same audio and video material multiple times and collecting subjective quality ratings, the randomness and error of individual ratings can be reduced, improving the stability and reliability of subjective quality ratings. Averaging across rounds and users eliminates individual rating differences and fluctuations in individual ratings, resulting in a more objective and representative average subjective quality rating. By flexibly configuring the number of users (M), rounds (N), audio and video type (I), video distortion level (J), and audio distortion level (K), the system can adapt to the needs of experiments of varying scales, ensuring the comprehensiveness and scalability of rating data. Controlling subjective quality ratings reduces external interference and improves rating consistency and comparability. By stabilizing user status through gray screen periods, playing audio and video material for a fixed duration, and maintaining rating reliability through short breaks, users can be guided to focus on audio and video quality assessment, improving rating accuracy and effectiveness. Furthermore, subjective quality ratings quantify user perceptions and provide structured ground truth labels for training multimodal audio and video quality of experience evaluation models.
[0033] By playing the same audio and video material multiple times and collecting subjective quality ratings, combined with dual-dimensional averaging of rounds and users, we can effectively reduce the error in single ratings and obtain more representative subjective quality rating results. Flexible configuration of the number of users, rounds, audio and video types, and distortion level parameters can adapt to the subjective quality rating needs of different scales, ensuring data comprehensiveness and scalability. Through standardized subjective quality rating environment control and scientific scoring process design, users can be guided to focus on evaluation, external interference can be reduced, and the accuracy of subjective quality ratings can be improved. Structured subjective quality rating labels provide a reliable quantitative perceptual benchmark for the training of multimodal audio and video experience quality evaluation models, forming a complete closed loop from data collection to model training.
[0034] The model training module 300 is used to capture the bidirectional dependency between the video features and audio features of each annotated audio and video material, and to use multiple annotated audio and video materials to train a multimodal audio and video experience quality evaluation model. The multimodal audio and video experience quality evaluation model is used to comprehensively evaluate the quality scores of audio and video from both audio and video perspectives.
[0035] Specifically, the bidirectional dependency relationship characterizes the mutual influence and constraint between the audio features in the audio track of each annotated audio and video material and the video features in the audio track of each annotated audio and video material. When using multiple annotated audio and video materials for model training to obtain a multimodal audio and video experience quality evaluation model, by capturing the bidirectional dependency relationship between the video and audio features of each annotated audio and video material for training, the multimodal audio and video experience quality evaluation model can more accurately assess the overall quality of the audio and video.
[0036] The embodiment of the present application proposes an audio and video experience quality evaluation system based on multimodal learning. Through the separation of audio tracks and joint distortion processing of multi-scene audio and video materials, combined with the parametric control of constant quality factor and audio bit rate, subjective quality rating samples covering multiple quality level combinations are generated to ensure sample diversity and scene representativeness; multi-round playback, two-dimensional average rating and standardized experimental process design are adopted to effectively eliminate single rating errors and individual deviations, and obtain highly reliable subjective rating labels; by constructing a multimodal audio and video experience quality evaluation model to capture the two-way dependency relationship of audio and video features, a comprehensive evaluation of the overall audio and video quality from the two dimensions of audio and video is achieved. Through the collaborative work of the three modules of audio and video material processing, subjective quality rating annotation and model training, technical problems such as the inability of traditional single-modal evaluation models to capture the synergy of audio and video, high subjective rating costs, and disconnection between objective evaluation and user perception are solved. The accuracy, comprehensiveness and automation level of audio and video quality evaluation are improved, providing an efficient and reliable decision-making basis for communication system coding optimization and bandwidth resource allocation.
[0037] In some embodiments, the audio and video material processing module 100 is further configured to extract the audio track and video track of each annotated audio and video material.
[0038] The model training module 300 is used to: Through the video encoder in the model to be trained, video features are extracted from the video track of each labeled audio and video material to obtain the video features of each labeled audio and video material.
[0039] The audio encoder in the model to be trained is used to extract audio features from the audio track of each annotated audio and video material to obtain the audio features of each annotated audio and video material.
[0040] The multimodal fusion module in the training model aims to capture the bidirectional dependency between the video and audio features of each annotated audio and video material. The video and audio features of each annotated audio and video material are fused to obtain the fused audio and video features for each annotated audio and video material. Based on these fused audio and video features, a predicted quality score is obtained for each annotated audio and video material.
[0041] Based on the predicted quality score of each annotated audio and video material and the annotations carried by the annotated audio and video, the model parameters of the training model are updated to obtain a multimodal audio and video experience quality evaluation model.
[0042] Specifically, the model training module 300 includes a video encoder, an audio encoder, and a multimodal fusion module. The video encoder is used to extract the video features of each video track with annotated audio and video materials, and the audio encoder is used to extract the audio features of the audio track of each annotated audio and video material. The extracted audio features and video features are simultaneously sent to the multimodal fusion module. The multimodal fusion module captures the bidirectional dependency between video features and audio features, deeply interacts and fuses the two modal features, and generates audio fusion features and video fusion features that contain both video information and audio information. Based on the audio fusion features and video fusion features, the predicted quality score of each annotated audio and video material is predicted. The predicted quality score of the annotated audio and video material and the subjective quality score carried by the annotated audio and video material are used to update the model parameters of the training model to obtain a multimodal audio and video experience quality evaluation model.
[0043] For example, refer to Figure 3 and Figure 4 The video encoder consists of a convolutional neural network backbone, a Transformer module (i.e., a transformer module), and a fully connected layer. The convolutional neural network backbone is used to extract video features for each video track with annotated audio and video material. It consists of five convolutional layers with ReLU activation functions. Each convolutional layer uses a 3*3 convolution kernel with a stride of 2, and is followed by max pooling and batch normalization. The Transformer module is used to extract bidirectional dependencies between video and audio features. It consists of two Transformer encoder layers, each with 8 attention heads and a hidden dimension of 512. The positional encoding in the Transformer encoder layer is used to identify the correlation and mutual influence between video features in the temporal dimension. The fully connected layer is used to adjust the dimensionality of the feature vector to adapt to multimodal fusion and improve the expressiveness of video features. The video encoder efficiently extracts video features, preserves temporal dependencies, and improves the compatibility of video features in multimodal fusion.
[0044] The audio encoder consists of a Mel feature extraction layer, a convolutional neural network layer, and a fully connected layer. The Mel feature extraction layer converts each audio track of annotated audio and video material into a Mel spectrogram, providing an audio representation consistent with human auditory perception. The convolutional neural network layer includes four 3*3 convolution kernels and a ReLU activation function. After batch normalization, it enters the fully connected layer, mapping the audio features to the same dimensions as the video features. The audio features of the video are aligned with the video features, providing a comparable feature space for multimodal fusion.
[0045] Table 1 shows the performance comparison results of the method provided in one embodiment of the present application and other multimodal methods on the test set. In the comparison method, this experiment adopted the ConvLSTM+SVM encoder method and the C3D+CNN-Audio encoder method, wherein the ConvLSTM+SVM encoder method refers to the use of the ConvLSTM network and the support vector machine method to extract video features and audio features respectively, and the C3D+CNN-Audio encoder method refers to the use of the C3D method to extract video features, while using CNN to extract features for the audio Mel; In addition, this experiment also adopted four multimodal fusion methods MCB, MRRF, CBAM and AVFF that are different from cross-attention for performance comparison. It can be seen that the video feature and audio feature extraction method provided by this application is superior to other methods in performance.
[0046] Table 1 Performance comparison of this method and other multimodal methods on the test set
[0047] The multimodal fusion module includes a cross-attention unit and a gating unit. The audio and video features of each audio track of annotated audio and video material are simultaneously fed into the cross-attention unit, which then outputs a feature vector and feeds it into the gating unit. After normalization, the predicted quality score for each annotated audio and video material is obtained. By calculating bidirectional attention between audio and video features, the enhancement effect of audio features on video features and the constraint effect of video features on audio features are dynamically modeled, achieving deep modal interaction.
[0048] The embodiments of the present application achieve a leap from single-modality to multi-modal collaboration in audio and video quality evaluation through the joint extraction of spatiotemporal features of the video encoder, the joint modeling of time-frequency-depth features of the audio encoder, and the bidirectional dependency capture and gated fusion of the multimodal fusion module, significantly improving the model's perception of joint distortion scenarios and providing an efficient and automated decision-making basis for communication system bandwidth optimization.
[0049] Table 2 shows the results of the audio and video ablation experiment provided by an embodiment of the present application. Referring to Table 2, it can be seen that the performance changes after removing the key components in the model, where "w / o" means that the component is removed in the method. The results show that the audio encoder, cross-attention unit, and dynamic gating unit all play a significant role in performance improvement. This also further illustrates the importance of audio in video quality evaluation and the superiority of the cross-attention fusion method.
[0050] Table 2 Audio and video ablation experiment results
[0051] In some embodiments, the multimodal fusion module in the model to be trained is used to fuse the video features and audio features of each annotated audio and video material with the goal of capturing the bidirectional dependency between the video features and audio features of each annotated audio and video material, thereby obtaining audio fusion features and video fusion features of each annotated audio and video material, including: For each annotated audio and video material, the following steps are performed through the cross-attention unit in the multimodal fusion module: The Q transformation matrix, the K transformation matrix, and the V transformation matrix are combined to determine the Q, K, and V values of the video features and audio features of the annotated audio and video material.
[0052] Specifically, the Q-transform matrix is a weight matrix that maps input features to query vectors. It is used to represent the attention that audio features require for video features, or vice versa. The Q-value is the vector obtained by linearly transforming audio or video features using the Q-transform matrix. It represents the information that the current feature hopes to obtain from other features.
[0053] The K transformation matrix is a weight matrix that maps input vectors to key vectors, representing the key information of audio and video features for query matching. The K value is the vector obtained by linearly transforming the audio or video features using the K transformation matrix. It represents the key information of the current feature and is used to match query requirements.
[0054] The V transformation matrix is a weight matrix that maps input vectors to value vectors. It is used to represent the actual content of audio and video features as the attention-weighted output. The V value is the vector obtained by linearly transforming the audio or video features using the V transformation matrix. It represents the actual content of the audio and video features as the attention-weighted output.
[0055] For example, Figure 5 The structure diagram of the cross attention unit provided by an embodiment of the present application is shown. Figure 5 , determine the Q, K, and V values of the video and audio features of the annotated audio and video materials according to the following formula:
[0056]
[0057]
[0058]
[0059]
[0060]
[0061] in, Represents video features, Represents audio features, 、 and They are Q transformation matrix, K transformation matrix and V transformation matrix respectively, and Represent the query matrices calculated from video features and audio features respectively, and Represent the key matrices calculated from video features and audio features, respectively. and Represent the value matrices calculated from video features and audio features respectively.
[0062] According to the Q value of the video feature of the annotated audio and video material and the K value of the audio feature of the annotated audio and video material, the attention score of the video feature to the audio feature is calculated, and, according to the Q value of the audio feature of the annotated audio and video material and the K value of the video feature of the annotated audio and video material, the attention score of the audio feature to the video feature is calculated.
[0063] Specifically, the dot product of the Q value and the K value is used to calculate the attention score to quantify the correlation between audio features and video features, and the V value is used as a weighted target to map the correlation to the actual content to achieve feature fusion.
[0064] For example, the attention score of the video feature to the audio feature is calculated according to the following formula:
[0065] in, represents the attention score of video features to audio features, d is the dimension of the feature, Represents the dot product similarity between the query and the key.
[0066] The attention score of audio features to video features is calculated according to the following formula:
[0067] in, represents the attention score of audio features to video features, d is the dimension of the feature, Represents the dot product similarity between the query and the key.
[0068] According to the attention scores of the video features to the audio features, the attention distribution of the video features to the audio features is determined; and, according to the attention scores of the audio features to the video features, the attention distribution of the audio features to the video features is determined.
[0069] Exemplarily, the attention distribution of video features to audio features is determined according to the following formula:
[0070] The attention distribution of audio features to video features is determined as follows:
[0071] in, represents the attention distribution of video features to audio features, It represents the attention distribution of audio features to video features. The softmax operation is used to normalize the attention distribution into a probability distribution so that the sum of all attention weights is 1.
[0072] According to the attention distribution of video features to audio features and the V value of audio features with labeled audio and video materials, audio features adjusted based on video features are obtained, and, according to the attention distribution of audio features to video features and the V value of video features with labeled audio and video materials, video features adjusted based on audio features are obtained.
[0073] Exemplarily, the audio features adjusted based on the video features and the video features adjusted based on the audio features are determined according to the following formula:
[0074]
[0075] in, Represents the video features adjusted based on audio features, Represents the audio features adjusted based on the video features.
[0076] Combined with the weight matrix for feature transformation, linear transformation is performed on the audio features adjusted based on the video features and the video features adjusted based on the audio features.
[0077] The audio features after linear transformation are determined as audio fusion features of each annotated audio and video material, and the video features after linear transformation are determined as video fusion features of each annotated audio and video material.
[0078] For example, the linear transformation is performed according to the following formula:
[0079]
[0080] Exemplarily, the fused video features and the fused audio features are expressed as follows:
[0081] in, represents the fused video features, represents the fused audio features, Represents video features, Represents audio features, Represents the weight matrix used for feature transformation, and C represents the fused video features And the fused audio features The relationship function between video features and audio features.
[0082] The embodiments of the present application capture the bidirectional dependency between audio features and video features so that the adjusted fusion features simultaneously contain interactive information from both modalities. This allows the trained multimodal audio and video quality of experience evaluation model to not only perceive the impact of video quality on audio perception when performing audio and video quality assessment, but also identify the feedback effect of audio quality on video perception, thereby improving the accuracy of the judgment of the comprehensive audio and video experience quality.
[0083] Furthermore, based on the predicted quality score of each annotated audio and video material and the annotations carried by the annotated audio and video, the model parameters of the training model are updated to obtain a multimodal audio and video experience quality evaluation model, including: Through the gating unit in the multimodal fusion module, the following steps are performed: Based on the audio fusion features and video fusion features of each annotated audio and video material, the gating weights of the audio modality and the gating weights of the video modality are generated. The gating weights of the audio modality are used to: characterize the contribution of the audio track of each annotated audio and video material to the predicted quality score of the annotated audio and video material; the gating weights of the video modality are used to: characterize the contribution of the video track of each annotated audio and video material to the predicted quality score of the annotated audio and video material.
[0084] Based on the audio fusion features and video fusion features of each annotated audio and video material, the gating weight of the audio modality and the gating weight of the video modality, a predicted quality score of each annotated audio and video material is obtained.
[0085] Specifically, the gating unit consists of a fully connected layer and a sigmoid function. The gating unit adaptively adjusts the contribution of each modality by generating gating weights. The gating weight of the audio modality is used to characterize the contribution of the audio track of each annotated audio and video material to the predicted quality score of the annotated audio and video material. The gating weight of the video modality is used to characterize the contribution of the video track of each annotated audio and video material to the predicted quality score of the annotated audio and video material. Based on the audio fusion features and video fusion features of each annotated audio and video material, the gating weight of the audio modality and the gating weight of the video modality, the predicted quality score of each annotated audio and video material is obtained. Based on the predicted quality score of each annotated audio and video material and the annotations carried by the annotated audio and video, the model parameters of the model to be trained are updated to obtain a multimodal audio and video experience quality evaluation model.
[0086] For example, the fused video features And the fused audio features The gating unit in the multimodal fusion module adjusts the fused video features. The gate weights and fused audio features The gating weights are used to make more accurate quality score predictions.
[0087] After the fusion of video features And the fused audio features After average pooling, a two-layer fully connected network with ReLU activation function is input, and then normalized by sigmoid activation function to obtain the fused video features. The gate weights and fused audio features Specifically, it is calculated according to the following formula:
[0088]
[0089] in, Represents the fused video features The gate weight, Represents the fused audio features , sigmoid and ReLU represent different activation functions, FC represents the fully connected layer, and mean represents the calculated average.
[0090] Next, use the fused video features The gate weight And the fused audio features The gate weight , for the fused video features And the fused audio features Perform weighted summation to obtain the final fusion feature vector. The final fusion feature vector is calculated according to the following formula:
[0091] in, Represents the final fused feature vector.
[0092] The final fused feature vector is input into a fully connected header to obtain the predicted quality score, and the MES loss function is used to update the model parameters of the training model to obtain a multimodal audio and video experience quality evaluation model. The specific formula is as follows:
[0093]
[0094] Among them, FC represents a fully connected header, represents the prediction quality score, represents the MES loss function for updating model parameters, Indicates the annotation score carried by each annotated audio or video material.
[0095] The embodiment of the present application implements dynamic modal contribution adjustment through a gating unit, thereby improving the accuracy, robustness, and generalization capability of the multimodal audio and video experience quality evaluation model, and optimizing training efficiency and model interpretability.
[0096] Furthermore, it also includes: The actual subjective quality scores of audio and video materials at all audio distortion levels under the same video distortion level are compared to determine the degree of influence of audio modality on subjective quality scores.
[0097] The actual subjective quality scores of audio and video materials at all video distortion levels under the same audio distortion level are compared to determine the impact of video modality on subjective quality scores.
[0098] According to the influence of the video modality and the influence of the audio modality on the subjective quality score, the initialization weights of the audio modality and the video modality are determined.
[0099] The gating unit is initialized according to the initialization weights of the audio modality and the video modality respectively.
[0100] Specifically, the impact refers to the magnitude of score fluctuations caused by the same modality at the same distortion level. This can be achieved by calculating the standard deviation or range of the scores, quantifying the independent contribution of each modality to the subjective quality score. Initialization weights refer to the initial parameters used in the gating unit to adjust the contribution of the audio and video modalities. This can be achieved by normalizing the ratio of the impact of the two modalities, ensuring that the weight distribution matches the actual effect of the modalities.
[0101] For example, Figure 6 The effect of the audio and video quality provided by an embodiment of the present application on the user's subjective rating is shown in FIG. 1 , where the user's subjective quality rating is compared under the levels of no distortion and severe distortion. Figure 6 As shown. By comparing the real subjective quality score differences of audio and video materials at all audio distortion levels under the same video distortion level, the impact of audio quality changes on user perception can be quantified. When the video quality maintains CRF=37, the greater the drop in score when the audio bitrate drops from 64kbps to 16kbps, the greater the impact of the audio modality on the subjective quality score. Similarly, when the audio bitrate is fixed at 32kbps, the impact of the video modality can be calculated by comparing the score changes corresponding to different video CRF values. The bimodal influence is normalized, for example, the audio influence is divided by the sum of the bimodal influence to obtain the initialization weight of the audio modality, and the remainder is used as the initialization weight of the video modality. Based on this weight, the fully connected layer parameters of the gated unit are initialized, so that the model has a reasonable modal fusion tendency at the beginning of training, avoiding the convergence direction uncertainty problem caused by random initialization.
[0102] For example, for a distortion-free video, adding the original audio increases the subjective score from 0.987 to 0.993, indicating that the original audio slightly improves the user's video quality experience. When audio quality degrades, it negatively impacts user perception, and higher audio distortion levels increase this negative impact. For example, when the audio bitrate is 32kbps, or an audio distortion level of 1, the average subjective score drops to 0.901; when the audio bitrate is 16kbps, or an audio distortion level of 2, the subjective score drops to 0.784.
[0103] For severely distorted videos, adding the original audio increased the average subjective rating from 0.248 to 0.316, demonstrating that it also improves the user experience. However, when audio quality degrades, the improvement is minimal. For example, at an audio distortion level of 1, the subjective rating increased from 0.248 to 0.251. It can even further degrade the experience. For example, at an audio distortion level of 2, the subjective rating decreased from 0.248 to 0.222.
[0104] Therefore, during audio and video quality evaluation, poor audio quality can lead to a decline in the overall quality of experience, while high-quality audio can complement the video information and improve the overall quality of experience. By constructing a modal influence assessment mechanism based on real subjective quality scores and using objective experimental data to derive initialization weights, the initial state of the gating unit is made closer to the actual modal contribution distribution, effectively shortening the model convergence cycle, improving model training efficiency, reducing the number of parameter adjustments during training, and enhancing the consistency between model predictions and user subjective perception.
[0105] In some embodiments, the invention further includes an EEG signal acquisition module, which is used to: Obtain an EEG signal of the mth user during the nth round when watching an audio / video material of the ith audio / video type with the jth video distortion level and the kth audio distortion level.
[0106] Averaging is performed from the round dimension to obtain the average EEG signal of the mth user when watching the audio and video material of the ith audio and video type with the jth video distortion level and the kth audio distortion level.
[0107] Averaging is performed from the user dimension to obtain the average EEG signal of the audio and video material of the i-th audio and video type with the j-th video distortion level and the k-th audio distortion level.
[0108] The model training module 300 is also used to: The average EEG signal of the audio and video material of the i-th audio and video type with the j-th video distortion level and the k-th audio distortion level is used to train an audio and video quality rating model, which is used to evaluate the audio quality level and video quality level of the audio and video.
[0109] Specifically, averaging from the round dimension refers to averaging the EEG signals generated by the same user watching the same audio and video material multiple times. Specifically, it can be implemented by arithmetic averaging or weighted averaging methods to eliminate noise interference caused by random factors during individual viewing. Averaging from the user dimension refers to averaging the EEG responses generated by different users watching the same material. Specifically, it can be implemented by group mean calculation methods to eliminate the impact of individual physiological differences on EEG characteristics. The audio and video quality rating model refers to a cross-modal mapping model constructed based on a neural network. Specifically, it can be implemented by a structure combining a convolutional neural network and a fully connected layer to extract characteristic patterns associated with audio distortion levels and video distortion levels from the average EEG signal.
[0110] For example, when a user watches audio and video materials of a specific quality combination, the EEG signal acquisition module synchronously records its original physiological response data. In the case where each user watches the same audio and video material multiple times, the multiple recorded data of the same user are first averaged to effectively eliminate signal fluctuations caused by attention fluctuations or environmental interference. The average EEG data of different users are further averaged twice to eliminate data deviations caused by differences in individual EEG characteristics. The EEG data set after double averaging can reflect the stable physiological response characteristics of a group of users under a specific quality level combination. The model training module 300 associates the processed EEG data with known audio distortion levels and video distortion levels for training, so that the model can identify the EEG feature change rules corresponding to different quality levels, and establish an objective evaluation standard from EEG signals to audio and video quality levels.
[0111] The embodiment of the present application uses a double averaging mechanism to effectively suppress random noise and individual differences while retaining quality-related features, making the mapping relationship between EEG signals and quality levels more reliable and generalizable. It achieves an objective assessment of audio and video quality levels based on group EEG characteristics, solving the problem of poor stability of traditional subjective scoring methods. By establishing a correlation model between EEG signals and audio distortion levels and video distortion levels, it is possible to accurately identify the specific modes that cause quality degradation, providing a quantifiable evaluation basis for audio and video quality optimization.
[0112] In some embodiments, it also includes: an audio and video evaluation module and a low quality cause analysis module.
[0113] The audio and video material processing module 100 is also used to extract the audio track and video track of the target audio and video.
[0114] The EEG signal acquisition module is also used to collect EEG signals of users while they are watching target audio and video.
[0115] The audio and video evaluation module is used to input the audio track and video track of the target audio and video into the multimodal audio and video experience quality evaluation model to obtain the quality score of the target audio and video, and to input the EEG signals of the user while watching the target audio and video into the audio and video quality rating model to obtain the audio quality level and video quality level of the target audio and video.
[0116] The low quality cause analysis module is used to determine the low quality cause of the target audio and video according to the audio quality level and video quality level of the target audio and video and the quality score of the target audio and video.
[0117] Specifically, the audio and video quality rating model is a machine learning model that analyzes the objective quality of audio and video based on EEG signals. Specifically, a convolutional neural network is used to extract the spatiotemporal features of EEG signals and, through a regression layer, predict the independent quality levels of audio and video. This model captures users' subconscious perception of audio and video through neurophysiological signals, compensating for individual differences in subjective ratings.
[0118] The low-quality cause analysis module establishes a mapping rule system between quality levels and comprehensive scores. Specifically, it uses a threshold comparison method. When the audio quality level falls below a preset threshold, audio distortion is determined to be the primary cause. When the video quality level does not meet the standard, video distortion is determined to be the primary cause. When both are abnormal, a composite distortion is determined. This quantitative analysis enables precise tracing of quality defects.
[0119] The audio and video material processing module 100 separates the target audio and video into independent audio and video tracks, and the EEG signal acquisition module synchronously records the neural response signals of the user when watching. In the audio and video evaluation module, the multimodal model performs feature fusion on the two tracks to generate an overall quality score, and the quality rating model parses the alpha wave energy characteristics that represent audio clarity and the theta wave energy characteristics that reflect video smoothness from the EEG signals, and outputs independent audio and video quality levels respectively. The low-quality cause analysis module establishes a correspondence between the quality level and the comprehensive score. When the audio quality level is lower than the set threshold, it triggers the audio defect judgment, and when the video quality level does not meet the standard, it triggers the video defect judgment. When both are triggered at the same time, a composite defect report is generated.
[0120] The embodiment of the present application constructs a two-dimensional evaluation system for audio and video quality by integrating EEG signal analysis and multimodal feature analysis. It achieves accurate positioning of audio and video quality problems and can clearly distinguish quality defects caused by audio distortion, video distortion or combined distortion. In a video conferencing scenario, when it is detected that the audio quality level is continuously lower than the threshold, the audio encoding parameters can be optimized first; in a streaming media transmission scenario, when the video quality level is abnormal, the video compression strategy can be adjusted in a targeted manner. This solution effectively solves the problem of difficulty in tracing the source of quality defects in multimodal scenarios, and provides a clear direction for the optimization and improvement of audio and video systems.
[0121] In some embodiments, an audio and video compression module is further included, and the audio and video compression module is used to: Determine the target bit rate of the target audio and video based on the quality score of the target audio and video.
[0122] Compress the target audio and video according to the target bit rate and send the compressed audio and video.
[0123] Among them, the lower the quality score of audio and video, the higher the corresponding target bit rate.
[0124] Specifically, the quality score refers to the quantitative assessment of the perceived quality of audio and video material using a multimodal audio and video experience quality evaluation model. This can be achieved using a normalized numerical range or a graded scoring system, reflecting the user's perception of the overall quality of the audio and video. The target bitrate refers to a compression parameter that is dynamically adjusted based on the quality score. This can be achieved using a preset linear mapping rule or a nonlinear mapping table. For example, the quality score is divided into multiple intervals and a corresponding bitrate threshold is assigned to each interval, so that the compression process matches the corresponding bitrate level based on the score.
[0125] The multimodal audio and video quality of experience evaluation model calculates the quality score of the target audio and video, reflecting the impact of audio and video distortion levels on the user experience. Based on the inverse correlation between quality score and bitrate, audio and video with low quality scores are assigned a higher target bitrate to reduce information loss during compression and avoid further deterioration of the user experience due to secondary distortion. Audio and video with high quality scores are assigned a lower target bitrate to reduce data redundancy and optimize transmission efficiency while ensuring basic quality. The audio and video compression module performs encoding operations based on the target bitrate, generating compressed audio and video data and sending it to the target device.
[0126] For example, the bitrate mapping rule may be in the form of a piecewise function, for example, when the quality score is lower than 0.3, the highest bitrate level is adopted, when the score is between 0.3 and 0.7, the middle bitrate level is adopted, and when the score is higher than 0.7, the lowest bitrate level is adopted.
[0127] By introducing a dynamic correlation mechanism between quality score and bitrate, the system proactively identifies low-quality audio and video during the compression phase and implements protective compression strategies, effectively suppressing the quality avalanche effect caused by over-compression while achieving more efficient bandwidth utilization for high-quality audio and video. This solves the problem of further degradation in the quality of experience caused by compression during the transmission of low-quality audio and video, optimizes the transmission efficiency of high-quality audio and video, and achieves a dynamic balance between quality preservation and bandwidth utilization.
[0128] In some embodiments, a video synthesis module and a video synthesis model fine-tuning module are also included.
[0129] The video synthesis module is used to: Input images and audio into the video synthesis model to generate target audio and video.
[0130] The video synthesis model fine-tuning module is used to: According to the audio quality level and video quality level of the target audio and video and the quality score of the target audio and video, the model parameters of the video synthesis model are fine-tuned multiple times until the audio quality level and video quality level of the target audio and video and the quality score of the target audio and video obtained using the video synthesis model all reach the target audio quality level, target video quality level and target quality score.
[0131] Specifically, the video synthesis model refers to an audio and video generation algorithm model built based on a deep learning framework, which can be implemented using a generative adversarial network or an autoregressive model. Its function is to convert discrete image and audio inputs into time-synchronized audio and video streams. The video synthesis model fine-tuning module refers to a functional unit that includes a parameter optimization algorithm, which can be implemented using backpropagation combined with a gradient descent method. The parameter adjustment signal is generated by analyzing the deviation between the quality assessment result and the target value. The audio quality level and the video quality level refer to quantitative indicators output by an independent modal evaluation network, which can be implemented using a hierarchical probability distribution output by a classification network, and are used to reflect the degree of distortion of the audio track and the video track respectively. The quality score refers to the comprehensive evaluation value output by the multimodal audio and video experience quality evaluation model, which can be implemented using a scalar value output by a regression network, and is used to characterize the overall perceptual quality under the synergy of audio and video.
[0132] The quality of audio and video synthesis is controlled by building a closed-loop optimization system. After the original synthesis model generates audio and video, the audio quality assessment network and the video quality assessment network each output a quality rating for each modality, while the multimodal audio and video quality of experience evaluation model outputs a comprehensive quality score. These three evaluation metrics together constitute a quality feedback signal, driving the video synthesis model fine-tuning module to calculate the direction and magnitude of model parameter adjustments. During each iteration, the video synthesis model fine-tuning module compares the current quality metrics with the preset target values. If any metric falls short, the convolutional kernel weights and attention mechanism parameters of the synthesis model are updated based on the error gradient. For example, if the audio quality level falls below the target, the video synthesis model fine-tuning module will focus on adjusting the parameters of the neural network layer in the audio generation path. If the quality score falls short but the quality of a single modality meets the requirements, the parameters of the cross-modal feature fusion module will be adjusted to optimize audio and video synchronization. Through multiple iterative updates, the video synthesis model gradually adapts to multi-dimensional quality constraints, ultimately generating audio and video content that meets audio quality, video quality, and overall quality of experience requirements.
[0133] This embodiment of the application introduces a multi-dimensional quality feedback mechanism that simultaneously considers the synergistic impact of single-modality quality levels and cross-modality quality scores during model fine-tuning. This allows the optimization process to simultaneously correct for defects in independent modalities and improve audio and video perceptual consistency. In video conferencing scenarios, the quality score dynamically balances the optimization strength of lip synchronization parameters and noise reduction modules, achieving quality improvements that are more consistent with human perception. This effectively addresses the issue of poor matching between audio and video synthesis quality and multimodal evaluation criteria.
[0134] The embodiment of the present application dynamically adjusts the synthesis model parameters according to the real-time quality evaluation results, optimizing the audio and video collaborative performance while ensuring that the quality of a single modality meets the standard. For example, in the automatic generation of short videos, it can ensure that the generated video maintains high-definition image quality while maintaining precise synchronization between the background music and the rhythm of the picture, and that the overall viewing experience reaches the preset quality threshold. In addition, by establishing a multi-dimensional quality constraint mechanism, the problem of quality imbalance between modalities that may be caused by traditional single-objective optimization is avoided, providing a reliable quality control method for multimodal content generation.
[0135] The above is a detailed introduction to the multimodal audio and video experience quality evaluation system based on multimodal learning provided by this application. This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core idea of this application; at the same time, for those skilled in the art, according to the ideas of this application, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as limiting this application.
Claims
1. A multimodal audio and video experience quality evaluation system based on multimodal learning, characterized by: include: an audio and video material processing module, configured to perform distortion processing on audio and video materials of multiple audio and video types to obtain audio and video materials of multiple quality level combinations of the multiple audio and video types, wherein a quality level combination includes an audio distortion level and a video distortion level; A subjective quality score labeling module is used to play the audio and video materials of each quality level combination of each audio and video type, obtain the user's subjective quality score, and label the audio and video materials of each quality level combination of each audio and video type with the real subjective quality score to obtain multiple labeled audio and video materials; The model training module is used to capture the bidirectional dependency between the video and audio features of each annotated audio and video material, and use the multiple annotated audio and video materials to train a multimodal audio and video experience quality evaluation model. The multimodal audio and video experience quality evaluation model is used to comprehensively evaluate the quality score of audio and video from both audio and video perspectives.
2. The system according to claim 1, wherein The audio and video material processing module is further used to extract the audio track and video track of each annotated audio and video material; The model training module is used to: Performing video feature extraction on the video track of each of the annotated audio and video materials using a video encoder in the model to be trained to obtain video features of each of the annotated audio and video materials; Extracting audio features from the audio track of each annotated audio or video material using the audio encoder in the to-be-trained model to obtain audio features of each annotated audio or video material; By means of the multimodal fusion module in the to-be-trained model, the video features and audio features of each annotated audio and video material are fused with the goal of capturing the bidirectional dependency between the video features and the audio features of each annotated audio and video material, thereby obtaining audio fusion features and video fusion features of each annotated audio and video material; Obtaining a predicted quality score for each of the annotated audio and video materials based on the audio fusion features and the video fusion features of each of the annotated audio and video materials; Based on the predicted quality score of each annotated audio and video material and the annotations carried by the annotated audio and video, the model parameters of the model to be trained are updated to obtain the multimodal audio and video experience quality evaluation model.
3. The system according to claim 2, wherein: Through the multimodal fusion module in the to-be-trained model, with the goal of capturing the bidirectional dependency between the video features and the audio features of each annotated audio and video material, the video features and the audio features of each annotated audio and video material are fused to obtain the audio fusion features and the video fusion features of each annotated audio and video material, including: For each of the video and audio features of the annotated audio and video material, the following steps are performed by the cross attention unit in the multimodal fusion module: Determine the Q, K, and V values of the video features and audio features of the annotated audio and video material by combining the Q transformation matrix, the K transformation matrix, and the V transformation matrix; Calculating an attention score of the video feature to the audio feature based on the Q value of the video feature of the annotated audio and video material and the K value of the audio feature of the annotated audio and video material, and calculating an attention score of the audio feature to the video feature based on the Q value of the audio feature of the annotated audio and video material and the K value of the video feature of the annotated audio and video material; Determining an attention distribution of the video feature to the audio feature based on the attention score of the video feature to the audio feature, and determining an attention distribution of the audio feature to the video feature based on the attention score of the audio feature to the video feature; Obtaining, based on the attention distribution of the video feature to the audio feature and the V value of the audio feature of the annotated audio and video material, an audio feature adjusted based on the video feature; and, based on the attention distribution of the audio feature to the video feature and the V value of the video feature of the annotated audio and video material, obtaining a video feature adjusted based on the audio feature; In combination with a weight matrix for feature transformation, linearly transform the audio features adjusted based on the video features and the video features adjusted based on the audio features respectively; The audio features after linear transformation are determined as the audio fusion features of each annotated audio and video material, and the video features after linear transformation are determined as the video fusion features of each annotated audio and video material.
4. The system according to claim 2, wherein: Based on the predicted quality score of each annotated audio and video material and the annotations carried by the annotated audio and video, the model parameters of the to-be-trained model are updated to obtain the multimodal audio and video experience quality evaluation model, including: The following steps are performed by the gating unit in the multimodal fusion module: Based on the audio fusion features and video fusion features of each annotated audio and video material, generating a gating weight of the audio modality and a gating weight of the video modality, the gating weight of the audio modality is used to: characterize the contribution of the audio track of each annotated audio and video material to the predicted quality score of the annotated audio and video material, and the gating weight of the video modality is used to: characterize the contribution of the video track of each annotated audio and video material to the predicted quality score of the annotated audio and video material; Based on the audio fusion features and video fusion features of each annotated audio and video material, the gating weight of the audio modality and the gating weight of the video modality, a predicted quality score of each annotated audio and video material is obtained.
5. The system according to claim 4, wherein: Also includes: Comparing the actual subjective quality scores of audio and video materials at all audio distortion levels under the same video distortion level to determine the degree of influence of the audio modality on the subjective quality scores; Comparing the actual subjective quality scores of audio and video materials at all video distortion levels under the same audio distortion level to determine the degree of influence of the video modality on the subjective quality scores; Determining initialization weights of the audio modality and the video modality, respectively, based on the degree of influence of the video modality on the subjective quality score and the degree of influence of the audio modality on the subjective quality score; The gating unit is initialized according to respective initialization weights of the audio modality and the video modality.
6. The system according to claim 1, wherein: The subjective quality score marking module is also used to: Play the audio and video material of the i-th audio and video type with the j-th video distortion level and the k-th audio distortion level for N rounds; Obtaining, in the nth round, the subjective quality score of the audio and video material of the ith audio and video type with the jth video distortion level and the kth audio distortion level by the mth user; Averaging across rounds to obtain the average subjective quality score of the mth user for the audio and video material of the ith audio and video type with the jth video distortion level and the kth audio distortion level; Averaging from the user dimension, obtaining the average subjective quality score of the audio and video material of the i-th audio and video type with the j-th video distortion level and the k-th audio distortion level; The value of n ranges from 1 to N, the value of m ranges from 1 to M, the value of i ranges from 1 to I, the value of j ranges from 1 to J, and the value of k ranges from 1 to K.
7. The system according to claim 6, wherein: It also includes an EEG signal acquisition module, which is used to: Obtaining an electroencephalogram signal of the mth user during the nth round when watching the audio and video material of the ith audio and video type with the jth video distortion level and the kth audio distortion level; Averaging the round dimensions to obtain an average EEG signal of the mth user during the process of watching the audio and video material of the ith audio and video type with the jth video distortion level and the kth audio distortion level; Averaging from the user dimension, obtaining an average EEG signal of the audio and video material of the i-th audio and video type with the j-th video distortion level and the k-th audio distortion level; The model training module is also used to: The audio and video quality rating model is trained using the average EEG signal of the audio and video material with the jth video distortion level and the kth audio distortion level of the i-th audio and video type. The audio and video quality rating model is used to evaluate the audio quality level and video quality level of the audio and video.
8. The system according to claim 7, wherein: Also includes: Audio and video evaluation module and low quality cause analysis module; The audio and video material processing module is further used to extract the audio track and video track of the target audio and video; The EEG signal acquisition module is further used to acquire EEG signals of the user while watching the target audio and video; The audio and video evaluation module is configured to input the audio track and video track of the target audio and video into the multimodal audio and video experience quality evaluation model to obtain a quality score of the target audio and video, and input the EEG signals of the user during viewing of the target audio and video into the audio and video quality rating model to obtain an audio quality grade and a video quality grade of the target audio and video; The low-quality cause analysis module is used to determine the low-quality cause of the target audio and video according to the audio quality level and video quality level of the target audio and video and the quality score of the target audio and video.
9. The system according to claim 8, wherein It also includes an audio and video compression module, which is used to: Determining a target bit rate of the target audio and video according to the quality score of the target audio and video; Compressing the target audio and video according to the target bit rate, and sending the compressed audio and video; Among them, the lower the quality score of audio and video, the higher the corresponding target bit rate.
10. The system according to claim 8, wherein It also includes a video synthesis module and a video synthesis model fine-tuning module; The video synthesis module is used for: Inputting the image and audio into the video synthesis model to generate the target audio and video; The video synthesis model fine-tuning module is used to: According to the audio quality level and video quality level of the target audio and video and the quality score of the target audio and video, the model parameters of the video synthesis model are fine-tuned multiple times until the audio quality level and video quality level of the target audio and video and the quality score of the target audio and video obtained using the video synthesis model all reach the target audio quality level, target video quality level and target quality score.
Citation Information
Patent Citations
Video transmission method, system and device based on visual quality of images
CN101895752A
Overall user experience quality assessment method of VR audio and video
CN108683909A
Method, device and system for monitoring audio and video quality and electronic equipment
CN113382232A
No-reference audio and video quality evaluation method based on gated recurrent neural network
CN113473117A
Video playing quality evaluation method based on electroencephalogram characteristics and device thereof
CN113662565A