A video quality evaluation method and device based on interactive fusion of audio and video features
Through the method of interactive fusion of audio and video features, the problem of ignoring the quality of video auditory perception in the prior art is solved, and a more comprehensive video quality evaluation is achieved, which improves the accuracy and comprehensiveness of the evaluation.
Patent Information
- Application Number
- CN202510194992.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-02-21
AI Technical Summary
The prior art focuses on the visual perception quality of videos, ignores the auditory perception quality of videos, resulting in the inability to comprehensively evaluate video quality.
By obtaining the original audio and video data, preprocessing is performed to extract video frame data and audio spectrum data, combining the video feature extraction model and the audio feature extraction model, video frame features and audio frame features are generated, and multi-level interactive fusion is performed through the feature interaction fusion model, and the quality evaluation model is finally input to output the quality evaluation score of audio and video.
The interactive fusion of audio and video features is realized, and the video quality can be evaluated more accurately. It takes into account the space-time characteristics of the video and the spectrum characteristics of the audio, providing a more comprehensive video quality evaluation.
Smart Images

Figure CN119693860B_ABST
Abstract
Description
Technical Field
[0001] The present invention generally relates to the field of video analysis, and more specifically, to a video quality evaluation method and device based on interactive fusion of audio and video features. Background Art
[0002] With the development of mobile Internet technology and the rise of the short video industry, users can easily create and share videos. Videos shot by users themselves are everywhere in people's production and life, and have an important impact on society. At the same time, the booming development of short video services has also led to an exponential growth in the demand for high-quality video services.
[0003] Traditional objective video quality assessment mainly relies on simple image or video frame analysis methods, which usually start from the perspective of signal processing and focus on low-level features such as pixel differences. Most of these early methods are based on the comparison of statistical and physical features, and representative methods include peak signal-to-noise ratio (PSNR), structural similarity (SSIM), video multi-method evaluation fusion (VMAF), etc.
[0004] However, current existing technologies only focus on the visual perception quality of videos, and emphasize on detecting and evaluating the image distortion of videos, thereby ignoring the auditory perception quality of videos, which is not conducive to a comprehensive evaluation of video quality. Summary of the invention
[0005] According to the present invention, a video quality evaluation scheme based on interactive fusion of audio and video features is provided. The scheme fully combines the spectral features of the audio sequence and the spatiotemporal features of the video, and can more accurately evaluate the video quality.
[0006] The present invention provides a video quality evaluation based on interactive fusion of audio and video features. The method comprises:
[0007] Acquire an original data set, where the original data set includes video data and audio data;
[0008] Preprocessing the original data set to obtain video frame data and audio spectrum data;
[0009] Inputting the video frame data into a video feature extraction model to obtain video frame features; and inputting the audio spectrum data into an audio feature extraction model to obtain audio frame features;
[0010] Inputting the video frame features and the audio frame features into a feature interaction fusion model to obtain audio and video interaction fusion frame features;
[0011] The audio and video interactive fusion frame features are input into a quality evaluation model, the quality evaluation score of the audio and video is output, and the quality evaluation score of the audio and video is used as a video quality evaluation result.
[0012] Further, the original data set is preprocessed to obtain video frame data and audio spectrum data, including:
[0013] Extract video frames at equal intervals, and make a difference between adjacent video frames to obtain a video frame difference map, and then divide the video frame and the video frame difference map of each video frame into K image blocks, respectively, to obtain a video frame image block and a video frame difference image block of each video frame, and the video frame image blocks and video frame difference image blocks of all video frames constitute video frame data;
[0014] The audio data is converted into an audio spectrum graph, and then the audio spectrum graph is segmented into audio spectrum images corresponding to the video frame data to obtain audio spectrum data.
[0015] Furthermore, the video frame data is input into the video feature extraction model to obtain the video frame features, including:
[0016] Inputting the video frame image block and the video frame difference image block of each video frame in the video frame data into the first Vision Mamba module respectively, to obtain the video frame image block features and the video frame difference image block features of each video frame;
[0017] The average value of the video frame image block features and the average value of the video frame difference image block features of each video frame are calculated, and then the average value of the video frame image block features and the average value of the video frame difference image block features are concatenated to obtain the video frame features.
[0018] Furthermore, the video frame features and the audio frame features are input into a feature interaction fusion model to obtain audio and video interaction fusion frame features, including:
[0019] The feature interaction fusion model includes three identical Multimodal Transformer modules:
[0020] The input data of the first Multimodal Transformer module is the video frame features , audio frame features , output the first level fusion frame features ;
[0021] The input data of the second Multimodal Transformer module is the video frame features Fusion frame features with the first level The first-level video frame fusion features obtained by addition , audio frame features Fusion frame features with the first level The first-level audio frame fusion features obtained by adding , output the second level fusion frame features ;
[0022] The input data of the third Multimodal Transformer module is the first-level video frame fusion feature Fusion frame features with the second level The second-level video frame fusion features obtained by addition , first-level audio frame fusion features Fusion frame features with the second level The second-level audio frame fusion features obtained by adding , output the third-level fusion frame features , and the third-level fusion frame features As the audio and video interactive fusion frame feature .
[0023] Further, the quality evaluation model includes: a score branch and a weighted branch;
[0024] The score branch includes: a first fully connected layer, a first ReLU activation function, and a second fully connected layer;
[0025] The weighted branch includes: a third fully connected layer, a second ReLU activation function, a fourth fully connected layer, and a first Sigmoid activation function.
[0026] Furthermore, the audio and video interaction fusion frame features Input the quality evaluation model, output the quality evaluation score of the audio and video, and use the quality evaluation score of the audio and video as the video quality evaluation result, including:
[0027] Fusion of audio and video interaction frame features Input the score branch and weighted branch of the quality evaluation model respectively, multiply the output results of the score branch and the weighted branch to obtain the weighted score of each frame, sum the weighted scores of all frames, and output the quality evaluation score of the audio and video as the video quality evaluation result.
[0028] Furthermore, the quality evaluation model is optimized by a mean square error loss function:
[0029] The loss value of the quality evaluation model is calculated through the mean square error loss function, and then the quality evaluation model parameters are updated through the mean square error loss function.
[0030] Furthermore, the mean square error loss function includes: ;
[0031] in, Loss value; is the total number of audio and video; is the i-th audio and video; It is the subjective evaluation score of the dataset. is the quality evaluation score predicted by this method.
[0032] In a second aspect of the present invention, a video quality evaluation device based on interactive fusion of audio and video features is provided. The device comprises:
[0033] An acquisition module, used to acquire an original data set, wherein the original data set includes video data and audio data;
[0034] A preprocessing module, used to preprocess the original data set to obtain video frame data and audio spectrum data;
[0035] A feature extraction module, used to input video frame data into a video feature extraction model to obtain video frame features; and input audio spectrum data into an audio feature extraction model to obtain audio frame features;
[0036] A feature fusion module, used for inputting the video frame features and the audio frame features into a feature interaction fusion model to obtain audio and video interaction fusion frame features;
[0037] The evaluation module is used to input the audio and video interactive fusion frame features into a quality evaluation model, output the quality evaluation score of the audio and video, and use the quality evaluation score of the audio and video as the video quality evaluation result.
[0038] In a third aspect of the present invention, an electronic device is provided. The electronic device has at least one processor; and a memory connected to the at least one processor in communication; the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform the method of the first aspect of the present invention.
[0039] Compared with the prior art, the present invention has the following beneficial technical effects:
[0040] Through the multi-level interactive fusion of audio features and video features, it is possible to mine the semantic and interactive relationships within and between modalities, achieve fine-grained feature fusion, and improve the overall performance of the model.
[0041] It should be understood that the contents described in the summary of the invention are not intended to limit the key or important features of the embodiments of the present invention, nor are they intended to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] The above and other features, advantages and aspects of the embodiments of the present invention will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:
[0043] Figure 1 A flowchart of a video quality assessment method based on interactive fusion of audio and video features according to an embodiment of the present invention is shown;
[0044] Figure 2 A flow chart of a feature interaction fusion model according to an embodiment of the present invention is shown;
[0045] Figure 3 A flow chart of a quality assessment model according to an embodiment of the present invention is shown;
[0046] Figure 4 A block diagram of a video quality assessment device based on interactive fusion of audio and video features according to an embodiment of the present invention is shown;
[0047] Figure 5 A block diagram of an exemplary electronic device capable of implementing embodiments of the present invention is shown.
[0048] Among them, 500 is an electronic device, 501 is a computing unit, 502 is a ROM, 503 is a RAM, 504 is a bus, 505 is an I / O interface, 506 is an input unit, 507 is an output unit, 508 is a storage unit, and 509 is a communication unit. DETAILED DESCRIPTION
[0049] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0050] In addition, the term "and / or" in this article is only a description of the association relationship between the associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.
[0051] In the present invention, both the video frame data and the video frame difference information are input into the video feature extraction model to obtain the video features, the original audio data is converted into a spectrum data form that can be processed by the model and contains rich time information, and is input into the audio feature extraction model to obtain the audio features; the video features and audio features are input into the feature interaction fusion model to obtain the audio and video interaction fusion features; the audio and video features are input into the quality evaluation model to obtain the weighted average audio and video quality score. The above method can realize the mining of semantics and interaction relationships within and between modalities through the multi-level interactive fusion of audio features and video features, realize fine-grained feature fusion, and improve the overall performance of the model.
[0052] Figure 1 A flow chart of a video quality assessment method based on interactive fusion of audio and video features according to an embodiment of the present invention is shown.
[0053] The method includes:
[0054] S101. Acquire an original data set, where the original data set includes video data and audio data.
[0055] This embodiment can obtain multimodal features by integrating video data and audio data, can capture details of more dimensions, and provide more accurate and meaningful results; and when processing multimodal features, a method of interactive fusion of auditory features and visual features is adopted, considering their mutual relationship and potential interaction, which can deeply understand the video quality.
[0056] S102: Preprocess the original data set to obtain video frame data and audio spectrum data.
[0057] In this embodiment, the original data set is preprocessed to obtain video frame data and audio spectrum data, including:
[0058] Video frames are extracted at equal intervals, and adjacent video frames are subtracted to obtain a video frame difference map. The video frame and the video frame difference map of each video frame are then divided into K image blocks to obtain a video frame image block and a video frame difference image block for each video frame. The video frame image blocks and video frame difference image blocks of all video frames constitute video frame data.
[0059] The audio data is converted into an audio spectrum graph, and then the audio spectrum graph is segmented into audio spectrum images corresponding to the video frame data to obtain audio spectrum data.
[0060] Specifically, the video frame difference map is obtained by directly subtracting adjacent video frames, wherein the video frame difference map of the current video frame is obtained by subtracting the current video frame from the video frame of the previous frame; the resolution of the video frame image block and the video frame difference image block is 224*224; and the audio spectrum map is a two-dimensional spectrum map.
[0061] In this embodiment, the complete audio data is converted into a two-dimensional spectrogram by short-time Fourier transform, and the two-dimensional spectrogram is then cropped into audio spectrum image blocks corresponding to the video frame data. All audio spectrum image blocks constitute the audio spectrum data, wherein the resolution of the audio spectrum image blocks is 224*224.
[0062] Considering that the quality of video motion representation is very important for video quality perception, this embodiment helps to supplement the temporal information of video frame data and obtain richer visual features through feature extraction of video frame difference maps; the fusion of video frames and video frame difference maps helps to capture the long-term spatiotemporal dependency information of video sequences and improve the model accuracy.
[0063] This embodiment converts the audio data into an audio spectrogram by converting the entire sequence into frequency domain information and then cutting it into blocks, which can obtain complete time series information, help obtain complete spatiotemporal information, and supplement visual features; at the same time, the audio spectrogram can obtain rich global speech information from the perspective of the frequency domain, and the global speech information includes implicit and non-implicit semantic information.
[0064] S103, inputting the video frame data into a video feature extraction model to obtain video frame features; and inputting the audio spectrum data into an audio feature extraction model to obtain audio frame features.
[0065] In this embodiment, the video frame data is input into the video feature extraction model to obtain the video frame features, including:
[0066] The video frame image block and the video frame difference image block of each video frame in the video frame data are respectively input into the first Vision Mamba module to obtain the video frame image block features and the video frame difference image block features of each video frame.
[0067] The average value of the video frame image block features and the average value of the video frame difference image block features of each video frame are calculated, and then the average value of the video frame image block features and the average value of the video frame difference image block features are concatenated to obtain the video frame features.
[0068] Specifically, the video frame feature is formed by the concatenation of the average value of the video frame image block features of all video frames and the average value of the video frame difference image block features.
[0069] Video frame features can obtain rich spatial information, and frame difference features can obtain temporal information and motion information between frames. The two complement each other to capture the spatiotemporal characteristics of the video and effectively reflect the visual quality perception information of the video.
[0070] In this embodiment, inputting the audio spectrum data into the audio feature extraction model to obtain the audio frame features includes: inputting the audio spectrum data into the second Vision Mamba module to obtain the audio frame features.
[0071] Specifically, the working steps of the Vision Mamba module are as follows:
[0072] Step 1: linearly map the input two-dimensional audio spectrum data or image blocks in the video frame data into a one-dimensional feature vector.
[0073] Step 2: Add position embedding and labeling to the vector obtained in step 1.
[0074] Step 3: Input the vector marked in step 2 into the Vision Mamba model to obtain the audio frame features.
[0075] The Vision Mamba model includes several identical encoders, which are connected in series in sequence; the input data of the first encoder is the vector after adding the mark, the input data of other encoders is the calculation result of the previous encoder, and the calculation result of the last encoder is used as the audio frame feature. The encoder includes a normalization layer, a one-dimensional convolution layer and a bidirectional state space model SSM (State Space Models) in sequence.
[0076] S104, inputting the video frame features and the audio frame features into a feature interaction fusion model to obtain audio and video interaction fusion frame features, including:
[0077] like Figure 2 As shown, the feature interaction fusion model includes three identical Multimodal Transformer modules:
[0078] The input data of the first Multimodal Transformer module is the video frame features , audio frame features , output the first level fusion frame features .
[0079] The input data of the second Multimodal Transformer module is the video frame features Fusion frame features with the first level The first-level video frame fusion features obtained by addition , audio frame features Fusion frame features with the first level The first-level audio frame fusion features obtained by adding , output the second level fusion frame features .
[0080] The input data of the third Multimodal Transformer module is the first-level video frame fusion feature Fusion frame features with the second level The second-level video frame fusion features obtained by adding , first-level audio frame fusion features Fusion frame features with the second level The second-level audio frame fusion features obtained by adding , output the third-level fusion frame features , and the third-level fusion frame features As the audio and video interactive fusion frame feature .
[0081] This embodiment, through the interactive fusion of video and audio modalities, not only takes into account the internal features of each modality, but also can effectively mine the internal features and cross-modal relationships to obtain common semantic information between modalities; at the same time, multi-level cross-modal fusion realizes multi-level fine-grained fusion between modalities, which can increase the fusion depth of features and improve the overall performance of the model.
[0082] S105: Input the audio and video interactive fusion frame features into a quality evaluation model, output the quality evaluation score of the audio and video, and use the quality evaluation score of the audio and video as the video quality evaluation result.
[0083] In this embodiment, the quality evaluation model includes: a score branch and a weighted branch.
[0084] The score branch includes: a first fully connected layer, a first ReLU activation function, and a second fully connected layer.
[0085] The weighted branch includes: a third fully connected layer, a second ReLU activation function, a fourth fully connected layer, and a first Sigmoid activation function.
[0086] Specifically, the input dimension of the first and third fully connected layers is 384, and the output dimension is 128; the input dimension of the second and fourth fully connected layers is 128, and the output dimension is 1; the function of the ReLU activation function is to perform a nonlinear transformation on the input data; the function of the Sigmoid activation function is to map the input data to the range of (0, 1) to obtain the quality weight of each frame.
[0087] The quality evaluation model of this embodiment can effectively integrate the different impacts of different frames on the overall quality of audio and video, provide more reasonable quality scores, and improve the overall performance of the model.
[0088] In this embodiment, if Figure 3 As shown, the audio and video interaction is fused with frame features Input the score branch and weighted branch of the quality evaluation model respectively, multiply the output results of the score branch and the weighted branch to obtain the weighted score of each frame, sum the weighted scores of all frames, and output the quality evaluation score of the audio and video as the video quality evaluation result.
[0089] Among them, audio and video interactive fusion frame features is a feature set. By traversing the feature set and inputting each data in the set into the quality evaluation model, a weighted score for each frame can be obtained.
[0090] The quality evaluation model of this embodiment is more in line with the visual and auditory saliency perception characteristics of the human eye and ear. By weighting different frames, it can more accurately predict the perceived score of each frame for the overall audio and video quality, provide more accurate prediction results, and improve overall prediction accuracy.
[0091] In this embodiment, the quality evaluation model can be optimized by using a mean squared error loss function (MSE): the loss value of the quality evaluation model is calculated by using the mean squared error loss function, and then the quality evaluation model parameters are updated by using the mean squared error loss function.
[0092] Specifically, the mean square error loss function includes: ;
[0093] in, Loss value; is the total number of audio and video; is the i-th audio and video; It is the subjective evaluation score of the dataset. is the quality evaluation score predicted by this method.
[0094] Specifically, the quality evaluation model parameters are updated through the mean square error loss function, including:
[0095] The Adam algorithm optimizes the model through multiple iterations of the training data to reduce the loss function value of the network prediction quality score, so that the network predicted image quality score is close to the subjective quality score; when the training iteration loss function does not change significantly, the model parameters are saved, the training ends, and the model training is completed. The Adam algorithm optimizes the model, and the initial learning rate is 10 -5 , the decay factor is 0.1, it decays every 10 epochs, and the batch size is 12.
[0096] In the mass regression task, the mean square error calculation is simple, intuitive and easy to implement, has high computational efficiency, and has good robustness.
[0097] According to the embodiments of the present invention, the present invention has the following advantages and effects compared with the prior art:
[0098] (1) The complementary combination of frame difference features and frame-level features can obtain the spatiotemporal information of the video, provide rich visual quality perception features, and improve the overall performance of the model.
[0099] (2) By utilizing the complete time series of audio information, we can mine rich audio semantic information. At the same time, by analyzing the audio information from the perspective of the frequency domain, we can mine rich global audio features that complement the visual features.
[0100] (3) The multi-level interactive fusion of audio features and video features can realize the mining of semantic and interactive relationships within and between modalities, achieve fine-grained feature fusion, and improve the overall performance of the model.
[0101] (4) The weighted average quality evaluation model can effectively obtain the quality perception of the overall audio and video for each video frame and audio spectrum, provide more accurate quality score evaluation accuracy, and improve the overall prediction accuracy of the model.
[0102] It should be noted that, for the above-mentioned method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should know that the present invention is not limited by the described order of actions, because according to the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily required by the present invention.
[0103] The above is an introduction to a method embodiment. The following is a further explanation of the scheme of the present invention through a device embodiment having the same inventive concept as the method in the aforementioned embodiment.
[0104] like Figure 4 As shown, the device 400 includes:
[0105] The acquisition module 410 is used to acquire an original data set, where the original data set includes video data and audio data.
[0106] The preprocessing module 420 is used to preprocess the original data set to obtain video frame data and audio spectrum data.
[0107] The feature extraction module 430 is used to input the video frame data into the video feature extraction model to obtain the video frame features; and input the audio spectrum data into the audio feature extraction model to obtain the audio frame features.
[0108] The feature fusion module 440 is used to input the video frame features and the audio frame features into a feature interaction fusion model to obtain audio and video interaction fusion frame features.
[0109] The evaluation module 450 is used to input the audio and video interactive fusion frame features into a quality evaluation model, output the quality evaluation score of the audio and video, and use the quality evaluation score of the audio and video as the video quality evaluation result.
[0110] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the described module can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0111] In the technical solution of the present invention, the acquisition, storage and application of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0112] According to an embodiment of the present invention, the present invention further provides an electronic device and a readable storage medium.
[0113] Figure 5 A schematic block diagram of an electronic device 500 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.
[0114] The electronic device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the electronic device 500 can also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0115] Multiple components in the electronic device 500 are connected to the I / O interface 505, including: an input unit 506, such as a keyboard, a mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a disk, an optical disk, etc.; and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows the electronic device 500 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0116] The computing unit 501 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 501 performs the various methods and processes described above, such as methods S101 to S105. For example, in some embodiments, methods S101 to S105 may be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as a storage unit 508. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of the methods S101 to S105 described above may be performed. Alternatively, in other embodiments, the computing unit 501 may be configured to execute methods S101 - S105 in any other appropriate manner (eg, by means of firmware).
[0117] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0118] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partially on the machine, partially on the machine as a stand-alone software package and partially on a remote machine, or entirely on a remote machine or server.
[0119] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps described in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution of the present invention can be achieved, and this document does not limit this.
[0120] The above specific implementations do not constitute a limitation on the protection scope of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A video quality evaluation method based on interactive fusion of audio and video features, characterized in that: include: Acquire an original data set, where the original data set includes video data and audio data; The original data set is preprocessed to obtain video frame data and audio spectrum data, including: Extract video frames at equal intervals, and make a difference between adjacent video frames to obtain a video frame difference map, and then divide the video frame and the video frame difference map of each video frame into K image blocks, respectively, to obtain a video frame image block and a video frame difference image block of each video frame, and the video frame image blocks and the video frame difference image blocks of all video frames constitute video frame data; convert the audio data into an audio spectrum map, and then divide the audio spectrum map into audio spectrum images corresponding to the video frame data to obtain audio spectrum data; Inputting the video frame data into a video feature extraction model to obtain video frame features; and inputting the audio spectrum data into an audio feature extraction model to obtain audio frame features; The step of inputting the video frame data into the video feature extraction model to obtain the video frame features includes: Inputting the video frame image block and the video frame difference image block of each video frame in the video frame data into the first Vision Mamba module respectively, obtaining the video frame image block features and the video frame difference image block features of each video frame; calculating the average value of the video frame image block features and the average value of the video frame difference image block features of each video frame, and then splicing the average value of the video frame image block features and the average value of the video frame difference image block features to obtain the video frame features; The video frame features and the audio frame features are input into a feature interaction fusion model to obtain audio and video interaction fusion frame features, including: the feature interaction fusion model includes three identical Multimodal Transformer modules: the input data of the first Multimodal Transformer module is the video frame features , audio frame features , output the first level fusion frame features ; The input data of the second Multimodal Transformer module is the video frame feature Fusion frame features with the first level The first-level video frame fusion features obtained by addition , audio frame features Fusion frame features with the first level The first-level audio frame fusion features obtained by adding , output the second level fusion frame features ; The input data of the third Multimodal Transformer module is the first-level video frame fusion feature Fusion frame features with the second level The second-level video frame fusion features obtained by addition , first-level audio frame fusion features Fusion frame features with the second level The second-level audio frame fusion features obtained by adding , output the third-level fusion frame features , and the third-level fusion frame features As the audio and video interactive fusion frame feature ; The audio and video interactive fusion frame features are input into a quality evaluation model, the quality evaluation score of the audio and video is output, and the quality evaluation score of the audio and video is used as a video quality evaluation result.
2. The method according to claim 1, characterized in that The quality evaluation model includes: a score branch and a weighted branch; The score branch includes: a first fully connected layer, a first ReLU activation function, and a second fully connected layer; The weighted branch includes: a third fully connected layer, a second ReLU activation function, a fourth fully connected layer, and a first Sigmoid activation function.
3. The method according to claim 2, characterized in that The audio and video interaction is fused with frame features Input the quality evaluation model, output the quality evaluation score of the audio and video, and use the quality evaluation score of the audio and video as the video quality evaluation result, including: Fusion of audio and video interaction frame features Input the score branch and weighted branch of the quality evaluation model respectively, multiply the output results of the score branch and the weighted branch to obtain the weighted score of each frame, sum the weighted scores of all frames, and output the quality evaluation score of the audio and video as the video quality evaluation result.
4. The method according to claim 3, characterized in that It also includes optimizing the quality evaluation model through a mean square error loss function: The loss value of the quality evaluation model is calculated through the mean square error loss function, and then the quality evaluation model parameters are updated through the mean square error loss function.
5. The method according to claim 4, characterized in that The mean square error loss function includes: ; in, Loss value; is the total number of audio and video; is the i-th audio and video; It is the subjective evaluation score of the dataset. is the quality evaluation score predicted by this method.
6. A video quality evaluation device based on interactive fusion of audio and video features, characterized in that: include: An acquisition module, used to acquire an original data set, wherein the original data set includes video data and audio data; The preprocessing module is used to preprocess the original data set to obtain video frame data and audio spectrum data, including: Extract video frames at equal intervals, and make a difference between adjacent video frames to obtain a video frame difference map, and then divide the video frame and the video frame difference map of each video frame into K image blocks, respectively, to obtain a video frame image block and a video frame difference image block of each video frame, and the video frame image blocks and the video frame difference image blocks of all video frames constitute video frame data; convert the audio data into an audio spectrum map, and then divide the audio spectrum map into audio spectrum images corresponding to the video frame data to obtain audio spectrum data; A feature extraction module, used to input video frame data into a video feature extraction model to obtain video frame features; and input audio spectrum data into an audio feature extraction model to obtain audio frame features; Inputting the video frame image block and the video frame difference image block of each video frame in the video frame data into the first Vision Mamba module respectively, obtaining the video frame image block features and the video frame difference image block features of each video frame; calculating the average value of the video frame image block features and the average value of the video frame difference image block features of each video frame, and then splicing the average value of the video frame image block features and the average value of the video frame difference image block features to obtain the video frame features; A feature fusion module is used to input the video frame features and audio frame features into a feature interaction fusion model to obtain audio and video interaction fusion frame features, including: the feature interaction fusion model includes three identical MultimodalTransformer modules: the input data of the first Multimodal Transformer module is the video frame feature , audio frame features , output the first level fusion frame features ; The input data of the second Multimodal Transformer module is the video frame feature Fusion frame features with the first level The first-level video frame fusion features obtained by addition , audio frame features Fusion frame features with the first level The first-level audio frame fusion features obtained by adding , output the second level fusion frame features ; The input data of the third Multimodal Transformer module is the first-level video frame fusion feature Fusion frame features with the second level The second-level video frame fusion features obtained by addition , first-level audio frame fusion features Fusion frame features with the second level The second-level audio frame fusion features obtained by adding , output the third-level fusion frame features , and the third-level fusion frame features As the audio and video interactive fusion frame feature ; The evaluation module is used to input the audio and video interactive fusion frame features into a quality evaluation model, output the quality evaluation score of the audio and video, and use the quality evaluation score of the audio and video as the video quality evaluation result.
7. An electronic device comprising at least one processor; and a memory connected to the at least one processor in communication; characterized in that: The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Audio processing control system for audio and video
CN116320575A
Video quality evaluation method, device and equipment based on audio and video feature fusion
CN118646929A