Video call user experience quality evaluation method and device, equipment and medium
By obtaining frame image sequences from the called party of a video call, extracting and fusing multimodal features, and combining real video call data with an expert-scored quality evaluation model, we solve the problem that existing methods cannot fully reflect the quality of user experience and achieve more accurate video call quality evaluation.
Patent Information
- Application Number
- CN202510893616.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-09-05
AI Technical Summary
Existing methods for evaluating the user experience quality of video calls rely on traditional subjective scoring mechanisms, which cannot fully reflect the user's subjective experience quality and lack a comprehensive evaluation of video content, fluency, clarity, and fidelity.
By obtaining frame image sequences from the called party of a video call, multimodal features (video content perception, smoothness, clarity, and fidelity) are extracted, and feature fusion is performed using a cross-modal multi-perceptual feature fuser. The system then uses a pre-trained quality evaluation model, combined with video call data under real mobile network conditions and expert subjective scores, for evaluation.
It achieves a more comprehensive and accurate evaluation of the user experience quality of video calls, improves the accuracy and reliability of the evaluation, and can comprehensively consider the characteristics of multiple aspects of video calls.
Smart Images

Figure CN120602640A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of video call technology, and in particular to a method, device, equipment and medium for evaluating the quality of user experience of a video call. Background Art
[0002] With the widespread adoption of mobile internet and the rapid development of 5G technology (fifth-generation mobile communications), the demand for video calls is growing. In China, WeChat, an instant messaging app launched by Tencent, is currently the most widely used app supporting online video calls. In today's era of ultra-high-definition video, people's demands for video calls are no longer limited to simply recognizable images. Instead, they demand higher standards for video smoothness, clarity, and fidelity (fidelity refers to the degree of color distortion in the image). Therefore, evaluating the quality of experience (QoE) of video calls is crucial to ensuring the overall quality of video calls.
[0003] Currently, major telecom operators in China have established mature methods and standards for evaluating network and voice quality for VoNR (Voice over New Radio, i.e., voice calls over 5G networks) and VoLTE (Voice over Long-Term Evolution, i.e., voice calls over 4G LTE networks) video calls. However, comprehensive and unified methods and standards for evaluating the user experience quality of video calls have yet to be established. Furthermore, existing methods for evaluating the user experience quality of video calls rely on traditional subjective scoring mechanisms, which fail to comprehensively reflect the user's subjective experience quality. Summary of the Invention
[0004] In view of this, embodiments of the present invention provide a method, apparatus, device, and medium for evaluating the quality of user experience of a video call, so as to overcome the above-mentioned problems or at least partially solve the above-mentioned problems.
[0005] A first aspect of an embodiment of the present invention provides a method for evaluating the quality of user experience of a video call, which is applied to a video call scenario and is used to evaluate the quality of user experience of a called party in a video call. The method includes: Obtain a frame image sequence from a target video of a called party of a video call; Extracting first multimodal features from the frame image sequence using a plurality of feature extractors, wherein the first multimodal features include at least: a video content perception feature, a video smoothness feature, a video clarity feature, and a video fidelity feature; fusing the first multimodal features through a cross-modal multi-sensory feature fuser to obtain a first multi-sensory feature; The target video is evaluated according to the first multi-perceptual features using a pre-trained quality evaluation model to obtain a predicted score. The quality evaluation model is obtained by training a neural network using real video call data of multiple video calls between the same calling end and different called ends under real mobile network conditions as sample data. The sample data carries subjective scores, which are subjective scores given by multiple experts to the real video call data.
[0006] Optionally, the video content perception feature is obtained by the following steps: Calculating the average value of all pixels in each feature map of multiple frames included in the frame image sequence through a spatial global average pooling layer of a video content-aware feature extractor to obtain an average value feature vector of the frame image sequence; Calculating the standard deviation of all pixels in each feature map of the multiple frames included in the frame image sequence through a spatial global standard deviation pooling layer of a video content-aware feature extractor to obtain a standard deviation feature vector of the frame image sequence; Concatenating the mean value feature vector and the standard deviation feature vector to obtain a concatenated feature vector; Performing dimensionality reduction on the concatenated feature vector using a fully connected neural network of a video content-aware feature extractor to obtain the video content-aware feature; Among them, the video content-aware feature extractor includes: an improved ResNet50-SGAS model and a fully connected neural network connected to the output end of the improved ResNet50-SGAS model; the improved ResNet50-SGAS model is obtained by replacing the fully connected layer and the global average pooling layer in ResNet50 with a spatial global average pooling layer and a spatial global standard deviation pooling layer, and the fully connected neural network includes an input layer, a Dropout layer, and a fully connected layer.
[0007] Optionally, the video fluency feature is obtained by a video fluency feature extractor performing the following steps: determining the number of lost frames of the video call according to frame numbers of the multiple frames of images included in the frame image sequence and the number of original video frames of the calling end of the video call; Determining an average optical flow amplitude of each frame image according to the optical flow between each adjacent frame in the plurality of frame images included in the frame image sequence; determining a frame difference score for each frame image according to a total pixel sum of a frame difference image between each frame image and an adjacent frame image of the frame image in the plurality of frame images included in the frame image sequence; For the frame image sequence, when the average optical flow amplitude of any frame image among the multiple frames included in the frame image sequence is less than a first threshold and the frame difference score is less than a second threshold, determining that a freeze occurs in the target video, and recording the number of freezes in the target video and the duration of each freeze; The number of lost frames, the number of freezes, and the freeze duration are determined as video smoothness features of the target video.
[0008] Optionally, the video definition feature is obtained by a video definition feature extractor performing the following steps: Converting each frame image included in the frame image sequence into a grayscale image, and calculating the variance value of the Laplace transform of the target video through the Laplace operator; Perform convolution calculation on the grayscale image using a Sobel operator to obtain a Sobel gradient energy value of the target video; Calculating the total variation value of the gradient of each frame image in the frame image sequence; The variance value of the Laplace transform, the Sobel gradient energy value, and the total variation value of the gradient are determined as the video clarity feature.
[0009] Optionally, the video fidelity feature is obtained by a video fidelity feature extractor performing the following steps: Determining a color richness value of the target video according to each frame image included in the frame image sequence; Determining a standard deviation value of the saturation of the target video according to each frame image included in the frame image sequence; A standard deviation value of the color richness value and the saturation is determined as the video fidelity feature.
[0010] Optionally, the cross-modal multi-sensory feature fuser includes a cross-modal attention unit and a position encoding unit, and before fusing the first multi-modal features to obtain the first multi-sensory features, includes: Constructing a cross-modal attention unit, which includes: a feature projection layer, a multi-head attention layer, residual connections and layer normalization, and a feedforward network; Constructing a position encoding unit, wherein the position encoding unit uses a sine function and a cosine function to perform position encoding on each frame image in the frame image sequence; Respectively copying the video smoothness feature, the video clarity feature, and the video fidelity feature in the time dimension according to the number of frames of the multiple images included in the frame image sequence; The fusing the first multimodal features to obtain a first multi-sensory feature includes: Performing position coding on each frame image in the frame image sequence by the position coding unit; Through the cross-modal attention unit, the video content perception feature in the first multimodal feature and the video smoothness feature, video clarity feature and video fidelity feature replicated in the time dimension are feature fused to obtain a first multi-perception feature.
[0011] Optionally, before evaluating the target video according to the first multi-perceptual features using a pre-trained quality evaluation model to obtain a prediction score, the method further includes: Real video call data of video calls between the same calling end and different called ends under real mobile network conditions, and extracting real frame image sequences from the real video call data; Obtaining subjective scores from multiple experts on the real video call data; Extracting second multimodal features from the real frame image sequence using the multiple feature extractors, the second multimodal features comprising at least: video content perception features, video smoothness features, video clarity features, and video fidelity features; fusing the second multimodal features by the cross-modal multi-sensory feature fuser to obtain a second multi-sensory feature; Constructing a quality assessment model to be trained, the quality assessment model includes: a GRU layer and a temporal pooling layer. The GRU layer is used to capture long-term dependencies in video calls, and the temporal pooling layer is used to generate a final predicted score for user experience quality through the subjective inspiration of the temporal pooling layer. Inputting the second multi-sensory feature as sample data into the quality assessment model to be trained, and using the subjective score as label data to obtain a prediction score for the sample data output by the quality assessment model to be trained; Setting a prediction effect indicator, wherein the prediction effect indicator is used to evaluate the prediction effect of the quality evaluation model of the video call; The quality evaluation model to be trained is trained through multiple iterations, and it is determined that the training of the quality evaluation model is completed according to the prediction effect index.
[0012] A second aspect of an embodiment of the present invention provides a device for evaluating the quality of user experience of a video call, which is applied to a video call scenario and is used to evaluate the quality of user experience of a called party in a video call. The device includes: A frame image sequence acquisition module is used to acquire a frame image sequence from a target video of a called party of a video call; a feature extraction module, configured to extract first multimodal features from the frame image sequence using a plurality of feature extractors, wherein the first multimodal features include at least: a video content perception feature, a video smoothness feature, a video clarity feature, and a video fidelity feature; a feature fusion module, configured to fuse the first multimodal features through a cross-modal multi-sensory feature fuser to obtain a first multi-sensory feature; A quality assessment module is used to evaluate the target video according to the first multi-perceptual features using a pre-trained quality assessment model to obtain a predicted score; the quality assessment model is obtained by training a neural network using real video call data of multiple video calls between the same calling end and different called ends under real mobile network conditions as sample data, and the sample data carries subjective scores, which are subjective scores given by multiple experts to the real video call data.
[0013] According to a third aspect of an embodiment of the present invention, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the method described in the first aspect.
[0014] According to a fourth aspect of the embodiments of the present invention, a computer-readable storage medium is provided, on which a computer program is stored, wherein when the computer program is executed by a processor, the method described in the first aspect is implemented.
[0015] Beneficial effects of the present invention: An embodiment of the present invention provides a method, apparatus, device and medium for evaluating the quality of user experience of a video call, including: obtaining a frame image sequence from a target video of a called end of a video call; extracting a first multimodal feature from the frame image sequence through multiple feature extractors, the first multimodal feature including at least a video content perception feature, a video fluency feature, a video clarity feature and a video fidelity feature; fusing the first multimodal feature through a cross-modal multi-perceptual feature fuser to obtain a first multi-perceptual feature; evaluating the target video based on the first multi-perceptual feature using a pre-trained quality evaluation model to obtain a prediction score; the quality evaluation model is obtained by training a neural network using real video call data of multiple video calls between the same calling end and different called ends under real mobile network conditions as sample data, the sample data carrying subjective scores, and the subjective scores being subjective scores of multiple experts on the real video call data.
[0016] The technical solution of the present invention captures a frame sequence from the target video of the called party in a video call and extracts multimodal features (including video content perception features, video smoothness features, video clarity features, and video fidelity features). These features are then fused using a cross-modal multi-perceptual feature fuser. Finally, a pre-trained quality assessment model is used to evaluate the performance and generate a predicted score. This method comprehensively considers multiple aspects of a video call, including image content, smoothness, clarity, and color fidelity, enabling a more comprehensive and accurate assessment of the user experience quality of the video call. By training the network using real video call data under real mobile network conditions as sample data and subjective ratings as labels, the network learns how experts score, thereby improving the accuracy and reliability of the video call user experience quality assessment. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0018] In order to more clearly illustrate the technical solution of the present invention, the following briefly introduces the drawings required for use in the description of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0019] Figure 1 This is a flow chart of a method for evaluating the quality of user experience of a video call provided by one embodiment of the present invention; Figure 2 This is a framework diagram of a quality evaluation system in a method for evaluating the quality of video call user experience provided by one embodiment of the present invention; Figure 3 This is an example diagram of each frame image in a frame image sequence in a method for evaluating the quality of user experience of a video call provided by an embodiment of the present invention; Figure 4 This is a schematic diagram of feature fusion in a method for evaluating the quality of user experience of a video call provided by one embodiment of the present invention; Figure 5 This is a schematic diagram of a framework of a device for evaluating the quality of user experience of a video call provided by one embodiment of the present invention. DETAILED DESCRIPTION
[0020] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features therein may be combined with each other.
[0021] In related technologies, the following problems often exist in evaluating the user experience quality of video calls: (1) Existing QoS (Quality of Service) indicators can only reflect network-level performance and cannot directly reflect users’ subjective experience.
[0022] (2) Most existing video quality assessment methods are based on full-reference (FR-VQA, which requires reference to the entire target video to evaluate the quality of the distorted video) or reduced-reference (RR-VQA, which requires reference to part of the target video to evaluate the quality of the distorted video). However, users cannot see the target video when using video calls. Therefore, no-reference (NR-VQA, which only requires the distorted video itself to evaluate its quality) methods are more applicable.
[0023] (3) Existing methods lack comprehensive evaluation of video content, video smoothness, video clarity and fidelity, and cannot fully reflect the user's subjective experience quality.
[0024] Based on this, the present invention proposes a method, device, equipment and medium for evaluating the user experience quality of video calls. In this method, it is possible to monitor, evaluate and ensure the user experience quality of video calls under real mobile network conditions from the aspects of video content, video fluency, video clarity and fidelity, and independently construct a model training based on real video call data with subjective scores. At the same time, based on the user's multi-perceptual features, an improved ResNet50-SGAS model is proposed to extract content-perceptual features of video calls, and an algorithm combining cross-modal attention mechanism and time-embedded position coding is selected to improve the fusion effectiveness of various perceptual features.
[0025] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0026] Figure 1 This is a flow chart of a method for evaluating the quality of user experience of a video call provided by one embodiment of the present invention. Figure 2 This is a framework diagram of a quality evaluation system in a method for evaluating the quality of video call user experience provided by an embodiment of the present invention. Figure 1 The present invention provides a method for evaluating the quality of user experience of a video call, which is applied to a video call scenario and is used to evaluate the quality of user experience of a called party in a video call. The method includes steps S11 to S14: Step S11 : obtaining a frame image sequence from a target video of the called party of the video call.
[0027] refer to Figure 2 The video call user experience quality evaluation method of the present invention is to Figure 2 The quality evaluation system shown is implemented, and the quality evaluation system includes multiple feature extractors (video content perception feature extractor, video smoothness feature extractor, video clarity feature extractor and video fidelity feature extractor), a cross-modal multi-perceptual feature fusion device and a quality evaluation model.
[0028] In this embodiment, in order to evaluate the user experience quality of a video call, it is first necessary to obtain a target video from the called end of the video call. Then, one frame image is extracted from the obtained target video at regular time intervals (for example, every 1 second), and frame numbers are added to each extracted frame image. All frame images containing the frame numbers are used as the frame image sequence of the target video.
[0029] Step S12: extracting first multimodal features from the frame image sequence through multiple feature extractors, where the first multimodal features include at least: video content perception features, video smoothness features, video clarity features, and video fidelity features.
[0030] In this embodiment, in order to comprehensively evaluate the quality of a video call, multiple feature extractors are used to extract multimodal features, namely, first multimodal features, from a frame image sequence. The first multimodal features refer to multimodal features extracted in the process of actually evaluating the user experience quality of a video call.
[0031] Among them, multimodal features include: Video content-aware features: extracted through the improved ResNet50-SGAS model and fully connected neural network, used to reflect the content information of the pictures in the target video.
[0032] Video smoothness features: These features are extracted through frame number recognition, optical flow calculation, and frame difference calculation. They are used to reflect the smoothness of the target video, including the number of lost frames, the number of freezes, and the duration of freezes.
[0033] Video clarity feature: The clarity of each frame is calculated using Laplace transform and Sobel operator to reflect the sharpness and details of the image in the original data.
[0034] Video fidelity features: This feature reflects the color diversity and uniformity of the target video by calculating the standard deviation of color richness and saturation.
[0035] Step S13: fusing the first multimodal features through a cross-modal multi-sensory feature fuser to obtain a first multi-sensory feature.
[0036] In this embodiment, to comprehensively consider the extracted features of different modalities, a multi-sensory feature fuser is used to fuse the extracted first multimodal features. The fusion process of the first multimodal features includes feature projection, a multi-head attention mechanism, residual connections and layer normalization, a feedforward network, and positional encoding. Through these steps, the features of different modalities contained in the first multimodal features are effectively fused to obtain the first multi-sensory features, providing a more comprehensive input for subsequent quality evaluation.
[0037] Step S14: Evaluate the target video based on the first multi-perceptual features using a pre-trained quality evaluation model to obtain a predicted score; the quality evaluation model is obtained by training a neural network using real video call data of multiple video calls between the same calling end and different called ends under real mobile network conditions as sample data, and the sample data carries subjective scores, which are subjective scores of multiple experts on the real video call data.
[0038] In this embodiment, a pre-trained quality assessment model is used to evaluate the fused first multi-sensory features to obtain a predicted score. This quality assessment model is trained using sample data from real video call data collected under real mobile network conditions. This sample data includes target videos from multiple video calls between the same calling end and different called ends, and each sample data includes subjective ratings from multiple experts. This quality assessment model, trained in this way, learns how to score target videos, obtaining a predicted score, thereby more accurately evaluating the user experience quality of video calls.
[0039] Through the above embodiment, a frame image sequence is obtained from the target video of the called end of a video call, and multimodal features (including video content perception features, video smoothness features, video clarity features, and video fidelity features) are extracted. These features are then fused using a cross-modal multi-perceptual feature fuser. Finally, a pre-trained quality assessment model is used to evaluate the quality of the video call experience, resulting in a predicted score. This method comprehensively considers multiple aspects of a video call, including image content, smoothness, clarity, and color fidelity, thereby more comprehensively and accurately evaluating the user experience quality of the video call. By training the network using real video call data under real mobile network conditions as sample data and subjective ratings as labels, the network can learn how experts score, thereby further improving the accuracy and reliability of the video call user experience quality assessment.
[0040] In combination with the above embodiment, the present invention further provides another method for evaluating the quality of user experience of a video call. In this method, the video content perception features are obtained through the following steps S21 to S24: In step S21 , the spatial global average pooling layer of the video content-aware feature extractor is used to calculate the average value of all pixels in each feature map of the multiple frame images included in the frame image sequence to obtain an average value feature vector of the frame image sequence.
[0041] In this embodiment, a video content-aware feature extractor is used to extract video content-aware features. The video content-aware feature extractor includes two parts: an improved ResNet50-SGAS model and a fully connected neural network connected to the output end of the improved ResNet50-SGAS model.
[0042] The improved ResNet50-SGAS model is obtained by replacing the fully connected layers and global average pooling layers in ResNet50 (50 indicates that it has 50 layers and is a residual neural network) with spatial global average pooling layers and spatial global standard deviation pooling layers.
[0043] The frame image sequence is input into the improved ResNet50-SGAS model. Through the spatial global average pooling layer, a feature map is extracted from each frame image in the multiple frames contained in the frame image sequence to obtain multiple feature maps. Then, all pixel values in each extracted feature map are averaged to obtain the pixel average value of each feature map. The pixel average values of all feature maps are then combined to obtain an average feature vector. The average feature vector can reflect the content information of each frame image in the frame image sequence.
[0044] The calculation process of the pixel average of each feature map is shown in the following formula: ; In this formula, N For each feature map (feature map, refers to the image obtained after the convolutional neural network model extracts the frame image features), the spatial dimension (spatial dimension refers to the width and height of the feature map), is the first i pixel values.
[0045] Step S22, calculating the standard deviation of all pixels in each feature map of the multiple frame images included in the frame image sequence through the spatial global standard deviation pooling layer of the video content-aware feature extractor, and obtaining a standard deviation feature vector of the frame image sequence.
[0046] In this embodiment, a spatial global standard deviation pooling layer is used to calculate the standard deviation of all pixel values within each of the multiple feature maps extracted from the multiple frames contained in the frame image sequence. This calculation yields the pixel standard deviation for each feature map. The pixel standard deviations of all feature maps are then combined to form a standard deviation feature vector. The standard deviation feature vector reflects the degree of variation in pixel values within each frame, i.e., the degree of feature discreteness. This allows the capture of details and variation within each frame, providing a richer feature representation for subsequent feature fusion.
[0047] The calculation process of the pixel standard deviation of each feature map is shown in the following formula: ; In this formula, N For each feature map (feature map, refers to the image obtained after the convolutional neural network model extracts the frame image features), the spatial dimension (spatial dimension refers to the width and height of the feature map), is the first i pixel values, Mean is the pixel average of each feature map.
[0048] Through the above two steps, the improved ResNet50-SGAS model can better retain the detailed information of the features by combining spatial global average pooling and spatial global standard deviation pooling, preventing a large number of features of each frame image contained in the frame image sequence from being discarded by global average pooling, thereby retaining the deep change information of each frame image contained in the frame image sequence, so that the obtained features contain more effective information.
[0049] Step S23: concatenate the mean value feature vector and the standard deviation feature vector to obtain a concatenated feature vector.
[0050] In this embodiment, in order to comprehensively consider the average value characteristics and standard deviation characteristics of each frame image contained in the frame image sequence, the average value feature vector and the standard deviation feature vector obtained in the above steps are spliced to obtain a spliced feature vector. The spliced feature vector includes the global average characteristics (represented by the average value feature vector) and change characteristics (represented by the standard deviation feature vector) of each frame image contained in the frame image sequence, which can more comprehensively reflect the feature information contained in the target video.
[0051] For example, suppose there is a 10-second target video, and one frame image is extracted per second to obtain a frame image sequence containing 10 frames of images. Then, the frame image sequence is input into the improved ResNet50-SGAS model to obtain the average value feature vector with a shape of (10, 2048) after passing through the spatial global average pooling layer and the standard deviation feature vector with a shape of (10, 2048) after passing through the spatial global standard deviation pooling layer, where 10 represents the time step of the frame image sequence, and 2048 represents that a 2048-dimensional feature vector is extracted from each frame image in the frame image sequence. Then, the average value feature vector and the standard deviation feature vector are concatenated into a concatenated feature vector with a shape of (10, 4096).
[0052] Step S24, reducing the dimension of the concatenated feature vector by a fully connected neural network of a video content-aware feature extractor to obtain the video content-aware feature; the fully connected neural network includes an input layer, a dropout layer, and a fully connected layer.
[0053] In this embodiment, a fully connected neural network is used to further process the concatenated feature vector. The fully connected neural network includes an input layer, a dropout layer, and a fully connected layer. Specifically, the fully connected neural network may include one input layer, one dropout layer, and two fully connected layers.
[0054] Among them, the input layer is used to receive the spliced feature vector as input, the Dropout layer is used to prevent overfitting, and the fully connected layer is used to perform further feature extraction and dimensionality reduction on the input spliced feature vector, and finally output a more compact feature vector, which is the video content-aware feature, which can more efficiently represent the feature information of the content in the target video and is suitable for subsequent quality evaluation tasks.
[0055] Through the above-mentioned embodiments, the improved ResNet50-SGAS model replaces the original fully connected layer and global average pooling layer with a spatial global standard deviation pooling layer, which can better preserve feature details. This improvement enables the model to more comprehensively capture the characteristics of video content, including not only the global average features but also the degree of feature variation. This provides a richer feature representation for video call quality evaluation, especially under complex network conditions. It can better capture the changes and details of the target video content, thereby more accurately evaluating the user experience quality.
[0056] In combination with the above embodiment, the present invention further provides another method for evaluating the quality of user experience of a video call. In this method, the video fluency feature is obtained by a video fluency feature extractor performing the following steps S31 to S35: Step S31 : determining the number of lost frames of the video call according to the frame numbers of the multiple frames of images included in the frame image sequence and the number of original video frames of the calling end of the video call.
[0057] In this embodiment, in order to evaluate the smoothness of a video call, it is first necessary to determine the number of lost frames during the video call.
[0058] Figure 3 This is an example diagram of each frame image in a frame image sequence in a method for evaluating the user experience quality of a video call provided by an embodiment of the present invention. Specifically, Python can be used to process the target video, extract one frame image from the acquired target video at a certain time interval (for example, every 1 second), and add frame numbers to the extracted frame images respectively. All frame images containing frame numbers are used as the frame image sequence of the target video. For example, the target video has a total of 21 frames, and 5 frames are extracted. The frame numbers corresponding to the extracted frames should be 1, 6, 11, 16, and 21. Therefore, frame numbers can be sequentially added to each frame of the target video, and then the frames can be extracted.
[0059] like Figure 3 As shown, as an example, the frame image includes: a person speaking continuously in the foreground, a dynamic natural landscape full of flowers and plants in the background, the frame number is displayed in the upper left corner of the video screen, and all the target images are valid frame images, and there is no black screen frame (a black screen frame refers to a frame image in which the entire screen is displayed as a single solid black color, and this type of frame image is an invalid frame image).
[0060] Specifically, the number of lost frames is determined by identifying the frame number of each frame in the frame image sequence and summing them to obtain the target video frame number. This number is then combined with the original video frame number from the calling end to calculate the number of lost frames during the video call. The frame number identifies the order of each frame, while the original video frame number serves as the baseline frame number for the video call. By comparing the actual frame number received by the called end with the original video frame number from the calling end, the number of lost frames can be accurately calculated.
[0061] For example, the calculation process of the number of lost frames is shown in the following formula: ; In this formula, M represents the number of original video frames, and N represents the number of target video frames.
[0062] Step S32: determining an average optical flow amplitude of each frame image according to the optical flow between each adjacent frame in the plurality of frame images included in the frame image sequence.
[0063] In this embodiment, the smoothness of the video call is further evaluated by calculating the optical flow between each frame and adjacent frames to determine the average optical flow amplitude of each frame. Optical flow is a method used in computer vision to describe the movement of objects or scenes in an image between consecutive frames.
[0064] Specifically, the optical flow between each frame and adjacent frames can be calculated using the cv2.calcOpticalFlowFarneback function in the OpenCV library in Python. The magnitude and direction of the optical flow are then calculated using the cv2.cartToPolar function. The average magnitude of these optical flow vectors is then calculated to obtain the average optical flow magnitude for each frame. A smaller average optical flow magnitude indicates less motion in that frame, and the video call is likely to be smoother.
[0065] Step S33 , determining a frame difference score for each frame image according to the sum of all pixels of the frame difference image between each frame image and its adjacent frame image in the plurality of frame images included in the frame image sequence.
[0066] In this embodiment, in order to more comprehensively evaluate the smoothness of the video call, it is necessary to calculate the frame difference score between each frame image and the adjacent frames. The frame difference score is obtained by calculating the pixel difference between each frame image and the adjacent frames.
[0067] Specifically, a frame difference algorithm is used to calculate the frame difference image between each frame and the adjacent frames. For example, the cv2.absdiff function in the OpenCV library in Python is used to calculate the frame difference between each frame and the adjacent frames. (Frame difference is a technique used in motion and change detection that primarily identifies areas of motion or change by comparing the differences between adjacent frames in a video sequence.) This ensures that each video frame generates a frame difference image of the same size. All pixel values in the frame difference image are then summed to obtain a frame difference score. A smaller frame difference score indicates less change between the frame and the adjacent frames, resulting in a smoother video call.
[0068] Step S34, for the frame image sequence, when the average amplitude of the optical flow of any frame image in the multiple frame images contained in the frame image sequence is less than the first threshold and the frame difference number is less than the second threshold, it is determined that the target video has a freeze, and the number of freezes of the target video and the duration of each freeze are recorded.
[0069] In this embodiment, a video freeze determiner is constructed, and the video freeze determiner determines whether the target video of the video call freezes according to a preset threshold.
[0070] Specifically, when the average amplitude of the optical flow of a frame image is less than the first threshold and the frame difference score is less than the second threshold, the frame image is considered to have a freeze, and the number of freezes and the duration of each freeze are recorded. In addition, the total freeze time, the accumulated remaining freeze time, etc. are also recorded. The freeze duration is obtained by using the cv2.VideoCapture function of OpenCV in Python to obtain the actual frame rate of the target video and converting it. The first and second thresholds can be set according to actual conditions.
[0071] For example, when the video freeze detector determines that the average amplitude value of the optical flow is less than a first threshold, such as 0.1, and the frame difference score is less than a second threshold, such as 5000, it is determined that the target video has experienced one freeze.
[0072] Then use the cv2.VideoCapture function of OpenCV in Python to obtain the actual frame rate of the target video, convert it to the duration after the freeze occurs, and repeat this step multiple times to calculate: the number of freezes, the first freeze time t1, the second freeze time t2, the third freeze time t3, the fourth freeze time t4, the fifth freeze time or the accumulated remaining freeze time t5, and the total freeze time.
[0073] Among them, the more times the video freezes and the longer the freezes last, the worse the smoothness of the video call.
[0074] Step S35: Determine the number of lost frames, the number of freezes, and the freeze duration as video smoothness features of the target video.
[0075] In this implementation, the number of dropped frames, number of freezes, and duration of freezes calculated in the above steps are combined to form the video fluency feature of the video call. This video fluency feature comprehensively reflects the fluency of the video call and provides an important basis for subsequent quality evaluation.
[0076] The above-described embodiment, through detailed analysis of video call frame image sequences, extracts key features such as the number of lost frames, average optical flow amplitude, frame difference score, number of freezes, and freeze duration. These features together constitute the video fluency feature. These features can accurately assess the fluency of video calls, especially under complex network conditions, effectively detecting freezes and frame drops during video calls, thereby significantly improving the accuracy and reliability of video call quality assessment.
[0077] In combination with the above embodiment, the present invention further provides another method for evaluating the quality of user experience of a video call. In this method, the video clarity feature is obtained by performing the following steps S41 to S44 by a video clarity feature extractor: Step S41 : converting each frame image included in the frame image sequence into a grayscale image, and calculating the variance value of the Laplace transform of the target video through the Laplace operator.
[0078] In this embodiment, to evaluate the clarity of a video call, each frame in the frame sequence is first converted to a grayscale image. This simplifies calculations because color images have three channels (RGB), while grayscale images have only one channel. Next, a Laplace operator is used to convolve the grayscale image, calculating the Laplace transform result. The variance of the Laplace transform is then calculated, which provides a measure of image clarity. A larger variance indicates more pronounced edges and details, and therefore higher clarity.
[0079] The specific formula is as follows: Define a Laplacian operator: , use the operator to perform convolution operation with the grayscale image, and then perform variance calculation to obtain the variance value of Laplace transform;
[0080] In this formula, is the first i The result of the Laplace transform of pixels is is the average value of the Laplace transform of all pixels in the grayscale image, and N is the total number of pixels contained in the grayscale image.
[0081] Step S42: performing convolution calculation on the grayscale image using a Sobel operator to obtain a Sobel gradient energy value of the target video.
[0082] In this embodiment, to further evaluate the clarity of a video call, a Sobel operator is convolved with a grayscale image to obtain the Sobel gradient energy value. The Sobel operator is used to detect image edges and calculate the image's gradient in the x and y directions. The gradient energy value reflects the edge strength and, therefore, the image's clarity.
[0083] The specific formula is as follows: Define the Sobel operators in the x and y directions respectively: 、 , use the Sobel operator to perform convolution operation with the grayscale image, and then calculate the gradient energy of the result of the convolution operation to obtain the Sobel gradient energy value.
[0084]
[0085] in, and are the gradients in the x and y directions respectively.
[0086] Step S43: Calculate the total variation value of the gradient of each frame image in the frame image sequence.
[0087] In this embodiment, to more comprehensively evaluate the clarity of a video call, the total variation of the gradient of each frame is calculated. This total variation reflects the texture complexity of each frame in the sequence. The more complex the texture, the richer the image details.
[0088] The specific calculation formula for the total variation value of the gradient of each frame image is as follows: ; in, I(i,j) Indicates that the image is at the coordinate point (i,j) The pixel value at i 、 j Indicates the horizontal and vertical coordinates.
[0089] Step S44 : determining the variance value of the Laplace transform, the Sobel gradient energy value, and the total variation value of the gradient as the video definition feature.
[0090] In this embodiment, the Laplace transform variance, Sobel gradient energy, and total gradient variation calculated in the above steps are combined to form the video clarity features of the video call. These features comprehensively reflect the clarity of the video call and provide an important basis for subsequent quality evaluation.
[0091] Through the above-described embodiment, by analyzing the video call frame image sequence in detail, key features such as the Laplace transform variance, Sobel gradient energy, and total gradient variation were extracted. Together, these features constitute the video clarity feature. These features can accurately assess the clarity of video calls, effectively detecting blur and distortion, especially under complex network conditions. This method can significantly improve the accuracy and reliability of video call quality assessment.
[0092] In combination with the above embodiment, the present invention further provides another method for evaluating the quality of user experience of a video call. In this method, the video fidelity feature is obtained by performing the following steps S51 to S53 by a video fidelity feature extractor: Step S51 : determining the color richness value of the target video according to each frame image included in the frame image sequence.
[0093] In this embodiment, in order to evaluate the color fidelity of a video call, it is necessary to calculate the color richness value of each frame of the image. The color richness value reflects the diversity and intensity of the colors in the image. The color richness value can be calculated using the following formula: ; in, represents the standard deviation of the red-green channel color component, Represents the standard deviation of the yellow-blue channel color component, represents the standard deviation of the red-green channel color component, Represents the standard deviation of the yellow-blue color component, where yellow is the sum of red and green.
[0094] Step S52: determining a standard deviation value of the saturation of the target video according to each frame image included in the frame image sequence.
[0095] In this embodiment, in order to further evaluate the color fidelity of the video call, it is necessary to calculate the saturation standard deviation value of each frame image. The saturation standard deviation value reflects the color uniformity of each frame image included in the frame image sequence.
[0096] The saturation standard deviation can be calculated using the following formula: ; in, Indicates the i The saturation value of the samples, Indicates the i The maximum RGB value of samples, represents the number of samples, Step S53: Determine the standard deviation of the color richness value and the saturation value as the video fidelity feature.
[0097] In this embodiment, the color richness value and saturation standard deviation value calculated in the above steps are combined to form the video fidelity features of the video call. These two features can comprehensively reflect the color quality of the video call and provide an important basis for subsequent quality evaluation.
[0098] Through the above-described embodiment, by analyzing the frame image sequence of a video call in detail, key features such as color richness and saturation standard deviation were extracted. These features together constitute the video fidelity features. These features can accurately evaluate the color quality of a video call, especially under complex network conditions. They can effectively detect color distortion and unevenness in video calls, thereby significantly improving the accuracy and reliability of video call quality assessment.
[0099] In combination with the above embodiments, the present invention further provides another method for evaluating the quality of user experience of a video call. In this method, the cross-modal multi-sensory feature fuser includes a cross-modal attention unit and a position encoding unit. Before executing step S13 of "fusing the first multi-modal features to obtain the first multi-sensory features", steps S61 to S63 are further included. Step S13 specifically includes steps S13-1 and S13-2: Step S61: construct a cross-modal attention unit, which includes: a feature projection layer, a multi-head attention layer, a residual connection and layer normalization, and a feedforward network.
[0100] In this embodiment, in order to achieve the fusion of the first multimodal features, it is necessary to first construct a cross-modal attention unit, which includes the following key parts: 1. Feature projection layer: The features of different modes contained in the first multimodal feature are mapped to the same feature space through linear transformation to facilitate subsequent feature fusion operations.
[0101] 2. Multi-head attention layer: Through the multi-head attention mechanism, the interaction between features of different modalities contained in the first multimodal feature is enhanced, so that features in different positions can be paid attention to simultaneously. By using dynamic features as queries (Query) and static features as keys (Key) and values (Value), attention output is obtained by performing attention calculations. Dynamic features refer to the extracted video content perception features, and static features refer to the extracted video clarity features, video fidelity features, and video smoothness features. It should be noted that the "static" in static features refers to the fact that the form of the feature data extracted by the model is relatively static and does not "dynamically" change over time. This is different from the traditional sense of video smoothness.
[0102] ; Among them, Q, K, and V are query, key, and value respectively. is the dimension of the key.
[0103] 3. Residual connection and layer normalization: Use the nn.LayerNorm function of the torch library in Python to perform layer normalization and add a residual connection.
[0104] 4. Feedforward network: The fused features are further processed through the feedforward network to extract more abstract feature representations. After obtaining the output of the feedforward network, the output result is repeatedly input into the feedforward network to obtain the final attention output.
[0105] Step S62: constructing a position encoding unit, wherein the position encoding unit uses a sine function and a cosine function to perform position encoding on each frame image in the frame image sequence.
[0106] Specifically, this class encodes each position pos and each dimension i according to the following calculation formula:
[0107] in, is the input dimension of the model, PE represents position encoding, pos represents each position, i Represents each dimension.
[0108] In this embodiment, in order to facilitate understanding of the relative relationship of the first multimodal features in the time series, a position encoding unit can be constructed. The unit generates position encoding by using sine and cosine functions to provide a unique representation for each time step, so that the position of the feature in each frame image is determined.
[0109] In step S63 , the video smoothness feature, the video definition feature, and the video fidelity feature are respectively copied in the time dimension according to the number of frames of the multiple images included in the frame image sequence.
[0110] In this embodiment, in order to align the features of different modalities contained in the first multimodal feature in the temporal dimension, the video smoothness feature, video clarity feature, and video fidelity feature need to be replicated in the temporal dimension. Specifically, if the frame image sequence contains 10 frames, the video smoothness feature, video clarity feature, and video fidelity feature will be replicated 10 times to form a time series. In this way, the feature vector of each time step will contain information from different modalities, providing consistent temporal alignment for subsequent feature fusion.
[0111] The step S13 of “fusing the first multimodal features to obtain a first multi-sensory feature” includes: Step S13 - 1 : performing position coding on each frame image in the frame image sequence by the position coding unit.
[0112] In this embodiment, a position encoding unit is first used to position-encode each frame in the frame image sequence. The position encoding is registered as a buffer and added to the feature vector of each time step. It does not participate in backpropagation, thereby reducing the amount of computation and maintaining the stability of the feature position information. This facilitates understanding the order and relative relationship of the features of different modalities contained in the first multimodal feature in the time series.
[0113] In step S13-2, the cross-modal attention unit is used to fuse the video content perception feature in the first multimodal feature and the video smoothness feature, video clarity feature, and video fidelity feature replicated in the time dimension to obtain a first multi-perception feature.
[0114] Figure 4 This is a schematic diagram of feature fusion in a method for evaluating the quality of user experience of a video call provided by an embodiment of the present invention, with reference to Figure 4 In this embodiment, a cross-modal attention unit is used to fuse video content perception features with temporally replicated video smoothness, clarity, and fidelity features. The multi-head attention mechanism simultaneously focuses on features from different modalities, and the fused features are further processed by a feedforward network through residual connections and layer normalization. This process is repeated until the first multi-perceptual feature is finally obtained.
[0115] Through the above embodiment, by constructing a cross-modal attention unit and a position encoding unit, the different modal features contained in the first multimodal feature are effectively fused. The cross-modal attention unit enhances the interaction between different modal features through a multi-head attention mechanism, and the position encoding unit helps the model understand the relative relationship in the time series. By replicating the video smoothness feature, video clarity feature, and video fidelity feature in the time dimension, the temporal alignment of features of different modalities is ensured. Ultimately, by fusing these features through the cross-modal attention unit, the first multi-perception feature obtained can fully reflect the user experience quality of the video call, thereby significantly improving the accuracy and reliability of the video call quality evaluation.
[0116] In combination with the above embodiment, the present invention further provides another method for evaluating the quality of user experience of a video call. In this method, before executing step S14 of "using a pre-trained quality evaluation model to evaluate the target video based on the first multi-sensory features to obtain a predicted score", steps S71 to S78 are further included: Step S71 , obtaining real video call data of video calls between the same calling end and different called ends under real mobile network conditions, and extracting a real frame image sequence from the real video call data.
[0117] In this embodiment, to train the quality assessment model, we first need to collect real video call data from video calls between the same calling party and different called parties under real mobile network conditions. From this real video call data, we extract one frame at a regular interval (for example, every second) to form a real frame image sequence. This real frame image sequence will serve as sample data for training the quality assessment model.
[0118] Step S72: Obtain subjective scores of multiple experts on the real video call data.
[0119] In this embodiment, in order to evaluate the quality of the video call, it is necessary to obtain subjective scores of multiple experts on the real target video.
[0120] Specifically, subjective ratings are derived from a subjective experiment involving multiple experts, with a rating range of [1, 5]. The average of these scores is used as the final subjective score, reflecting the experts' subjective perception of video call quality. These subjective ratings serve as labels for training data to guide model learning.
[0121] For example, the relationship between subjective ratings and subjective feelings is shown in Table 1: Table 1: Correspondence between subjective ratings and subjective feelings
[0122] Referring to Table 1, the higher the subjective score given by the expert, the better the expert's subjective feeling about the real video call data.
[0123] Step S73: extracting second multimodal features from the real frame image sequence through the multiple feature extractors, where the second multimodal features include at least: video content perception features, video smoothness features, video clarity features, and video fidelity features.
[0124] In this embodiment, multiple feature extractors are used to extract second multimodal features from real-world frame image sequences. These second multimodal features are those extracted during the training phase to evaluate the user experience quality of video calls. These second multimodal features also include video content perception features, video fluency features, video clarity features, and video fidelity features. The extraction process for these second multimodal features can be referenced above for the extraction of the first multimodal features and will not be repeated here.
[0125] Step S74: The second multimodal features are fused by the cross-modal multi-sensory feature fuser to obtain second multi-sensory features.
[0126] In this embodiment, the second multi-sensory feature refers to a multi-sensory feature obtained by fusing the second multimodal feature during the training phase to evaluate the user experience quality of a video call. The extraction process for the second multi-sensory feature can be referenced to the extraction process for the first multi-sensory feature described above and will not be repeated here.
[0127] Step S75: Construct a quality evaluation model to be trained. The quality evaluation model includes: a GRU layer and a time pooling layer. The GRU layer is used to capture long-term dependencies in video calls. The time pooling layer is used to generate a final prediction score for user experience quality through the subjective inspiration of the time pooling layer.
[0128] In this embodiment, the quality evaluation model to be trained is constructed as follows: First, using the Torch library in Python, we built a GRU (Gated Recurrent Unit) layer, a recurrent neural network used to process sequential data. Video calls require long-term dependencies, and the GRU's gating mechanism effectively captures these dependencies while maintaining low computational complexity.
[0129] The core update formula of GRU: ; in, represents the update gate, represents a candidate hidden state.
[0130] Next, we construct a subjectively inspired temporal pooling layer. Using the torch library in Python, we implement the following steps: initializing the maximum and minimum tensors, performing max pooling, performing weighted average pooling, performing normalization, and finally outputting the data. By combining max pooling and weighted average pooling, we pool the time series data using a subjectively inspired approach. While also taking into account the input's attenuation characteristics, we control the impact of the two pooling methods using the Beta parameter to generate a more representative temporal feature representation.
[0131] ; , .
[0132] in, Represents the final output video-level quality prediction score; represents the frame-level quality prediction score of the jth frame; τ represents the time window size, that is, the number of consecutive frames considered for each pooling; β represents the weight factor, which is used to control the influence of the two pooling methods; Represents the minimum pooling result within the time window; Represents the weighted average pooling result, with weights given by In addition, i represents the starting position (or reference position) of the sliding window; j represents the index of each frame in the time window, ranging from i to i+τ-1.
[0133] Step S76: Input the second multi-sensory feature as sample data into the quality assessment model to be trained, and use the subjective score as label data to obtain the predicted score of the output of the quality assessment model to be trained for the sample data.
[0134] In this embodiment, the second multi-perceptual feature is input into the GRU layer (Gated Recurrent Unit, a recursive neural network used to process sequence data. Video calls require the establishment of long-term dependencies. Due to the gating mechanism of GRU, it can effectively capture long-distance dependencies while maintaining low computational complexity) to obtain a frame-level user experience quality prediction score for each video. Subsequently, the feature is input into the subjectively inspired time pooling layer to obtain the final video-level user experience quality prediction score for each sample data (i.e., real video call data).
[0135] At the same time, during the training process, you can use the random.shuffle function in Python to randomly shuffle the index of the sample data in the dataset to ensure that the quality evaluation model will not be biased by the order of the sample data during training.
[0136] The model learning rate was then set to 0.00001. The ADAM optimizer was used to dynamically adjust the model learning rate during training, multiplying the learning rate by 0.8 after every 50 training rounds. The batch size was set to 16. The subjective ratings included in the sample data were used as label data for the training quality evaluation model, allowing for supervised learning.
[0137] Step S77: Setting a prediction effect index, wherein the prediction effect index is used to evaluate the prediction effect of the quality evaluation model of the video call.
[0138] In this embodiment, to evaluate the performance of the quality assessment model, prediction metrics are set to assess the video call quality assessment model's prediction effectiveness. These metrics include the Pearson Linear Correlation Coefficient (PLCC), the Spearman Rank Correlation Coefficient (SROCC), the Root Mean Square Error (RSME), and the Mean Absolute Error (MAE). These metrics are used together to evaluate the quality assessment model's prediction effectiveness. PLCC and SROCC measure the correlation between the predicted score and the subjective rating, while RSME and MAE measure the difference between the predicted score and the subjective rating.
[0139] Step S78, training the quality evaluation model to be trained through multiple iterations, and determining whether the quality evaluation model training is completed according to the prediction effect index.
[0140] In this embodiment, a preset threshold can be set for each indicator in the prediction effect index. When these prediction effect indicators reach the preset indicator threshold, the setting of the preset indicator threshold can refer to: PLCC is close to 1, SROCC is close to 1, and the smaller the RMSE and MAE values are, the more it indicates that the prediction effect of the quality evaluation model has met the requirements. At this time, it can be determined that the quality evaluation model training is completed.
[0141] The technical effects of the present invention will be described below through various experimental data: 1. Comparative experimental results of video content-aware feature extractors based on different CNNs: The experiment will test the performance of the ResNet50-SGAS model and compare its performance with the VGG16 model widely used in the industry. The experimental results are shown in Table 2.
[0142] The VGG16 model comes from the Visual Geometry Group of the University of Oxford in the UK. It uses more hidden layers and undergoes a large amount of image training. It is a pre-trained CNN model.
[0143] Table 2: Experimental results of different video content-aware feature extractors
[0144] The results show that the improved ResNet50-SGAS model of this invention demonstrates significant advantages over the traditional VGG16 model. The SROCC metric improved by 114.04%, the PLCC metric improved by 153.08%, the RMSE metric decreased by 44.22%, and the MAE metric decreased by 39.92%. The ResNet50-SGAS model significantly outperformed VGG16 in all performance metrics, particularly achieving exponential performance improvements in SROCC and PLCC. This demonstrates that the ResNet50-SGAS model is capable of better extracting video content-aware features and improving the accuracy of predicting the user experience quality score for video calls.
[0145] 2. Comparative Experiments of Different Feature Fusion Methods This paper proposes a cross-modal multi-perceptual feature fusion method. The experiment randomly selects some data sets for testing and compares them with the simple splicing method. The results are shown in Table 3.
[0146] Table 3: Experimental results of different feature fusion methods
[0147] The results show that the proposed cross-modal multi-sensory feature fuser outperforms the simple splicing method. The SROCC index improved by 0.34%, the PLCC index improved by 7.09%, the RMSE index decreased by 10.27%, and the MAE index decreased by 5.64%. The data show that the correlation, association, and consistency between features have all been improved to a certain extent, indicating that the method is effective.
[0148] 3. Comparative Experiment between the Quality Evaluation Model Proposed in the Present Invention and the Existing VMOS The existing VMOS evaluation score only considers the smoothness of video calls, with frame loss accounting for 60% and freezes and freeze duration accounting for 20% each. The proposed quality evaluation model was compared with the existing VMOS method. The performance test results are shown in Table 4.
[0149] Table 4: Experimental results of different evaluation models
[0150] The results show that compared with the existing VMOS, the quality evaluation model proposed in this paper has significant improvements in multiple indicators. The SROCC indicator increased by 96.03%, indicating that it has a stronger ability to capture the relationship between the predicted ranking and the actual ranking. The PLCC indicator increased by 44.04%, indicating that the linear relationship between the predicted user experience quality score and the subjective score is stronger. The RMSE indicator decreased by 81.09%, indicating that the prediction error is much smaller than that of the existing VMOS. The MAE indicator decreased by 85.40%, indicating that the error is smaller and the prediction performance is better than that of the existing VMOS.
[0151] The key points of the present invention are: 1. Comprehensively evaluate the user experience quality of video calls based on multiple user perception characteristics, enabling a more accurate assessment of the user experience under real mobile network conditions.
[0152] 2. We built an innovative cross-modal multi-sensory feature fuser, which combines the cross-modal attention mechanism and time-embedded position encoding to enable more effective fusion of the extracted feature data.
[0153] 3. This invention utilizes a deep learning neural network approach, taking into account both the subjective perception of the human eye (i.e., the scoring mechanism of the quality scoring model) and the objective indicators of video calls (i.e., multi-perceptual features), and proposes a new subjective and objective user experience quality evaluation framework. This framework includes video content perception, fluency, clarity, and picture clarity and fidelity. Compared with traditional subjective evaluation methods, this framework can evaluate the user experience quality of video calls more quickly, efficiently, accurately, and flexibly.
[0154] Figure 5 This is a schematic diagram of a video call user experience quality evaluation device provided by an embodiment of the present invention, with reference to Figure 5 Based on the same inventive concept, another embodiment of the present invention further provides a device for evaluating the user experience quality of a video call, which is applied to a video call scenario and is used to evaluate the user experience quality of a called party in a video call. The device includes: The frame image sequence acquisition module 11 is used to acquire a frame image sequence from the target video of the called end of the video call; A feature extraction module 12 is configured to extract first multimodal features from the frame image sequence using a plurality of feature extractors, wherein the first multimodal features include at least: a video content perception feature, a video fluency feature, a video clarity feature, and a video fidelity feature; a feature fusion module 13, configured to fuse the first multimodal features through a cross-modal multi-sensory feature fuser to obtain a first multi-sensory feature; The quality evaluation module 14 is used to evaluate the target video according to the first multi-perceptual features using a pre-trained quality evaluation model to obtain a predicted score; the quality evaluation model is obtained by training a neural network using real video call data of multiple video calls between the same calling end and different called ends under real mobile network conditions as sample data, and the sample data carries subjective scores, which are subjective scores of multiple experts on the real video call data.
[0155] Optionally, the feature extraction module 12 includes: A first calculation unit is configured to calculate the average value of all pixels in each feature map of multiple frame images included in the frame image sequence through a spatial global average pooling layer of a video content-aware feature extractor, so as to obtain an average value feature vector of the frame image sequence; a second calculation unit, configured to calculate the standard deviation of all pixels in each feature map of the plurality of frames included in the frame image sequence through a spatial global standard deviation pooling layer of a video content-aware feature extractor, to obtain a standard deviation feature vector of the frame image sequence; a concatenation unit, configured to concatenate the average feature vector and the standard deviation feature vector to obtain a concatenated feature vector; a dimensionality reduction unit, configured to reduce the dimensionality of the concatenated feature vector using a fully connected neural network of a video content-aware feature extractor to obtain the video content-aware feature; Among them, the video content-aware feature extractor includes: an improved ResNet50-SGAS model and a fully connected neural network connected to the output end of the improved ResNet50-SGAS model; the improved ResNet50-SGAS model is obtained by replacing the fully connected layer and the global average pooling layer in ResNet50 with a spatial global average pooling layer and a spatial global standard deviation pooling layer, and the fully connected neural network includes an input layer, a Dropout layer, and a fully connected layer.
[0156] Optionally, the feature extraction module 12 includes: a lost frame number determining unit, configured to determine the number of lost frames of the video call based on frame numbers of multiple frames of images included in the frame image sequence and the number of original video frames of the calling end of the video call; an optical flow average amplitude determination unit, configured to determine an optical flow average amplitude of each frame image according to the optical flow between each adjacent frame in the plurality of frames of images included in the frame image sequence; a frame difference score determining unit, configured to determine a frame difference score for each frame image based on a total pixel sum of a frame difference image between each frame image and an adjacent frame image of the frame image in a plurality of frame images included in the frame image sequence; a freeze determination unit, configured to, for the frame image sequence, determine that a freeze has occurred in the target video when the average amplitude of the optical flow of any frame image among the multiple frames included in the frame image sequence is less than a first threshold and the frame difference score is less than a second threshold, and record the number of freezes in the target video and the duration of each freeze; The video fluency feature determination unit is used to determine the number of lost frames, the number of freezes, and the freeze duration as the video fluency feature of the target video.
[0157] Optionally, the feature extraction module 12 includes: a third computing unit, configured to convert each frame image included in the frame image sequence into a grayscale image, and calculate a variance value of the Laplace transform of the target video using a Laplace operator; a gradient energy value determining unit, configured to perform convolution calculation on the grayscale image using a Sobel operator to obtain a Sobel gradient energy value of the target video; a fourth calculating unit, configured to calculate a total variation value of the gradient of each frame image in the frame image sequence; The video definition feature determining unit is configured to determine the variance value of the Laplace transform, the Sobel gradient energy value, and the total variation value of the gradient as the video definition feature.
[0158] Optionally, the feature extraction module 12 includes: a color richness value determining unit, configured to determine a color richness value of the target video according to each frame image included in the frame image sequence; a saturation standard deviation value determining unit, configured to determine a standard deviation value of the saturation of the target video based on each frame image included in the frame image sequence; The video fidelity feature determining unit is configured to determine a standard deviation value of the color richness value and the saturation as the video fidelity feature.
[0159] Optionally, the cross-modal multi-sensory feature fuser includes a cross-modal attention unit and a position encoding unit, and the device further includes: A first construction unit is configured to construct a cross-modal attention unit before fusing the first multimodal features to obtain a first multi-sensory feature, wherein the cross-modal attention unit includes: a feature projection layer, a multi-head attention layer, a residual connection and layer normalization, and a feedforward network; A second construction unit is configured to construct a position encoding unit, wherein the position encoding unit performs position encoding on each frame image in the frame image sequence using a sine function and a cosine function; a copying unit, configured to copy the video smoothness feature, the video clarity feature, and the video fidelity feature in a time dimension according to the number of frames of the plurality of images included in the frame image sequence; The feature fusion module 13 includes: a position encoding unit, configured to perform position encoding on each frame image in the frame image sequence by means of the position encoding unit; A feature fusion unit is used to perform feature fusion on the video content perception feature in the first multimodal feature and the video smoothness feature, video clarity feature and video fidelity feature replicated in the time dimension through the cross-modal attention unit to obtain a first multi-perception feature.
[0160] Optionally, the device further comprises: a data set construction unit for, before evaluating the target video using a pre-trained quality assessment model based on the first multi-perceptual features to obtain a prediction score, collecting real video call data of video calls between the same calling end and different called ends under real mobile network conditions, and extracting real frame image sequences from the real video call data; A subjective score obtaining unit, configured to obtain subjective scores of multiple experts on the real video call data; a second multimodal feature extraction unit, configured to extract second multimodal features from the real frame image sequence using the multiple feature extractors, wherein the second multimodal features include at least: a video content perception feature, a video smoothness feature, a video clarity feature, and a video fidelity feature; a second multi-sensory feature fusion unit, configured to fuse the second multi-modal features through the cross-modal multi-sensory feature fuser to obtain a second multi-sensory feature; A third construction unit is configured to construct a quality assessment model to be trained. The quality assessment model includes a GRU layer and a temporal pooling layer. The GRU layer is configured to capture long-term dependencies in video calls. The temporal pooling layer is configured to generate a final predicted score for user experience quality through the subjectively inspired temporal pooling layer. a training unit, configured to input the second multi-sensory feature as sample data into the quality assessment model to be trained, and use the subjective score as label data to obtain a prediction score for the sample data output by the quality assessment model to be trained; An indicator setting unit, configured to set a prediction effect indicator, wherein the prediction effect indicator is used to evaluate the prediction effect of the quality evaluation model of the video call; The training effect determination unit is used to train the quality evaluation model to be trained through multiple iterations, and determine whether the quality evaluation model training is completed according to the prediction effect index.
[0161] Based on the same inventive concept, another embodiment of the present invention further provides an electronic device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the video call user experience quality evaluation method as described in any of the above embodiments.
[0162] Based on the same inventive concept, another embodiment of the present invention further provides a computer program product, including a computer program, which is executed by a processor to implement the method for evaluating the quality of video call user experience as described in any of the above embodiments.
[0163] Based on the same inventive concept, another embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the method for evaluating the quality of video call user experience as described in any of the above embodiments is implemented.
[0164] As for the device, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0165] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0166] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, apparatus, or computer program products. Thus, embodiments of the present invention may take the form of a fully hardware embodiment, a fully software embodiment, or an embodiment combining software and hardware. Furthermore, embodiments of the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0167] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the process in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0168] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0169] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable terminal device to implement the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0170] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they become aware of the basic creative concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.
[0171] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or terminal device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or terminal device that includes the element.
[0172] The above is a detailed introduction to the method, device, equipment and medium for evaluating the user experience quality of a video call provided by the present invention. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method and core ideas of the present invention. At the same time, for those skilled in the art, according to the ideas of the present invention, there may be changes in the specific implementation methods and application scopes. In summary, the contents of this specification should not be understood as limiting the present invention.
Claims
1. A method for evaluating the quality of user experience of a video call, characterized in that: Applied to video call scenarios, the method is used to evaluate the user experience quality of the called party in the video call, including: Obtain a frame image sequence from a target video of a called party of a video call; Extracting first multimodal features from the frame image sequence using a plurality of feature extractors, wherein the first multimodal features include at least: a video content perception feature, a video smoothness feature, a video clarity feature, and a video fidelity feature; fusing the first multimodal features through a cross-modal multi-sensory feature fuser to obtain a first multi-sensory feature; The target video is evaluated according to the first multi-perceptual features using a pre-trained quality evaluation model to obtain a predicted score. The quality evaluation model is obtained by training a neural network using real video call data of multiple video calls between the same calling end and different called ends under real mobile network conditions as sample data. The sample data carries subjective scores, which are subjective scores given by multiple experts to the real video call data.
2. The method for evaluating the quality of user experience of a video call according to claim 1, wherein: The video content perception feature is obtained by the following steps: Calculating the average value of all pixels in each feature map of multiple frames included in the frame image sequence through a spatial global average pooling layer of a video content-aware feature extractor to obtain an average value feature vector of the frame image sequence; Calculating the standard deviation of all pixels in each feature map of the multiple frames included in the frame image sequence through a spatial global standard deviation pooling layer of a video content-aware feature extractor to obtain a standard deviation feature vector of the frame image sequence; Concatenating the mean value feature vector and the standard deviation feature vector to obtain a concatenated feature vector; Performing dimensionality reduction on the concatenated feature vector using a fully connected neural network of a video content-aware feature extractor to obtain the video content-aware feature; Among them, the video content-aware feature extractor includes: an improved ResNet50-SGAS model and a fully connected neural network connected to the output end of the improved ResNet50-SGAS model; the improved ResNet50-SGAS model is obtained by replacing the fully connected layer and the global average pooling layer in ResNet50 with a spatial global average pooling layer and a spatial global standard deviation pooling layer, and the fully connected neural network includes an input layer, a Dropout layer, and a fully connected layer.
3. The method for evaluating the quality of user experience of a video call according to claim 1, wherein: The video fluency feature is obtained by performing the following steps by a video fluency feature extractor: determining the number of lost frames of the video call according to frame numbers of the multiple frames of images included in the frame image sequence and the number of original video frames of the calling end of the video call; Determining an average optical flow amplitude of each frame image according to the optical flow between each adjacent frame in the plurality of frame images included in the frame image sequence; determining a frame difference score for each frame image according to a total pixel sum of a frame difference image between each frame image and an adjacent frame image of the frame image in the plurality of frame images included in the frame image sequence; For the frame image sequence, when the average optical flow amplitude of any frame image among the multiple frames included in the frame image sequence is less than a first threshold and the frame difference score is less than a second threshold, determining that a freeze occurs in the target video, and recording the number of freezes in the target video and the duration of each freeze; The number of lost frames, the number of freezes, and the freeze duration are determined as video smoothness features of the target video.
4. The method for evaluating the quality of user experience of a video call according to claim 1, wherein: The video definition feature is obtained by performing the following steps by a video definition feature extractor: Converting each frame image included in the frame image sequence into a grayscale image, and calculating the variance value of the Laplace transform of the target video through the Laplace operator; Perform convolution calculation on the grayscale image using a Sobel operator to obtain a Sobel gradient energy value of the target video; Calculating the total variation value of the gradient of each frame image in the frame image sequence; The variance value of the Laplace transform, the Sobel gradient energy value, and the total variation value of the gradient are determined as the video clarity feature.
5. The method for evaluating the quality of user experience of a video call according to claim 1, wherein: The video fidelity feature is obtained by performing the following steps by a video fidelity feature extractor: Determining a color richness value of the target video according to each frame image included in the frame image sequence; Determining a standard deviation value of the saturation of the target video according to each frame image included in the frame image sequence; A standard deviation value of the color richness value and the saturation is determined as the video fidelity feature.
6. The method for evaluating the quality of user experience of a video call according to claim 1, wherein: The cross-modal multi-sensory feature fuser includes a cross-modal attention unit and a position encoding unit. Before fusing the first multi-modal features to obtain the first multi-sensory features, the cross-modal multi-sensory feature fuser includes: Constructing a cross-modal attention unit, which includes: a feature projection layer, a multi-head attention layer, residual connections and layer normalization, and a feedforward network; Constructing a position encoding unit, wherein the position encoding unit uses a sine function and a cosine function to perform position encoding on each frame image in the frame image sequence; Respectively copying the video smoothness feature, the video clarity feature, and the video fidelity feature in the time dimension according to the number of frames of the multiple images included in the frame image sequence; The fusing the first multimodal features to obtain a first multi-sensory feature includes: Performing position coding on each frame image in the frame image sequence by the position coding unit; Through the cross-modal attention unit, the video content perception feature in the first multimodal feature and the video smoothness feature, video clarity feature and video fidelity feature replicated in the time dimension are feature fused to obtain a first multi-perception feature.
7. The method for evaluating the quality of user experience of a video call according to claim 1, wherein: Before evaluating the target video using the pre-trained quality evaluation model according to the first multi-perceptual features to obtain a prediction score, the method further includes: Real video call data of video calls between the same calling end and different called ends under real mobile network conditions, and extracting real frame image sequences from the real video call data; Obtaining subjective scores from multiple experts on the real video call data; Extracting second multimodal features from the real frame image sequence using the multiple feature extractors, the second multimodal features comprising at least: video content perception features, video smoothness features, video clarity features, and video fidelity features; fusing the second multimodal features by the cross-modal multi-sensory feature fuser to obtain a second multi-sensory feature; Constructing a quality assessment model to be trained, the quality assessment model includes: a GRU layer and a temporal pooling layer. The GRU layer is used to capture long-term dependencies in video calls, and the temporal pooling layer is used to generate a final predicted score for user experience quality through the subjective inspiration of the temporal pooling layer. Inputting the second multi-sensory feature as sample data into the quality assessment model to be trained, and using the subjective score as label data to obtain a prediction score for the sample data output by the quality assessment model to be trained; Setting a prediction effect indicator, wherein the prediction effect indicator is used to evaluate the prediction effect of the quality evaluation model of the video call; The quality evaluation model to be trained is trained through multiple iterations, and it is determined that the training of the quality evaluation model is completed according to the prediction effect index.
8. A device for evaluating the quality of user experience of a video call, characterized in that: Applied to video call scenarios, the device is used to evaluate the user experience quality of the called party in the video call, and includes: A frame image sequence acquisition module is used to acquire a frame image sequence from a target video of a called party of a video call; a feature extraction module, configured to extract first multimodal features from the frame image sequence using a plurality of feature extractors, wherein the first multimodal features include at least: a video content perception feature, a video smoothness feature, a video clarity feature, and a video fidelity feature; a feature fusion module, configured to fuse the first multimodal features through a cross-modal multi-sensory feature fuser to obtain a first multi-sensory feature; A quality assessment module is used to evaluate the target video according to the first multi-perceptual features using a pre-trained quality assessment model to obtain a predicted score; the quality assessment model is obtained by training a neural network using real video call data of multiple video calls between the same calling end and different called ends under real mobile network conditions as sample data, and the sample data carries subjective scores, which are subjective scores given by multiple experts to the real video call data.
9. An electronic device, characterized in that: The system comprises a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the method for evaluating the quality of video call user experience according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that A computer program is stored thereon, wherein when the computer program is executed by a processor, the method for evaluating the quality of video call user experience according to any one of claims 1 to 7 is implemented.