Video Content Quality Evaluation Method, Device, Equipment and Storage Medium Based on Big Data Analysis

By obtaining the multi-dimensional characteristics of the video and training the scoring model, the problem of incomplete video quality evaluation in the prior art is solved, and more comprehensive and accurate evaluation results are achieved.

CN119496924BActive Publication Date: 2025-08-05NANJING SUYI IND
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411522802.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-29
Publication Date
2025-08-05
Estimated Expiration
2044-10-29

AI Technical Summary

Technical Problem

The existing video quality evaluation methods based on deep learning models lack comprehensiveness and cannot fully cover the complex characteristics of video quality.

Method used

By obtaining the user score and standard score of the sample video set, video preprocessing is performed to obtain the comprehensive video definition, video frame stability, audio clarity and audio signal-to-noise ratio, the target feature vector is generated, and when the initial scoring model does not meet the preset conditions, the model is trained based on these feature vectors and scores until the preset conditions are met, and the target scoring model is obtained.

Benefits of technology

A comprehensive and accurate assessment of video quality is achieved to ensure the comprehensiveness and accuracy of the evaluation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119496924B_ABST
    Figure CN119496924B_ABST
Patent Text Reader

Abstract

The embodiments of the present invention disclose a video content quality assessment method, apparatus, device, and storage medium based on big data analysis, including: obtaining a sample video set, user ratings, and standard ratings for each video in the sample video set; performing video preprocessing on the sample video set to obtain the video comprehensive clarity, video frame stability, audio clarity, and audio signal-to-noise ratio of each video in the sample video set; obtaining a target feature vector for each video based on the video comprehensive clarity, video frame stability, audio clarity, and audio signal-to-noise ratio of each video in the sample video set; and when a pre-established initial scoring model does not meet preset conditions, training the initial scoring model based on each target feature vector, user ratings, and standard ratings for each video until the initial scoring model meets the preset conditions, thereby obtaining a target scoring model. The present invention enables the scoring model to comprehensively evaluate videos, ensuring that the evaluation results are more comprehensive and accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of video processing technology, and in particular to a video content quality assessment method, apparatus, device and storage medium based on big data analysis. Background Art

[0002] With the rapid development of digital video technology, video content has become one of the most important resources in the internet and media industries. The richness and diversity of video not only enhances user experience but also intensifies the demand for video quality assessment. Accurate and objective assessment of video quality is not only crucial for content creators and platform operators but also has a direct impact on optimizing user experience and improving satisfaction. With the advancement of artificial intelligence, an increasing number of research and applications are leveraging machine learning and deep learning models to evaluate video quality.

[0003] Current video evaluation methods based on deep learning models are mostly based on simple technical metrics. These metrics provide information about the technical aspects of a video. However, while these metrics can reflect some basic video characteristics, they often fail to fully cover all aspects of video quality. Therefore, this approach lacks comprehensiveness and fails to capture the complex characteristics of video quality. Summary of the Invention

[0004] The embodiments of the present invention provide a video content quality assessment method, apparatus, device and storage medium based on big data analysis, which can obtain a comprehensive target feature vector through evaluation indicators in multiple dimensions, so that the scoring model can comprehensively evaluate the video and ensure that the assessment results are more comprehensive and accurate.

[0005] In a first aspect, an embodiment of the present invention provides a video content quality assessment method based on big data analysis, comprising:

[0006] Obtain a sample video collection, and user ratings and standard ratings of each video in the sample video collection;

[0007] Performing video preprocessing on the sample video set to obtain comprehensive video clarity, video frame stability, audio clarity, and audio signal-to-noise ratio of each video in the sample video set;

[0008] Obtaining a target feature vector for each video based on comprehensive video clarity, video frame stability, audio clarity, and audio signal-to-noise ratio of each video in the sample video set;

[0009] When the pre-established initial scoring model does not meet the preset conditions, the initial scoring model is trained based on each target feature vector, the user rating of each video and the standard rating until the initial scoring model meets the preset conditions, thereby obtaining a target scoring model; wherein, the target scoring model is used to perform quality assessment on the video to be processed.

[0010] In a second aspect, an embodiment of the present invention provides a video content quality assessment device based on big data analysis, the device comprising:

[0011] A data acquisition module is used to obtain a sample video set, user ratings and standard ratings of each video in the sample video set;

[0012] a video processing module, configured to perform video preprocessing on the sample video set to obtain comprehensive video clarity, video frame stability, audio clarity, and audio signal-to-noise ratio of each video in the sample video set;

[0013] a vector determination module, configured to obtain a target feature vector for each video in the sample video set based on the comprehensive video clarity, video frame stability, audio clarity, and audio signal-to-noise ratio of each video;

[0014] The model training module is used to train the pre-established initial scoring model based on each target feature vector, the user rating of each video and the standard rating when the pre-established initial scoring model does not meet the preset conditions, until the initial scoring model meets the preset conditions, thereby obtaining a target scoring model; wherein the target scoring model is used to perform quality assessment on the video to be processed.

[0015] In a third aspect, an embodiment of the present invention further provides an electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, a video content quality assessment method based on big data analysis as described in any one of the embodiments of the present invention is implemented.

[0016] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a video content quality assessment method based on big data analysis as described in any one of the embodiments of the present invention.

[0017] In a fifth aspect, an embodiment of the present invention further provides a computer program product, including a computer program, which, when executed by a processor, implements a video content quality assessment method based on big data analysis as described in any one of the embodiments of the present invention.

[0018] In an embodiment of the present invention, a sample video set and user ratings and standard ratings of each video in the sample video set are obtained; video preprocessing is performed on the sample video set to obtain the comprehensive video clarity, video frame stability, audio clarity, and audio signal-to-noise ratio of each video in the sample video set; a target feature vector of each video is obtained based on the comprehensive video clarity, video frame stability, audio clarity, and audio signal-to-noise ratio of each video in the sample video set; when a pre-established initial rating model does not meet preset conditions, the initial rating model is trained based on each target feature vector, the user ratings, and standard ratings of each video until the initial rating model meets the preset conditions, thereby obtaining a target rating model. That is, in an embodiment of the present invention, a comprehensive target feature vector is obtained through comprehensive video clarity, video frame stability, audio clarity, and audio signal-to-noise ratio, so that the rating model can comprehensively evaluate the video, ensuring that the evaluation results are more comprehensive and accurate. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.

[0020] Figure 1 A first flow chart of a method for evaluating video content quality based on big data analysis provided by an embodiment of the present invention;

[0021] Figure 2 A second flow chart of a video content quality assessment method based on big data analysis provided by an embodiment of the present invention;

[0022] Figure 3 A schematic diagram of the structure of a video content quality assessment device based on big data analysis provided by an embodiment of the present invention;

[0023] Figure 4 A schematic structural diagram of an electronic device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0024] The present invention will be further described in detail below with reference to the accompanying drawings and examples. It will be understood that the specific embodiments described herein are intended only to illustrate the present invention and are not intended to limit the present invention. It should also be noted that, for ease of description, the accompanying drawings only illustrate portions relevant to the present invention, not all structures.

[0025] Figure 1This is the first flow chart of a method for evaluating video content quality based on big data analysis provided by an embodiment of the present invention. The method of the embodiment of the present invention can obtain a comprehensive target feature vector through evaluation indicators of multiple dimensions, so that the scoring model can comprehensively evaluate the video, ensuring that the evaluation results are more comprehensive and accurate. The method can be executed by a video content quality evaluation device based on big data analysis provided by an embodiment of the present invention, and the device can be implemented in software and / or hardware. The following embodiments will be described by taking the device integrated in an electronic device as an example. The electronic device can be a computer device or a server. Reference Figure 1 , the method may specifically include the following steps:

[0026] Step 101: Obtain a sample video set, and the user rating and standard rating of each video in the sample video set.

[0027] The sample video collection consists of display videos, which are videos presented to users, such as online short videos, film and television videos, or advertising videos. User ratings are the ratings of the display videos given by users who viewed them. Standard ratings are based on established standards and criteria for domain big data and evaluations of videos by domain professionals. Specifically, when a target scoring model that can accurately assess video quality is needed, a sample video collection can be received from the user, or directly obtained from the corresponding device through a pre-defined data interface.

[0028] Step 102: Perform video preprocessing on the sample video set to obtain the comprehensive video clarity, video frame stability, audio clarity, and audio signal-to-noise ratio of each video in the sample video set.

[0029] Video clarity refers to the overall visual clarity of a video image, including resolution, sharpness, and detail. Video frame stability refers to the stability of the image during video playback. Audio clarity measures the intelligibility and noise-free nature of the audio signal. The audio signal-to-noise ratio (SNR) measures the ratio of useful information to noise in an audio signal, expressed in decibels. A higher SNR indicates more useful information and less noise in the signal.

[0030] In an optional embodiment, after obtaining the sample video set, the video resolution of each video can be evaluated by measuring the horizontal and vertical pixel counts of the video. The sharpness analysis tool is used to calculate the sharpness value of the video frame, and the edge clarity of the image is evaluated based on the sharpness value. The comprehensive video clarity of the video is obtained based on the edge clarity. At the same time, after obtaining the sample video set, a pre-set streamer algorithm (such as the Lucas-Kanade method or the Farneback method) is used to track the feature points between the video frames, and the motion vector of the video is calculated based on its feature points to evaluate the stability between the video frames. At the same time, after obtaining the sample video set, an audio reading tool can be used to read the audio file to obtain the audio signal and sampling rate of the video. The audio signal is separated from the signal according to a pre-set filter (such as an adaptive filter) to obtain the signal power and noise power. After obtaining the signal power and noise power, the signal power and noise power are calculated, and the ratio of the signal power to the noise power is determined as the signal-to-noise ratio of the audio signal.

[0031] In an optional embodiment, after obtaining a sample video set, for each video in the sample video set, a preset number of frames of images are randomly extracted from the current video, and preliminary image processing is performed on the preset number of frames of images to obtain candidate images corresponding to the preset number of frames of images; wherein the preliminary image processing includes at least grayscale processing and denoising processing; the correlation coefficient between the clarity of each dimension and the standard score is calculated according to a preset correlation coefficient calculation formula; the clarity of each dimension is standardized based on the correlation coefficient to obtain a standardized clarity, and a target weight of the standardized clarity is obtained based on the standardized clarity, a preset initial weight, and a first adjustment coefficient; the clarity of each dimension is feature fused according to the target weight to obtain the comprehensive video clarity of the candidate image. For each video in the sample video set, the optical flow vector variance of the video frame of the current video is calculated, and the degree of change of the light field of the video frame of the current video is determined based on the optical flow vector variance; the evaluation standard for video frame stability is determined based on the degree of change of the light field, a preset basic threshold, and a second adjustment coefficient; and the video frame stability of the current video is obtained based on the evaluation standard for video frame stability. For each video in the sample video set, the audio signal of the current video is extracted from the current video; the audio signal is subjected to fast Fourier transform and Mel spectrum feature extraction to obtain the Mel spectrum feature of the audio signal; and the clarity of the audio signal is obtained based on the Mel spectrum feature; the audio signal is subjected to short-time Fourier transform to obtain the audio signal-to-noise ratio of the current video.

[0032] Step 103: Obtain a target feature vector for each video based on the comprehensive video clarity, video frame stability, audio clarity, and audio signal-to-noise ratio of each video in the sample video set.

[0033] Among them, the target feature vector is used to train the initial scoring model, so as to obtain a target scoring model that can accurately evaluate the video quality. In this solution, after obtaining the sample video set, the time series features corresponding to the video frames of each video are obtained according to the time series of each video frame in the video. The color histogram of the video frame of each video is extracted, and the video color features corresponding to the video frame are generated according to the color histogram. The texture features of the video frame are extracted using texture analysis technology (such as grayscale co-occurrence matrix), and the content features of the video frame are generated according to the video color features and texture features. After obtaining the time series features, content features, comprehensive video clarity, video frame stability, audio clarity and audio signal-to-noise ratio, these features are fused through feature splicing or a pre-set feature fusion algorithm (such as support vector regression algorithm) to obtain the target feature vector of each video.

[0034] Step 104: When the pre-established initial scoring model does not meet the preset conditions, the initial scoring model is trained based on each target feature vector, the user rating of each video, and the standard rating until the initial scoring model meets the preset conditions, thereby obtaining the target scoring model.

[0035] The target scoring model is used to assess the quality of the processed video. The initial scoring model is a pre-established, untrained scoring model. Pre-conditions may include failure to converge a pre-set loss function or failure to reach a pre-set number of iterations. The target scoring model is designed to accurately assess video quality.

[0036] The initial scoring model in this solution can be a support vector regression model. The support vector regression model is a regression algorithm based on the support vector machine. The support vector regression model can construct a robust regression model by maximizing the margin and minimizing the error. The initial scoring model includes a pre-set loss function. After obtaining the target feature vector, the target feature vector can be input into the initial scoring model, and the target feature vector is processed by the initial scoring model to output the output score corresponding to the target feature vector. The output score, the user score of each video, and the standard score are substituted into the loss function to obtain the loss function value. According to the loss function value, the model parameters of the initial scoring model are adjusted until the initial scoring model meets the preset conditions to obtain the target scoring model. After obtaining the target scoring model, when it is necessary to score the video data to be processed, the video data to be processed can be processed to obtain the target feature vector corresponding to the video data to be processed. The target feature vector is input into the target scoring model to obtain the score corresponding to the processed video data.

[0037] The technical solution of this embodiment obtains a sample video set, user ratings and standard ratings of each video in the sample video set; performs video preprocessing on the sample video set to obtain the comprehensive video clarity, video frame stability, audio clarity and audio signal-to-noise ratio of each video in the sample video set; obtains a target feature vector for each video based on the comprehensive video clarity, video frame stability, audio clarity and audio signal-to-noise ratio of each video in the sample video set; when a pre-established initial scoring model does not meet the preset conditions, the initial scoring model is trained based on each target feature vector, user ratings and standard ratings of each video until the initial scoring model meets the preset conditions, thereby obtaining a target scoring model. The technical solution of this embodiment obtains a comprehensive target feature vector through comprehensive video clarity, video frame stability, audio clarity and audio signal-to-noise ratio, so that the scoring model can comprehensively evaluate the video, ensuring that the evaluation results are more comprehensive and accurate.

[0038] Figure 2 This is a second flow chart of a method for evaluating video content quality based on big data analysis provided by an embodiment of the present invention. This embodiment is a refinement of the above embodiment. The specific method can be as follows: Figure 2 As shown, the method may include the following steps:

[0039] Step 201: Obtain a sample video set, and the user rating and standard rating of each video in the sample video set.

[0040] Step 202 : For each video in the sample video set, randomly extract a preset number of frame images from the current video, and perform preliminary image processing on the preset number of frame images to obtain candidate images corresponding to the preset number of frame images.

[0041] Among them, preliminary image processing includes at least grayscale processing and denoising. The preset number is set based on the domain big data and the specific experimental environment. In an optional embodiment, after obtaining the preset number of frame images, each frame image is converted into a grayscale image. Grayscale images only contain brightness information, not color information. Grayscale images help reduce the amount of calculation and highlight the edge and texture features of the image. A pre-set filter is used to remove noise from each frame image to obtain a candidate image corresponding to each frame image.

[0042] Step 203: Calculate the clarity of each dimension of the candidate image, and obtain the comprehensive video clarity of the candidate image based on the clarity calculation results of each dimension.

[0043] The comprehensive video clarity refers to the overall visual clarity of the video image, including resolution, sharpness, and detail. Dimensional clarity includes Laplace clarity, discrete differential operator clarity, and Fourier clarity. In one optional implementation, after obtaining a candidate image, the image is convolved with the Laplace operator to obtain an edge response image. The variance of the edge response image is calculated, i.e., the average of the sum of the squares of the differences between all pixel values and the image mean. The calculated variance value is determined as the Laplace clarity. Simultaneously, the image gradient of the candidate image is calculated using the Sobel operator. The Sobel operator is a discrete differential operator used for edge detection. The Sobel operator detects image edges by calculating the gradients of the image in the horizontal and vertical directions. After obtaining the candidate image, the candidate image is convolved with the Sobel horizontal operator and the Sobel vertical operator, respectively, to obtain gradient images in the horizontal and vertical directions. The gradient magnitude of each pixel in the gradient image is calculated by taking the square root of the sum of the squares of the horizontal and vertical gradient values. The standard deviation of the gradient amplitudes of all pixels is calculated to obtain the discrete differential operator sharpness of the candidate image. After obtaining the candidate image, a two-dimensional Fourier transform is performed on the candidate image to convert it from the spatial domain to the frequency domain, obtaining the corresponding frequency domain image. The amplitude spectrum of the frequency domain image is calculated, that is, the amplitude value of each frequency component. After obtaining the amplitude spectrum, the energy density of the amplitude spectrum is calculated, that is, the sum of the squares of the amplitude values of all frequency components, to obtain the Fourier sharpness.

[0044] After obtaining the clarity of each dimension of the candidate image, the comprehensive video clarity of the candidate image is obtained based on the clarity calculation results of each dimension. In this solution, optionally, the comprehensive video clarity of the candidate image is obtained based on the clarity calculation results of each dimension, including: calculating the correlation coefficient between the clarity of each dimension and the standard score according to a pre-set correlation coefficient calculation formula; standardizing the clarity of each dimension based on the correlation coefficient to obtain a standardized clarity, and obtaining a target weight of the standardized clarity based on the standardized clarity, a pre-set initial weight and a first adjustment coefficient; and performing feature fusion on the clarity of each dimension according to the target weight to obtain the comprehensive video clarity of the candidate image.

[0045] The correlation coefficient may be a Pearson correlation coefficient, which is a coefficient that measures the degree of linear correlation between two variables. After obtaining the clarity of each dimension of the candidate image, the correlation coefficient between the clarity of each dimension and the standard score is calculated according to the following formula:

[0046]

[0047] Among them, X i Indicates the features corresponding to the clarity of each dimension, Cov(X i, Y) represents feature X i The difference in help defense compared to the standard rating of Y; σ Y Represents the features X i and the standard deviation of the standard score; Correlation i Represents the correlation coefficient.

[0048] Normalize the clarity of each dimension based on the correlation coefficient:

[0049]

[0050] Importance i =| Correlation i |;

[0051] Among them, NormalizedImportance i represents the normalized clarity; j represents the dimension.

[0052] The target weight of the standardized clarity is obtained according to the standardized clarity, the preset initial weight and the first adjustment coefficient:

[0053] w i ′=w i ·(1+α·NormalizedImportance i );

[0054] Among them, w i represents the initial weight; α is the first adjustment coefficient; w i ′ represents the target weight of the normalized clarity; After obtaining the standardized target weight of the clarity, the clarity of each dimension is weighted averaged and fused according to the target weight to obtain the comprehensive video clarity of the candidate image.

[0055] The above steps can comprehensively and accurately obtain the comprehensive clarity of the video through the clarity calculation results of each dimension, laying the foundation for subsequently obtaining a comprehensive target feature vector.

[0056] Step 204 : For each video in the sample video set, calculate the optical flow vector variance of the video frame of the current video, and determine the degree of change of the flow light field of the video frame of the current video based on the optical flow vector variance.

[0057] The optical flow vector variance of a video frame is used to quantify the degree of change in the optical flow field. Based on the optical flow vector variance, the degree of change in the optical flow field can be directly determined. The optical flow field is a two-dimensional vector field that describes the instantaneous motion velocity vector information of each pixel in the image. The optical flow vector variance of the current video frame is calculated according to the following formula:

[0058]

[0059] Among them, v k Represents the optical flow vector of the kth frame; Represents the mean of the optical flow vector; N represents the total number of frames; MotionVariance represents the variance of the optical flow vector.

[0060] Step 205: Determine an evaluation standard for video frame stability according to the degree of change of the light flow field, a preset basic threshold, and a second adjustment coefficient; and obtain the video frame stability of the current video based on the evaluation standard for video frame stability.

[0061] The basic threshold is a threshold pre-determined based on domain big data and is used to measure the stability of video frames. The evaluation criteria for video frame stability are determined based on the degree of change in the light flow field, the pre-determined basic threshold, and the second adjustment coefficient:

[0062] Threshold dynamic =Base Threshold+β·Motion Variance;

[0063] Where Base Threshold represents the base threshold; β represents the second adjustment coefficient; and Motion Variance represents the variance of the light field. After obtaining the evaluation criteria, the video frame stability of the current video is obtained based on the video frame stability evaluation criteria:

[0064]

[0065] Stability represents the stability value, that is, the video frame stability of the current video. If Stability is greater than Threshold dynamic , the video is considered to have large motion fluctuations and is marked as unstable; if Stability is less than or equal to Threshold dynamic , the video is considered stable.

[0066] Step 206: Perform video preprocessing on the sample video set to obtain the audio clarity and audio signal-to-noise ratio of each video in the sample video set.

[0067] Audio clarity refers to the intelligibility and noise-free nature of an audio signal. The audio signal-to-noise ratio (SNR) measures the ratio of useful information to noise in an audio signal and can be expressed in decibels. A higher SNR indicates more useful information and less noise in the signal. In this solution, a sample video set is preprocessed to obtain the audio clarity of each video in the sample video set. This includes: extracting the audio signal of the current video from each video in the sample video set; performing fast Fourier transform and Mel spectrum feature extraction on the audio signal to obtain the Mel spectrum features of the audio signal; and determining the clarity of the audio signal based on the Mel spectrum features.

[0068] The Fast Fourier Transform (FFT) is an algorithm for efficiently computing the Discrete Fourier Transform (DFT). It exploits the periodicity and symmetry of the DFT to decompose a long sequence of DFTs into a shorter sequence of DFTs, significantly reducing the computational effort. Mel-spectral feature extraction converts the frequency axis of the audio signal to the Mel-frequency scale, then calculates the energy spectrum on this scale. Mel-spectral features are then obtained through a series of processing steps. Specifically, after obtaining a set of sample videos, the audio signal from each video is extracted and converted to a unified format. After obtaining the converted audio signal, a first-order difference filter is used to pre-emphasize the audio signal to enhance the high-frequency component. The spectrogram is mapped to the Mel-frequency scale: a triangular filter bank distributed along the Mel-frequency scale is used to smooth the spectrum, eliminate harmonics, and highlight the formants of the original sound. The spectrum on the Mel-frequency scale is logarithmically processed to compress the spectrum's dynamic range. The logarithmic Mel-spectrum is then generated based on the logarithmic operation. The Mel-spectrum after the logarithmic operation is then subjected to a discrete cosine transform to obtain the Mel-spectral features. Mel-spectrogram features can reflect the clarity of audio signals by capturing the energy distribution of audio signals at different frequencies. Therefore, after obtaining the Mel-spectrogram features of audio signals, the clarity of the audio signals can be obtained based on the Mel-spectrogram features.

[0069] By performing fast Fourier transform and Mel spectrum feature extraction on audio signals, the frequency components in the audio signal can be effectively analyzed and processed. This can highlight the important frequency parts of the audio signal while suppressing or removing unnecessary noise and interference, thus improving the efficiency of audio signal processing.

[0070] In this solution, video preprocessing is performed on the sample video set to obtain the audio signal-to-noise ratio of each video in the sample video set, including: performing short-time Fourier transform on the audio signal to obtain the audio signal-to-noise ratio of the current video.

[0071] The Short-Time Fourier Transform (STFT) is a method for converting signals from the time domain to the frequency domain. It can effectively analyze the frequency components of a signal at different time points. After obtaining the audio signal, the STFT is used to extract its spectral characteristics. These spectral characteristics reflect the frequency and energy distribution of the audio signal, and the audio signal-to-noise ratio (SNR) can be derived from these spectral characteristics. Once the SNR is determined, the pre-set standard SNR can be compared with the current audio signal's SNR to assess its quality.

[0072] By extracting the spectral characteristics of the audio signal through short-time Fourier transform, the frequency components in the audio signal can be effectively analyzed and processed, and the audio signal-to-noise ratio of the audio signal can be accurately obtained.

[0073] Step 207: Obtain a target feature vector for each video based on the comprehensive video clarity, video frame stability, audio clarity, and audio signal-to-noise ratio of each video in the sample video set.

[0074] Step 208: When the pre-established initial scoring model does not meet the preset conditions, the initial scoring model is trained based on each target feature vector, the user rating of each video, and the standard rating until the initial scoring model meets the preset conditions, thereby obtaining the target scoring model.

[0075] Among them, the initial scoring model is a pre-established untrained scoring model. The preset condition may be that the preset loss function fails to converge, or the number of iterations does not reach the preset number, etc. The target scoring model can accurately evaluate the video quality. In this solution, the initial scoring model is trained based on each target feature vector, the user rating of each video, and the standard rating, including: when the initial scoring model does not meet the preset conditions, a target feature vector is selected from each target feature vector as the current target feature vector, the current target feature vector is input into the initial scoring model, and the target output vector corresponding to the current target feature vector is obtained; the loss function value of the current target feature vector is calculated based on the target output vector, the user rating corresponding to the current target feature vector, the standard rating, and the loss function of the initial scoring model, and the initial scoring model is updated based on the loss function value.

[0076] Specifically, the initial scoring model includes a pre-set loss function. After obtaining the target feature vector, the target feature vector can be input into the initial scoring model, which processes the target feature vector and outputs the output score corresponding to the target feature vector. The output score, the user score of each video, and the standard score are substituted into the loss function to obtain the loss function value. The model parameters of the initial scoring model are adjusted according to the loss function value. The loss function is shown in the following formula:

[0077]

[0078] Among them, L total Represents the loss function; α1, β1, γ1, δ1 represent the weight coefficients of each item respectively; y i Indicates the actual rating; represents the output score; p i Represents user ratings. w j represents the weight of the jth feature; M represents the total number of features; Noise k (x i ) represents the k-th noise feature x i The influence of; λ represents the weight of the noise processing intensity; w represents the weight of the model.

[0079] The loss function of this solution includes user ratings and standard ratings. While improving the accuracy of the model input results, it also pays attention to the influence of user ratings on video ratings. This ensures that the model output results meet the rating standards of both professional fields and regular users (users who watch videos), thereby improving the accuracy of the model input results.

[0080] In the technical solution of this embodiment, a sample video set and user ratings and standard ratings for each video in the sample video set are obtained. For each video in the sample video set, a preset number of frames are randomly extracted from the current video and preliminary image processing is performed on the preset number of frames to obtain candidate images corresponding to the preset number of frames. The preliminary image processing includes at least grayscale processing and denoising. The clarity of each dimension of the candidate image is calculated, and the comprehensive video clarity of the candidate image is obtained based on the clarity calculation results of each dimension. For each video in the sample video set, the variance of the optical flow vector of the video frame of the current video is calculated, and the degree of change in the light flow field of the video frame of the current video is determined based on the variance of the optical flow vector. An evaluation criterion for video frame stability is determined based on the degree of change in the light flow field, a pre-set basic threshold, and a second adjustment coefficient. The video frame stability of the current video is obtained based on the evaluation criterion for video frame stability. Video preprocessing is performed on the sample video set to obtain the audio clarity and audio signal-to-noise ratio of each video in the sample video set. A target feature vector for each video in the sample video set is obtained based on the comprehensive video clarity, video frame stability, audio clarity, and audio signal-to-noise ratio. If the pre-established initial scoring model does not meet the preset conditions, the initial scoring model is trained based on each target feature vector, the user ratings of each video, and the standard ratings until the initial scoring model meets the preset conditions, thereby obtaining the target scoring model. In the technical solution of this embodiment, a comprehensive target feature vector is obtained by combining video clarity, video frame stability, audio clarity, and audio signal-to-noise ratio. This allows the scoring model to comprehensively evaluate the video, ensuring a more comprehensive and accurate evaluation result.

[0081] Figure 3 This is a schematic diagram of the structure of a video content quality assessment device based on big data analysis provided by an embodiment of the present invention. The device is suitable for executing a video content quality assessment method based on big data analysis provided by an embodiment of the present invention. Figure 3 As shown, the device may specifically include:

[0082] A data acquisition module 301 is used to acquire a sample video set, and user ratings and standard ratings of each video in the sample video set;

[0083] The video processing module 302 is configured to perform video preprocessing on the sample video set to obtain comprehensive video clarity, video frame stability, audio clarity, and audio signal-to-noise ratio of each video in the sample video set;

[0084] a vector determination module 303 for obtaining a target feature vector for each video in the sample video set based on the video comprehensive clarity, video frame stability, audio clarity, and audio signal-to-noise ratio of each video;

[0085] The model training module 304 is used to train the pre-established initial scoring model based on each target feature vector, the user rating of each video and the standard rating when the pre-established initial scoring model does not meet the preset conditions, until the initial scoring model meets the preset conditions, thereby obtaining a target scoring model.

[0086] Optionally, the video processing module 302 is specifically configured to: for each video in the sample video set, randomly extract a preset number of frames of images from the current video, and perform preliminary image processing on the preset number of frames of images to obtain candidate images corresponding to the preset number of frames of images; wherein the preliminary image processing includes at least grayscale processing and denoising processing;

[0087] The clarity of each dimension of the candidate image is calculated, and the comprehensive video clarity of the candidate image is obtained based on the clarity calculation results of each dimension.

[0088] Optionally, the video processing module 302 is further configured to: calculate the correlation coefficient between the clarity of each dimension and the standard score according to a preset correlation coefficient calculation formula;

[0089] Normalizing the clarity of each dimension based on the correlation coefficient to obtain a standardized clarity, and obtaining a target weight of the standardized clarity according to the standardized clarity, a preset initial weight, and a first adjustment coefficient;

[0090] The clarity of each dimension is subjected to feature fusion according to the target weight to obtain the comprehensive video clarity of the candidate image.

[0091] Optionally, the video processing module 302 is further configured to: calculate, for each video in the sample video set, an optical flow vector variance of a video frame of a current video, and determine a degree of change in the flow light field of the video frame of the current video based on the optical flow vector variance;

[0092] Determining an evaluation criterion for the stability of the video frame according to a degree of change of the light flow field, a pre-set basic threshold, and a second adjustment coefficient;

[0093] The video frame stability of the current video is obtained based on the evaluation standard of the video frame stability.

[0094] Optionally, the video processing module 302 is further configured to: for each video in the sample video set, extract an audio signal of the current video from the current video;

[0095] Performing fast Fourier transform and Mel spectrum feature extraction on the audio signal to obtain Mel spectrum features of the audio signal; and obtaining the clarity of the audio signal based on the Mel spectrum features;

[0096] Performing a short-time Fourier transform on the audio signal, and obtaining an audio signal-to-noise ratio of the current video based on the transformed audio signal according to a preset standard signal-to-noise ratio.

[0097] Optionally, the model training module 304 is specifically configured to: when the initial scoring model does not meet a preset condition, select a target feature vector from the target feature vectors as a current target feature vector, input the current target feature vector into the initial scoring model, and obtain a target output vector corresponding to the current target feature vector;

[0098] A loss function value of the current target feature vector is calculated based on the target output vector, the user rating corresponding to the current target feature vector, the standard rating, and the loss function of the initial rating model, and the initial rating model is updated based on the loss function value.

[0099] The apparatus for assessing video content quality based on big data analysis provided in an embodiment of the present invention can execute the method for assessing video content quality based on big data analysis provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects. Any details not fully described in this embodiment can be referred to the description of any method embodiment of the present invention.

[0100] Figure 4 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention, referring to Figure 4 , Figure 4 The electronic device 12 shown is only an example and should not limit the functions and scope of use of the embodiments of the present application. Figure 4 As shown, electronic device 12 is implemented as a general-purpose computing device. Components of electronic device 12 may include, but are not limited to, one or more processors or processing units 16, system memory 28, and a bus 18 that connects various system components (including system memory 28 and processing unit 16).

[0101] Bus 18 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processor, or a local bus using any of a variety of bus architectures. Examples of these architectures include, but are not limited to, an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MAC) bus, an Enhanced ISA bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus.

[0102] The electronic device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by the electronic device 12, including volatile and non-volatile media, removable and non-removable media.

[0103] The system memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. The electronic device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, the storage system 34 may be configured to read and write non-removable, non-volatile magnetic media ( Figure 4 Not shown, often called a "hard drive"). Although Figure 4 Not shown, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk"), and an optical drive for reading and writing to a removable non-volatile optical disk (e.g., a CD-ROM, DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to the bus 18 via one or more data medium interfaces. The memory 28 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the various embodiments of the present application.

[0104] A program / utility 40 having a set (at least one) of program modules 42 may be stored, for example, in memory 28. Such program modules 42 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data, each of which, or some combination thereof, may include an implementation of a network environment. Program modules 42 generally implement the functions and / or methods of the embodiments described herein.

[0105] The electronic device 12 may also communicate with one or more external devices 14 (e.g., a keyboard, a pointing device, a display 24, etc.), one or more devices that enable a user to interact with the electronic device 12, and / or any device that enables the electronic device 12 to communicate with one or more other computing devices (e.g., a network card, a modem, etc.). Such communication may be performed through an input / output (I / O) interface 22. Furthermore, the electronic device 12 may also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 20. As shown, the network adapter 20 communicates with the other modules of the electronic device 12 via the bus 18. It should be understood that although Figure 4 Not shown, other hardware and / or software modules may be used in conjunction with the electronic device 12, including but not limited to microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0106] The processing unit 16 executes various functional applications and data processing by running the program stored in the system memory 28, such as implementing a video content quality assessment method based on big data analysis provided by an embodiment of the present invention: obtaining a sample video set, a user score and a standard score of each video in the sample video set; performing video preprocessing on the sample video set to obtain the video comprehensive clarity, video frame stability, audio clarity and audio signal-to-noise ratio of each video in the sample video set; obtaining the target feature vector of each video based on the video comprehensive clarity, video frame stability, audio clarity and audio signal-to-noise ratio of each video in the sample video set; when a pre-established initial scoring model does not meet the preset conditions, the initial scoring model is trained based on each target feature vector, the user score and the standard score of each video until the initial scoring model meets the preset conditions, thereby obtaining a target scoring model.

[0107] An embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon. When executed by a processor, the program implements a video content quality assessment method based on big data analysis as provided in all embodiments of the present invention: obtaining a sample video set, user ratings, and standard ratings for each video in the sample video set; performing video preprocessing on the sample video set to obtain the video comprehensive clarity, video frame stability, audio clarity, and audio signal-to-noise ratio of each video in the sample video set; obtaining a target feature vector for each video based on the video comprehensive clarity, video frame stability, audio clarity, and audio signal-to-noise ratio of each video in the sample video set; and when a pre-established initial scoring model does not meet a preset condition, training the initial scoring model based on each target feature vector, user ratings, and standard ratings for each video until the initial scoring model meets the preset condition, thereby obtaining a target scoring model. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electronic device, apparatus, or component that is electrical, magnetic, optical, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection having one or more conductors, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction-executing electronic device, apparatus, or device.

[0108] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transfer a program for use by or in conjunction with an instruction-executing electronic device, apparatus, or device.

[0109] Program code embodied on a computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0110] Computer program code for performing the operations of the present invention can be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0111] An embodiment of the present invention further provides a computer program product, including a computer program, which, when executed by a processor, implements the text analysis method provided in any embodiment of the present application.

[0112] The computer program product may be implemented by writing computer program code for performing the operations of the present invention in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0113] Note that the above are only preferred embodiments of the present invention and the technical principles employed. Those skilled in the art will appreciate that the present invention is not limited to the specific embodiments herein, and that various obvious changes, readjustments, and substitutions are possible for those skilled in the art without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments and may include many other equivalent embodiments without departing from the scope of the present invention. The scope of the present invention is determined by the scope of the appended claims.

Claims

1. A video content quality assessment method based on big data analysis, characterized in that: The method comprises: Obtain a sample video collection, and user ratings and standard ratings of each video in the sample video collection; Performing video preprocessing on the sample video set to obtain comprehensive video clarity, video frame stability, audio clarity, and audio signal-to-noise ratio of each video in the sample video set; performing video preprocessing on the sample video set to obtain comprehensive video clarity of each video in the sample video set, including: calculating the clarity of each dimension of a candidate image corresponding to each video in the sample video set based on each video in the sample video set, and obtaining comprehensive video clarity of the candidate image based on the clarity calculation results of each dimension; the clarity of each dimension includes Laplace clarity, discrete differential operator clarity, and Fourier clarity; obtaining comprehensive video clarity of the candidate image based on the clarity calculation results of each dimension, including: calculating the correlation coefficient between the clarity of each dimension and the standard score according to a pre-set correlation coefficient calculation formula; standardizing the clarity of each dimension based on the correlation coefficient to obtain standardized clarity, and obtaining a target weight of the standardized clarity according to the standardized clarity, a pre-set initial weight, and a first adjustment coefficient; performing feature fusion on the clarity of each dimension according to the target weight to obtain comprehensive video clarity of the candidate image; Obtaining a target feature vector for each video based on comprehensive video clarity, video frame stability, audio clarity, and audio signal-to-noise ratio of each video in the sample video set; When the pre-established initial scoring model does not meet the preset conditions, the initial scoring model is trained based on each target feature vector, the user rating of each video and the standard rating until the initial scoring model meets the preset conditions, thereby obtaining a target scoring model; wherein, the target scoring model is used to perform quality assessment on the video to be processed.

2. The method according to claim 1, characterized in that Calculating the clarity of each dimension of a candidate image corresponding to each video in the sample video set includes: For each video in the sample video set, randomly extract a preset number of frames of images from the current video, and perform preliminary image processing on the preset number of frames of images to obtain candidate images corresponding to the preset number of frames of images; wherein the preliminary image processing includes at least grayscale processing and denoising processing; Calculate the clarity of each dimension of the candidate image.

3. The method according to claim 1, characterized in that Performing video preprocessing on the sample video set to obtain video frame stability of each video in the sample video set includes: For each video in the sample video set, calculating the optical flow vector variance of the video frame of the current video, and determining the degree of change of the flow light field of the video frame of the current video based on the optical flow vector variance; Determining an evaluation criterion for the stability of the video frame according to a degree of change of the light flow field, a pre-set basic threshold, and a second adjustment coefficient; The video frame stability of the current video is obtained based on the evaluation standard of the video frame stability.

4. The method according to claim 1, wherein Performing video preprocessing on the sample video set to obtain audio clarity and audio signal-to-noise ratio of each video in the sample video set includes: For each video in the sample video set, extracting the audio signal of the current video from the current video; Performing fast Fourier transform and Mel spectrum feature extraction on the audio signal to obtain Mel spectrum features of the audio signal; and obtaining the clarity of the audio signal based on the Mel spectrum features; Performing short-time Fourier transform on the audio signal to obtain an audio signal-to-noise ratio of the current video.

5. The method according to claim 1, wherein The initial scoring model is trained based on the target feature vectors, the user ratings of the videos, and the standard ratings, including: When the initial scoring model does not meet the preset conditions, selecting a target feature vector from the target feature vectors as the current target feature vector, inputting the current target feature vector into the initial scoring model, and obtaining a target output vector corresponding to the current target feature vector; A loss function value of the current target feature vector is calculated based on the target output vector, the user rating corresponding to the current target feature vector, the standard rating, and the loss function of the initial rating model, and the initial rating model is updated based on the loss function value.

6. A computer program product comprising a computer program, characterized in that When executed by a processor, the computer program implements a video content quality assessment method based on big data analysis as described in any one of claims 1 to 5.

7. A video content quality assessment device based on big data analysis, characterized in that: include: A data acquisition module is used to obtain a sample video set, user ratings and standard ratings of each video in the sample video set; a video processing module, configured to perform video preprocessing on the sample video set to obtain comprehensive video clarity, video frame stability, audio clarity, and audio signal-to-noise ratio of each video in the sample video set; Performing video preprocessing on the sample video set to obtain the comprehensive video clarity of each video in the sample video set, including: calculating the clarity of each dimension of the candidate image corresponding to each video based on each video in the sample video set, and obtaining the comprehensive video clarity of the candidate image based on the clarity calculation results of each dimension; the clarity of each dimension includes Laplace clarity, discrete differential operator clarity and Fourier clarity; obtaining the comprehensive video clarity of the candidate image based on the clarity calculation results of each dimension, including: calculating the correlation coefficient between the clarity of each dimension and the standard score according to a pre-set correlation coefficient calculation formula; standardizing the clarity of each dimension based on the correlation coefficient to obtain a standardized clarity, and obtaining a target weight of the standardized clarity according to the standardized clarity, a pre-set initial weight and a first adjustment coefficient; performing feature fusion on the clarity of each dimension according to the target weight to obtain the comprehensive video clarity of the candidate image; a vector determination module, configured to obtain a target feature vector for each video in the sample video set based on the comprehensive video clarity, video frame stability, audio clarity, and audio signal-to-noise ratio of each video; The model training module is used to train the pre-established initial scoring model based on each target feature vector, the user rating of each video and the standard rating when the pre-established initial scoring model does not meet the preset conditions, until the initial scoring model meets the preset conditions, thereby obtaining a target scoring model; wherein the target scoring model is used to perform quality assessment on the video to be processed.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, it implements a video content quality assessment method based on big data analysis as described in any one of claims 1 to 5.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, a video content quality assessment method based on big data analysis as described in any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • An intelligent image fusion method based on target feature driving

    CN109035188A

  • Video quality evaluation method and device, equipment and medium

    CN115037926A