Video quality analysis method and system based on visual technology

Through the video quality analysis method based on vision technology, the video quality is automatically analyzed using convolutional neural network and FAST feature point detection algorithm, solving the problems of low manual audit efficiency and insufficient objectivity in the existing technology, and achieving efficient and accurate video quality analysis.

CN120091124AInactive Publication Date: 2025-06-03ANHUI SHARETRONIC DATA TECHNOLOGY CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510174871.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-06-03
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing video platforms rely on manual review of video quality, are inefficient and are subject to the level and objectivity of the auditor.

Method used

The video quality analysis method based on vision technology is adopted, by obtaining historical video data, extracting video features and establishing a quality score model, and using convolutional neural network and FAST feature point detection algorithm, the quality of the video to be evaluated is automatically analyzed.

Benefits of technology

Efficient and accurate video quality analysis is achieved, reducing the dependence on the level of auditors, and improving the universality and objectivity of video quality analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120091124A_ABST
    Figure CN120091124A_ABST
Patent Text Reader

Abstract

The invention relates to the field of videos, and particularly discloses a video quality analysis method and system based on a visual technology, and the method comprises the steps: obtaining historical video data and a quality score of the historical video data, and extracting video features in the historical video data; establishing a quality score model through the video features; and obtaining a to-be-evaluated video, obtaining a to-be-evaluated feature vector through the convolutional neural network, substituting the to-be-evaluated feature vector into the quality score model to obtain the quality score of each subitem of the to-be-evaluated video, multiplying the quality score of each subitem by the weight factor of the subitem, and adding to obtain the quality score of the to-be-evaluated video. According to the method, the quality score model is constructed and the feature vector of the to-be-evaluated video is extracted in different modes, so that the inaccuracy of modeling and feature vector extraction in a single mode is avoided, the video quality is efficiently and accurately analyzed, the method is not limited by video contents, and the universality of video quality analysis can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of videos, and specifically to a method and system for video quality analysis based on vision technology. Background Art

[0002] When a video is published on a video platform, the video platform needs to review the quality of the video to ensure that the video quality meets people's viewing needs.

[0003] Existing video platforms all use manual methods to review the quality of videos. Although the review of video quality can be achieved, the work efficiency is low, and the types of video content are diverse. Moreover, it is limited by the level of the reviewers themselves and is not objective enough. Summary of the Invention

[0004] The purpose of the present invention is to provide a method and system for video quality analysis based on vision technology to solve the problems raised in the above background art.

[0005] To achieve the above purpose, the present invention provides the following technical solutions:

[0006] A method for video quality analysis based on vision technology, the method comprising:

[0007] Obtain historical video data and the quality scores of the historical video data, extract video features from the historical video data, and the video features include jitter features, loss features, frame rate features, video bit rate features, resolution features, delay features, and synchronization rate features;

[0008] Establish a coordinate system with a certain point of the same type of video feature, substitute all the same type of video features into the coordinate system, use the FAST feature point detection algorithm to select feature points, obtain multiple feature points, form feature vectors with adjacent feature points, divide all the feature vectors into training samples and validation samples, project the training samples into a three-dimensional space to obtain an initial model, substitute the validation samples into the initial model, and if the verification passes, the initial model is the quality score model;

[0009] Obtain the video to be evaluated and obtain the feature vector to be evaluated through a convolutional neural network, substitute the feature vector to be evaluated into the quality score model, obtain the quality scores of each sub-item of the video to be evaluated, and the sum of the quality scores of each sub-item multiplied by the weight factor of the sub-item is the quality score of the video to be evaluated.

[0010] As a further scheme of the present invention: the step of extracting video features from the historical video data includes:

[0011] Obtain the jitter parameter, loss parameter, frame rate parameter, video bit rate parameter, resolution parameter, delay parameter, and synchronization rate parameter in the historical video data;

[0012] Take the first derivative of the jitter parameter, loss parameter, frame rate parameter, video bit rate parameter, resolution parameter, delay parameter, and synchronization rate parameter with respect to time to obtain the jitter feature, loss feature, frame rate feature, video bit rate feature, resolution feature, delay feature, and synchronization rate feature.

[0013] As a further solution of the present invention: Before the adjacent feature points form a feature vector, it further includes:

[0014] Calculate the mean value of the distances between adjacent feature points and remove the feature points with large errors.

[0015] As a further solution of the present invention: The step of calculating the mean value and variance of the distances between adjacent feature points and removing the feature points with large errors includes:

[0016] Calculate the distances between adjacent feature points to obtain the mean value of the distances;

[0017] Subtract the mean value of the distances from the distances and divide by the mean value of the distances to calculate the distance difference percentage;

[0018] Compare the distance difference percentage with a preset threshold and remove the feature points with a distance difference percentage greater than the preset threshold.

[0019] As a further solution of the present invention: The step of obtaining the video to be evaluated and obtaining the feature vector to be evaluated through a convolutional neural network includes:

[0020] Obtain the video to be evaluated, input the video to be evaluated into the first convolutional layer of the convolutional neural network to obtain the first feature image;

[0021] Input the first feature image into the second convolutional layer to obtain the second feature image;

[0022] Pass the second feature image through the dropout layer to obtain the feature vector to be evaluated.

[0023] The technical solution of the present invention also provides a video quality analysis system based on vision technology, and the system includes:

[0024] An acquisition module, configured to acquire historical video data and the quality score of the historical video data, and extract video features from the historical video data, where the video features include jitter features, loss features, frame rate features, video bit rate features, resolution features, delay features, and synchronization rate features;

[0025] A model generation module, which is used to establish a coordinate system with a certain point of the same type of video features, substitute all the same type of video features into the coordinate system, use the FAST feature point detection algorithm to select feature points, obtain multiple feature points, form feature vectors with adjacent feature points, divide all the feature vectors into training samples and verification samples, project the training samples into a three-dimensional space to obtain an initial model, substitute the verification samples into the initial model, and if the verification passes, the initial model is the quality score model;

[0026] A score generation module, which is used to obtain the video to be evaluated and obtain the feature vector to be evaluated through a convolutional neural network, substitute the feature vector to be evaluated into the quality score model, obtain the quality scores of each sub-item of the video to be evaluated, and multiply the quality scores of each sub-item by the weight factor of the sub-item and add them up to obtain the quality score of the video to be evaluated.

[0027] As a further solution of the present invention: the acquisition module includes:

[0028] A video acquisition unit, which is used to acquire historical video data and the quality score of the historical video data;

[0029] A parameter acquisition unit, which is used to acquire the jitter parameter, loss parameter, frame rate parameter, video bit rate parameter, resolution parameter, delay parameter and synchronization rate parameter in the historical video data;

[0030] A feature acquisition unit, which is used to respectively calculate the first-order derivative of time for the jitter parameter, loss parameter, frame rate parameter, video bit rate parameter, resolution parameter, delay parameter and synchronization rate parameter to obtain the jitter feature, loss feature, frame rate feature, video bit rate feature, resolution feature, delay feature and synchronization rate feature.

[0031] As a further solution of the present invention: the model generation module includes:

[0032] A coordinate system unit, which is used to establish a coordinate system with a certain point of the same type of video features and substitute all the same type of video features into the coordinate system;

[0033] A feature point selection unit, which is used to select feature points by using the FAST feature point detection algorithm to obtain multiple feature points;

[0034] A removal unit, which is used to calculate the mean value of the distances between adjacent feature points and remove the feature points with larger errors;

[0035] A modeling unit, which is used to form feature vectors with adjacent feature points, divide all the feature vectors into training samples and verification samples, project the training samples into a three-dimensional space to obtain an initial model, substitute the verification samples into the initial model, and if the verification passes, the initial model is the quality score model.

[0036] As a further solution of the present invention: the removal unit includes:

[0037] A mean calculation unit, which is used to calculate the distance between adjacent feature points and obtain the mean value of the distances.

[0038] A difference calculation unit, which is used to subtract the mean value of the distances from the distances and divide the result by the mean value of the distances to calculate the distance difference percentage.

[0039] A comparison unit, which is used to compare the distance difference percentage with a preset threshold and remove the feature points with a distance difference percentage greater than the preset threshold.

[0040] As a further solution of the present invention: the score generation module includes:

[0041] A first convolution unit, which is used to obtain a video to be evaluated, input the video to be evaluated into the first convolution layer of a convolutional neural network, and obtain a first feature image.

[0042] A second convolution unit, which is used to input the first feature image into the second convolution layer and obtain a second feature image.

[0043] A feature vector generation unit, which is used to pass the second feature image through a dropout layer to obtain a feature vector to be evaluated.

[0044] A sub-score generation unit, which is used to substitute the feature vector to be evaluated into a quality score model to obtain the sub-item quality scores of the video to be evaluated.

[0045] A score generation unit, which is used to multiply the sub-item quality scores by the weight factors of the sub-items and add them up to obtain the quality score of the video to be evaluated.

[0046] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention adopts different methods to construct a quality score model and extract the feature vector of the video to be evaluated, avoiding the inaccuracy of modeling and extracting the feature vector by a single method, thereby accurately analyzing the video quality with high efficiency, not being restricted by the video content, and being able to effectively improve the universality of video quality analysis. Description of the Drawings

[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention.

[0048] Figure 1 It is a flow block diagram of a video quality analysis method based on vision technology.

[0049] Figure 2 It is a partial first sub-flow block diagram of a video quality analysis method based on vision technology.

[0050] Figure 3It is a block diagram of the third sub - process of the video quality analysis method based on vision technology.

[0051] Figure 4 It is a block diagram of the composition structure of the video quality analysis system based on vision technology.

[0052] Figure 5 It is a block diagram of the composition structure of the acquisition module in the video quality analysis system based on vision technology.

[0053] Figure 6 It is a block diagram of the composition structure of the model generation module in the video quality analysis system based on vision technology.

[0054] Figure 7 It is a block diagram of the composition structure of the removal unit in the video quality analysis system based on vision technology.

[0055] Figure 8 It is a block diagram of the composition structure of the score generation module in the video quality analysis system based on vision technology. Detailed implementation manners

[0056] Currently, it is still necessary for manual judgment on the situation captured by the camera to take measures. The monitoring personnel monitor multiple cameras simultaneously and rely on the experience of the monitoring personnel for judgment. It may judge unsafe behaviors as safe behaviors due to the experience of the monitoring personnel or the monitoring personnel may not be able to detect unsafe behaviors in time.

[0057] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0058] Embodiment 1: Figure 1 It is a flow chart of the video quality analysis method based on vision technology. In the embodiment of the present invention, a video quality analysis method based on vision technology, the method includes:

[0059] Step S100: Obtain historical video data and the quality score of the historical video data, extract video features from the historical video data, and the video features include jitter features, loss features, frame rate features, video bit rate features, resolution features, delay features and synchronization rate features;

[0060] Visual technology uses a computer to simulate the visual function of humans, extract information from the images of objective things, process and understand it, and finally apply it to actual detection, measurement, and control. Obtaining historical video data means observing historical videos and extracting video features from the historical video data. Knowing the historical video data and the corresponding quality scores can help us understand the quality evaluation criteria, determine what types of video features are related to video quality, and then extract these corresponding video features from the historical video data to facilitate the subsequent establishment of a quality evaluation model. There are many factors related to video quality, among which the main features include the degree of picture jitter, the degree of picture loss, the frame rate of the video, the video bit rate, the picture resolution, whether there is a delay in the picture, and whether the picture and sound are synchronized. The greater the degree of picture jitter, the worse the video quality; the greater the degree of picture loss, the more blurred the video and the worse the video quality; the greater the frame rate and video bit rate of the video, the clearer the video; the higher the video resolution, the more details the video contains and the higher the video quality; the less or shorter the delay in the video, the higher the video quality; the higher the synchronization rate of the video's picture and sound, the higher the video quality. These together constitute the criteria for judging video quality. Therefore, it is necessary to extract the jitter features, loss features, frame rate features, video bit rate features, resolution features, delay features, and synchronization rate features from the historical video data. The quality score includes the total video quality score and the quality scores of each sub-item, and then the weight factors of each sub-item can be determined to facilitate subsequent calculations.

[0061] Step S200: Establish a coordinate system with a certain point of the same type of video feature, substitute all the same type of video features into this coordinate system, use the FAST feature point detection algorithm to select feature points, obtain multiple feature points, form feature vectors with adjacent feature points, divide all feature vectors into training samples and validation samples, project the training samples into a three-dimensional space to obtain an initial model, substitute the validation samples into the initial model, and if the validation passes, the initial model is the quality score model.

[0062] Video features of the same type are grouped together for comparison. For example, there are multiple jitter features in each historical video. One of the jitter features is used as the origin of the coordinate system to form a jitter coordinate system. The remaining jitter features will be represented as different points in the jitter coordinate system. The FAST feature point detection algorithm has high computational efficiency and can select feature points in a very short time. The FAST feature point detection algorithm defines a feature point as a pixel point that is in a different region from enough pixel points in its surrounding area. For a grayscale image, that is, the gray value of this point is different from the gray values of enough pixel points in its surrounding area, then this pixel point is a feature point. The detailed calculation steps of this algorithm are as follows: Select a coordinate point from the picture and obtain the pixel value of this point. Next, determine whether this point is a feature point; Select a Bresenham circle with a radius of three centered at the selected point's coordinates (a discrete algorithm for calculating the trajectory of a circle, obtaining integer-level circle trajectory points). Generally, there are 16 points on this circle. Now select a threshold, assume it is t. The key step is that if there are N consecutive pixel points among these 16 points, and the difference between their brightness values and the pixel value of the center point is greater than or less than t, then this point is a feature point. (The value of n is generally 12 or 9). The FAST feature point detection algorithm proposes a method of non-maximum suppression to eliminate this situation. The specific method is as follows: 1. Calculate the response size (score function) VV for each detected feature point. Here, VV is defined as the sum of the absolute deviations between the center point and its surrounding 16 pixel points. 2. Consider two adjacent feature points and compare their VV values. 3. The point with a lower VV value will be deleted. The above is the principle of fast feature point detection. Two adjacent feature points can form a feature vector, so the feature vectors form the total sample. The total sample is randomly divided into training samples and validation samples. Preferably, the ratio of the number of training samples to the number of validation samples is 4:1, which can quickly realize the training and validation of the model and has high overall efficiency. Project all the training samples into three-dimensional space, and the initial model can be obtained using the principle of vector projection. At this time, the initial model does not mean that the model has been successfully established. The initial model still needs to be verified. Substitute the validation samples into the initial model. When the verification passes or the number of iterations reaches the requirement, the performance of the initial model meets the requirement, which is the jitter quality score model. According to the above similar operations, the missing quality score model, frame rate quality score model, video bitrate quality score model, resolution quality score model, delay quality score model, and synchronization rate quality score model can be obtained. At this time, all the quality score models of the sub-items can be obtained.

[0063] Step S300: Obtain the video to be evaluated and use a convolutional neural network to obtain the feature vector to be evaluated. Substitute the feature vector to be evaluated into the quality score model to obtain the quality scores of each sub-item of the video to be evaluated. Multiply the quality score of each sub-item by the weight factor of the sub-item and sum them up to obtain the quality score of the video to be evaluated.

[0064] The convolutional neural network includes an input layer, multiple convolutional layers, ReLU layers, mean-pooling layers, fully connected layers, and Dropout layers. The video to be evaluated extracts the feature vectors of each sub-item of the video to be evaluated through this convolutional neural network. Substitute the feature vector of each sub-item into the quality score model of the sub-item, and the quality score of the sub-item in the video to be evaluated can be obtained. Then multiply it by the weight factor of the sub-item to obtain the comprehensive quality score of the video to be evaluated. The quality score of the sub-item of the jitter feature is w1, the weight factor of the sub-item of the jitter feature is x1, the quality score of the sub-item of the lost feature is w2, the weight factor of the sub-item of the lost feature is x2, the quality score of the sub-item of the frame rate feature is w3, the weight factor of the sub-item of the frame rate feature is x3, the quality score of the sub-item of the video bit rate feature is w4, the weight factor of the sub-item of the video bit rate feature is x4, the quality score of the sub-item of the resolution feature is w5, the weight factor of the sub-item of the resolution feature is x5, the quality score of the sub-item of the delay feature is w6, the weight factor of the sub-item of the delay feature is x6, the quality score of the sub-item of the synchronization rate feature is w7, and the weight factor of the sub-item of the synchronization rate feature is x7. At this time, the comprehensive quality score of the video to be evaluated is w1*x1 + w2*x2 + w3*x3 + w4*x4 + w5*x5 + w6*x6 + w7*x7, and the video quality of the video to be evaluated can be judged, with high overall efficiency and not being restricted by the knowledge level of the reviewers.

[0065] Figure 2 It is a partial first sub-process block diagram of a video quality analysis method based on vision technology. The steps of extracting video features from historical video data include:

[0066] Obtain the jitter parameter, loss parameter, frame rate parameter, video bit rate parameter, resolution parameter, delay parameter, and synchronization rate parameter in the historical video data;

[0067] Determine which factors are used to judge video quality, and extract the relevant parameters of these factors in the historical video data for subsequent extraction of corresponding features.

[0068] Respectively take the first-order derivative of the jitter parameter, loss parameter, frame rate parameter, video bit rate parameter, resolution parameter, delay parameter, and synchronization rate parameter with respect to time to obtain the jitter feature, loss feature, frame rate feature, video bit rate feature, resolution feature, delay feature, and synchronization rate feature.

[0069] The first-order derivative of the jitter parameter in time is the sudden change rate of the jitter. The sudden change rate of jitter in high-quality videos is generally small, so that the video can be played continuously with high quality. Therefore, the sudden change rate of jitter can represent the jitter characteristics of the video. The loss characteristic is the sudden change rate of the degree of picture loss, the frame rate characteristic is the sudden change rate of the frame rate, the video bit rate characteristic is the sudden change rate of the video bit rate, the resolution characteristic is the sudden change rate of the resolution, the delay characteristic is the sudden change rate of the delay degree, and the synchronization rate characteristic is the sudden change rate of the synchronization rate.

[0070] Before the adjacent feature points form a feature vector, the following step is also included:

[0071] Calculate the mean distance between adjacent feature points and remove feature points with large errors.

[0072] Not all feature points are accurate, and some feature points with large errors need to be removed. At this time, the distance mean method can be used to filter out feature points with large errors and remove them, so that the created quality score model is more accurate.

[0073] The step of calculating the mean of the distances between adjacent feature points and removing feature points with large errors comprises:

[0074] Calculate the distance between adjacent feature points and get the mean of the distance;

[0075] Knowing the coordinates of the feature points in the coordinate system, we can calculate the distance between two feature points. For example, if the coordinates of the two feature points are (a1, b1) and (a2, b2), the square of the distance s between the two feature points is s. 2 =(a1-a2) 2 + (b1-b2) 2 , and then find its arithmetic square root to get the distance s between two feature points, find the distances between all adjacent feature points, and then divide it by the number of distances to get the mean of the distances.

[0076] The distance difference percentage was calculated by subtracting the distance mean and dividing by the distance mean;

[0077] The difference between each distance and the mean distance divided by the mean distance is the distance difference percentage. This parameter actually represents whether the difference between the distances is too large. When the distance difference is not large, the feature points of this distance source will also cluster together, and such feature points can represent its characteristics.

[0078] The distance difference percentage is compared with the preset threshold, and feature points whose distance difference percentage is greater than the preset threshold are removed.

[0079] When the percentage of distance difference exceeds the preset threshold, it indicates that the feature points from this distance source are discrete from the remaining feature points, and their relationship is not close. At this time, the feature points from this distance source are removed, and the remaining feature points can better represent the features of the video.

[0080] Figure 3 It is a partial third sub-process block diagram of the video quality analysis method based on vision technology. The steps of obtaining the video to be evaluated and obtaining the feature vector to be evaluated through a convolutional neural network include:

[0081] Obtain the video to be evaluated, input the video to be evaluated into the first convolutional layer of the convolutional neural network, and obtain the first feature image;

[0082] Crop the people or objects in the video to be evaluated, find the accurate area of the contour of the people or objects, then normalize the cropped image to a specific number of pixels, and then send the normalized image into the first convolutional layer. The first convolutional layer is connected to a ReLU layer and a mean-pooling layer. The first convolutional layer includes multiple convolutional kernels. The feature image after passing through the first convolutional layer is the feature image corresponding to the number of convolutional kernels of the first convolutional layer. This feature image is the first feature image. Compared with using sigmoid and tanh as activation functions, the computational amount is large. When calculating the error gradient by backpropagation, the derivative calculation amount is also large, and the sigmoid and tanh functions are prone to saturation, resulting in the disappearance of gradients, that is, when approaching convergence, the transformation is too slow, causing information loss. The ReLU layer will make the output of some neurons be 0, resulting in sparsity, which not only alleviates overfitting, but also is closer to the real neuron activation model, overcoming the disappearance of gradients. Without unsupervised pre-training (that is, training the first hidden layer of the network, then training the second layer, etc., and finally using the trained network parameter values as the initial values of the overall network parameters), it converges significantly faster than the sigmoid and tanh activation functions. The mean-pooling layer compresses the feature image and extracts the main features.

[0083] Input the first feature image into the second convolutional layer to obtain the second feature image;

[0084] The second convolutional layer includes multiple convolutional kernels. The feature image after passing through the second convolutional layer is the feature image corresponding to the number of convolutional kernels of the second convolutional layer. After passing through the second mean-pooling layer, it then enters two fully connected layers and the ReLU layer and dropout layer connected to the fully connected layers.

[0085] Pass the second feature image through the dropout layer to obtain the feature vector to be evaluated.

[0086] The dropout layer weakens the joint adaptability between neuron nodes, enhances the generalization ability. The dropout layer

[0087] When training the model, it randomly makes the weights of some hidden layer nodes in the network not work to prevent the model from overfitting. It is a regularization method to improve the generalization ability. As a result, the output feature vector to be evaluated is more accurate, and it can better show the characteristics of each sub-item of the video to be evaluated.

[0088] Embodiment 2: Figure 4 It is a block diagram of the composition structure of a video quality analysis system based on vision technology. In an embodiment of the present invention, a video quality analysis system based on vision technology, the system includes:

[0089] An acquisition module, used to acquire historical video data and the quality score of the historical video data, and extract video features from the historical video data. The video features include jitter features, loss features, frame rate features, video bit rate features, resolution features, delay features, and synchronization rate features;

[0090] A model generation module, used to establish a coordinate system with a certain point of the same type of video features, substitute all the same type of video features into the coordinate system, use the FAST feature point detection algorithm to select feature points, obtain multiple feature points, form feature vectors with adjacent feature points, divide all the feature vectors into training samples and verification samples, project the training samples into a three-dimensional space to obtain an initial model, and substitute the verification samples into the initial model. If the verification passes, the initial model is the quality score model;

[0091] A score generation module, used to acquire the video to be evaluated and obtain the feature vector to be evaluated through a convolutional neural network, substitute the feature vector to be evaluated into the quality score model, obtain the quality scores of each sub-item of the video to be evaluated, and the sum of multiplying the quality scores of each sub-item by the weight factor of the sub-item is the quality score of the video to be evaluated.

[0092] The acquisition module includes:

[0093] A video acquisition unit, used to acquire historical video data and the quality score of the historical video data;

[0094] A parameter acquisition unit, used to acquire the jitter parameter, loss parameter, frame rate parameter, video bit rate parameter, resolution parameter, delay parameter, and synchronization rate parameter in the historical video data;

[0095] A feature acquisition unit, used to respectively take the first derivative of the jitter parameter, loss parameter, frame rate parameter, video bit rate parameter, resolution parameter, delay parameter, and synchronization rate parameter with respect to time to obtain the jitter feature, loss feature, frame rate feature, video bit rate feature, resolution feature, delay feature, and synchronization rate feature.

[0096] The model generation module includes:

[0097] A coordinate system unit, which is used to establish a coordinate system with a certain point of the same type of video feature and substitute all the same type of video features into the coordinate system;

[0098] A feature point selection unit, which is used to select feature points by using the FAST feature point detection algorithm to obtain multiple feature points;

[0099] A removal unit, which is used to calculate the mean value of the distances between adjacent feature points and remove the feature points with larger errors;

[0100] A modeling unit, which is used to form feature vectors from adjacent feature points, divide all feature vectors into training samples and verification samples, project the training samples into a three-dimensional space to obtain an initial model, substitute the verification samples into the initial model, and if the verification passes, the initial model is the quality score model.

[0101] The removal unit includes:

[0102] A mean value calculation unit, which is used to calculate the distances between adjacent feature points and obtain the mean value of the distances;

[0103] A difference calculation unit, which is used to subtract the mean value of the distances from the distances and divide by the mean value of the distances to calculate the distance difference percentage;

[0104] A comparison unit, which is used to compare the distance difference percentage with a preset threshold and remove the feature points with a distance difference percentage greater than the preset threshold.

[0105] The score generation module includes:

[0106] A first convolution unit, which is used to obtain the video to be evaluated, input the video to be evaluated into the first convolution layer of the convolutional neural network to obtain a first feature image;

[0107] A second convolution unit, which is used to input the first feature image into the second convolution layer to obtain a second feature image;

[0108] A feature vector generation unit, which is used to pass the second feature image through a dropout layer to obtain a feature vector to be evaluated;

[0109] A sub-score generation unit, which is used to substitute the feature vector to be evaluated into the quality score model to obtain the quality scores of each sub-item of the video to be evaluated;

[0110] A score generation unit, which is used to multiply the quality scores of each sub-item by the weight factor of the sub-item and add them up to obtain the quality score of the video to be evaluated.

[0111] All functions that can be achieved by the video quality analysis method based on vision technology are completed by a computer device. The computer device includes one or more processors and one or more memories. At least one program code is stored in the one or more memories, and the program code is loaded and executed by the one or more processors to implement the functions of the user behavior prediction method based on big data.

[0112] The processor fetches instructions from the memory one by one, analyzes the instructions, and then completes corresponding operations according to the requirements of the instructions, generating a series of control commands to make each part of the computer act automatically, continuously and coordinately, becoming an organic whole, realizing the input of the program, the input of data, and the operation and output of results. All arithmetic operations or logical operations generated in this process are completed by the arithmetic unit; the memory includes a read-only memory (ROM), and the read-only memory is used to store computer programs, and a protection device is provided outside the memory.

[0113] Exemplarily, a computer program can be divided into one or more modules. One or more modules are stored in the memory and executed by the processor to complete the present invention. One or more modules can be a series of computer program instruction segments that can complete specific functions, and the instruction segments are used to describe the execution process of the computer program in the terminal device.

[0114] Those skilled in the art can understand that the description of the above service device is only an example and does not constitute a limitation on the terminal device. It may include more or fewer components than the above description, or combine some components, or different components. For example, it may include input and output devices, network access devices, buses, etc.

[0115] The so-called processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The above processor is the control center of the above terminal device, and uses various interfaces and lines to connect all parts of the entire user terminal.

[0116] The above-mentioned memory can be used to store computer programs and / or modules. By running or executing the computer programs and / or modules stored in the memory and invoking the data stored in the memory, the above-mentioned processor realizes various functions of the terminal device. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as an information collection template display function, a product information release function, etc.); the data storage area can store data created according to the use of the berth status display system (such as product information collection templates corresponding to different product types, product information to be released by different product providers, etc.). In addition, the memory can include high-speed random access memory, and can also include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0117] If the modules / units integrated in the terminal device are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the modules / units in the above-mentioned embodiment system of the present invention, it can also be completed by instructing relevant hardware through a computer program. The above-mentioned computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, it can realize the functions of the above-mentioned various system embodiments. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.

[0118] It should be noted that in this article, the term "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without more limitations, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including that element.

[0119] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be similarly included in the patent protection scope of the present invention.

Claims

1. A video quality analysis method based on visual technology, characterized in that: The method comprises: Obtain historical video data and quality scores of historical video data, and extract video feature points in historical video data. Video features include jitter features, loss features, frame rate features, video bit rate features, resolution features, delay features, and synchronization rate features. A coordinate system is established with a certain point of the same video feature, and all the same video features are substituted into the coordinate system. The FAST feature point detection algorithm is used to select feature points to obtain multiple feature points. Adjacent feature points form feature vectors, and all feature vectors are divided into training samples and verification samples. The training samples are projected into three-dimensional space to obtain the initial model. The verification samples are substituted into the initial model. If the verification passes, the initial model is the quality score model. The video to be evaluated is obtained and the feature vector to be evaluated is obtained through a convolutional neural network. The feature vector to be evaluated is substituted into the quality score model to obtain the quality scores of each sub-item of the video to be evaluated. The quality score of each sub-item is multiplied by the weight factor of the sub-item and the sum is the quality score of the video to be evaluated.

2. The video quality analysis method based on visual technology according to claim 1 is characterized in that: The step of extracting video features from historical video data comprises: Obtain jitter parameters, loss parameters, frame rate parameters, video bit rate parameters, resolution parameters, delay parameters and synchronization rate parameters in historical video data; The first-order derivatives of the jitter parameters, loss parameters, frame rate parameters, video bit rate parameters, resolution parameters, delay parameters and synchronization rate parameters are calculated respectively to obtain jitter characteristics, loss characteristics, frame rate characteristics, video bit rate characteristics, resolution characteristics, delay characteristics and synchronization rate characteristics.

3. The video quality analysis method based on visual technology according to claim 1 is characterized in that: Before the adjacent feature points form a feature vector, the following step is also included: Calculate the mean distance between adjacent feature points and remove feature points with large errors.

4. The video quality analysis method based on visual technology according to claim 3 is characterized in that: The step of calculating the mean of the distances between adjacent feature points and removing feature points with large errors comprises: Calculate the distance between adjacent feature points and get the mean of the distance; The distance difference percentage was calculated by subtracting the distance mean and dividing by the distance mean; The distance difference percentage is compared with the preset threshold, and feature points whose distance difference percentage is greater than the preset threshold are removed.

5. The video quality analysis method based on visual technology according to claim 1 is characterized in that: The steps of obtaining the video to be evaluated and obtaining the feature vector to be evaluated through a convolutional neural network include: Obtain a video to be evaluated, and input the video to be evaluated into the first convolution layer of the convolutional neural network to obtain a first feature image; Input the first feature image into the second convolution layer to obtain a second feature image; The second feature image is passed through the dropout layer to obtain the feature vector to be evaluated.

6. A video quality analysis system based on visual technology, characterized in that: The system comprises: An acquisition module is used to acquire historical video data and quality scores of historical video data, and extract video features from historical video data. Video features include jitter features, loss features, frame rate features, video bit rate features, resolution features, delay features, and synchronization rate features. The model generation module is used to establish a coordinate system with a certain point of the same video feature, substitute all the same video features into the coordinate system, use the FAST feature point detection algorithm to select feature points, obtain multiple feature points, form feature vectors with adjacent feature points, divide all feature vectors into training samples and verification samples, project the training samples into three-dimensional space, obtain the initial model, substitute the verification samples into the initial model, and if the verification passes, the initial model is the quality score model; The score generation module is used to obtain the video to be evaluated and obtain the feature vector to be evaluated through a convolutional neural network, substitute the feature vector to be evaluated into the quality score model, and obtain the quality score of each sub-item of the video to be evaluated. The quality score of each sub-item is multiplied by the weight factor of the sub-item and the sum is the quality score of the video to be evaluated.

7. The video quality analysis system based on visual technology according to claim 6 is characterized in that: The acquisition module comprises: A video acquisition unit, used to acquire historical video data and quality scores of the historical video data; A parameter acquisition unit, used to acquire jitter parameters, loss parameters, frame rate parameters, video bit rate parameters, resolution parameters, delay parameters and synchronization rate parameters in historical video data; The feature acquisition unit is used to calculate the first-order time derivative of the jitter parameter, loss parameter, frame rate parameter, video bit rate parameter, resolution parameter, delay parameter and synchronization rate parameter respectively to obtain jitter characteristics, loss characteristics, frame rate characteristics, video bit rate characteristics, resolution characteristics, delay characteristics and synchronization rate characteristics.

8. The video quality analysis system based on visual technology according to claim 6, characterized in that: The model generation module includes: A coordinate system unit is used to establish a coordinate system based on a certain point of the same video feature and substitute all the same video features into the coordinate system; A feature point selection unit, used to select feature points using a FAST feature point detection algorithm to obtain multiple feature points; The removal unit is used to calculate the mean of the distances between adjacent feature points and remove feature points with large errors; The modeling unit is used to form feature vectors from adjacent feature points, divide all feature vectors into training samples and verification samples, project the training samples into three-dimensional space to obtain the initial model, substitute the verification samples into the initial model, and if the verification passes, the initial model is the quality score model.

9. The video quality analysis system based on visual technology according to claim 8, characterized in that: The removal unit comprises: A mean calculation unit is used to calculate the distance between adjacent feature points and obtain the mean of the distance; a difference calculation unit, which is used to calculate the distance difference percentage by subtracting the distance from the mean and dividing it by the mean of the distance; The comparison unit is used to compare the distance difference percentage with a preset threshold value, and remove feature points whose distance difference percentage is greater than the preset threshold value.

10. The video quality analysis system based on visual technology according to claim 6, characterized in that: The score generation module comprises: A first convolution unit is used to obtain a video to be evaluated, and input the video to be evaluated into a first convolution layer of a convolutional neural network to obtain a first feature image; A second convolution unit, used for inputting the first feature image into a second convolution layer to obtain a second feature image; A feature vector generating unit, used for passing the second feature image through a dropout layer to obtain a feature vector to be evaluated; A sub-score generating unit, used for substituting the feature vector to be evaluated into the quality score model to obtain the quality scores of each sub-item of the video to be evaluated; The score generating unit is used to multiply the quality score of each sub-item by the weight factor of the sub-item and add them together to obtain the quality score of the video to be evaluated.

Citation Information

Patent Citations

  • Target tracking method applied to videos

    CN105469427A

  • Target tracking method based on FAST (Features from Accelerated Segment Test) corner point and pyramid KLT (Kanade-Lucas-Tomasi)

    CN106570888A

  • A fast matching algorithm for image feature points in printed circuit board detection

    CN109472770A

  • Silent living body detection method, device and equipment and storage medium

    CN111368731A

  • Video quality evaluation model training method, video quality evaluation method and device

    CN114125495A