A video content adaptive optimization system based on deep learning

Through the video content adaptive optimization system based on deep learning, using video feature data for dynamic encoding and adjustment, the problem of unsatisfactory video quality in the existing technology is solved, and efficient optimization and quality improvement of different types of videos are achieved.

CN119363994BActive Publication Date: 2025-05-06BEIJING IACTIVE NETWORK
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411907865.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-24
Publication Date
2025-05-06
Estimated Expiration
2044-12-24

AI Technical Summary

Technical Problem

Existing video encoding technology cannot dynamically adjust different types of input videos, resulting in unsatisfactory video quality, especially when facing different scenarios, frame rates and network states, video quality may have problems such as distortion, picture lag and delay.

Method used

A video content adaptive optimization system based on deep learning is proposed, including video extraction module, video decoding module, encoding adjustment module, video traceback module and video evaluation module. Video feature data is generated through motion vector analysis, scene detection algorithm and texture detail analysis, decoding and dynamic adjustments are performed based on the deep learning model, and video encoding parameters are optimized to improve video quality.

Benefits of technology

Adaptive optimization of different types of input videos is achieved, video quality is improved, video quality is ensured that the video can maintain high quality in complex motion states, fast scene switching and detailed texture content, adapt to the needs of different network environments and playback devices, and reduce the impact of network bandwidth on video quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119363994B_ABST
    Figure CN119363994B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of video optimization technology, and discloses a video content adaptive optimization system based on deep learning, including: a video extraction module, a video decoding module, a coding adjustment module, a video backtracking module and a video evaluation module, wherein the video extraction module is configured to obtain video feature data according to a motion complexity feature vector, a scene switching feature vector and a texture feature vector, the video decoding module is configured to decode a plurality of video feature data based on an initial deep learning model to obtain video adjustment parameters, the coding adjustment module is configured to obtain a first output video, the video backtracking module is configured to obtain a target deep learning model, and obtain a second output video based on the first output video adjustment parameters, and the video evaluation module is configured to determine whether the video quality of the second output video meets the standard, and if not, obtain a third output video according to the historical input video adjustment parameters. The present invention has good adaptability and optimization characteristics for different input videos.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video optimization, and in particular to a video content adaptive optimization system based on deep learning. Background Art

[0002] With the increasing popularity of video content, video coding technology plays an important role in various video applications. The current mainstream video coding technology uses fixed coding parameters to process video content. This fixed coding method cannot be optimized for different types of input videos, resulting in unsatisfactory video quality. In addition, when adjusting the coding in different scenes, frame rates and network conditions, the video quality of the output video will not only be affected by the initial video content itself, but also by factors such as the transmission network bandwidth, resulting in the inability to flexibly adjust the coding, resulting in video quality distortion, screen freezes and delays.

[0003] Therefore, how to provide a video content adaptive optimization system based on deep learning is a technical problem that technical personnel in this field urgently need to solve. Summary of the invention

[0004] In view of this, the present invention proposes a video content adaptive optimization system based on deep learning, aiming to solve the problem that different types of input videos cannot be dynamically adjusted for encoding and the video quality of the output video is poor.

[0005] The present invention proposes a video content adaptive optimization system based on deep learning, comprising:

[0006] Video extraction module, video decoding module, encoding adjustment module, video backtracking module and video evaluation module;

[0007] The video extraction module is configured to use motion vector analysis to obtain a motion complexity feature vector from the input video, use a scene detection algorithm to identify scene switching of the input video to obtain a scene switching feature vector, use texture detail analysis to obtain a texture feature vector, and obtain video feature data based on the motion complexity feature vector, the scene switching feature vector and the texture feature vector;

[0008] The video decoding module is configured to establish an initial deep learning model, and decode the plurality of video feature data based on the initial deep learning model to obtain video adjustment parameters;

[0009] The encoding adjustment module is configured to adjust the video adjustment parameter according to the network status of the input video, the scene of the input video and the frame rate of the input video to obtain a first output video;

[0010] The video tracing module is configured to extract edge image features of the first output video, calculate edge feature values ​​of the first output video, and determine whether the video quality of the first output video meets the standard according to the edge feature values;

[0011] If the target is not met, the initial deep learning model is optimized by a gradient descent algorithm to obtain a target deep learning model, the first output video is processed to obtain first video feature data, and the first video feature data is substituted into the target deep learning model to obtain a first output video adjustment parameter, and a second output video is obtained based on the first output video adjustment parameter;

[0012] The video evaluation module is configured to delimit the second output video into a grayscale frame area, extract the video frame of the grayscale frame area, determine the grayscale image brightness value of each pixel based on the video frame, determine the contrast based on all the grayscale image brightness values, and determine whether the video quality of the second output video meets the standards based on the contrast. If not, adjust the parameters based on the historical input video to obtain the third output video.

[0013] Further, the input video is analyzed by motion vector to obtain a motion complexity feature vector, the scene switching of the input video is identified by a scene detection algorithm to obtain a scene switching feature vector, and the texture detail analysis is used to obtain a texture feature vector, including:

[0014] The video extraction module extracts video pixels of the input video, and calculates the motion complexity feature vector according to the video pixels. The motion complexity feature vector is obtained by the following formula:

[0015] ;

[0016] in, represents the motion complexity feature vector, Indicates the total number of video pixels. represents the optical flow vector of the i-th video pixel, μ represents the mean of the optical flow vector;

[0017] The scene detection algorithm obtains a binary variable feature by comparing the scene differences between different frames of the input video, and if there is a scene difference, the binary variable feature is equal to 1, and if there is no scene difference, the binary variable feature value is equal to 0, thereby obtaining a scene switching feature vector;

[0018] The texture detail analysis captures the local texture features of the input video by generating binary values ​​from the grayscale values ​​of the video pixels, and obtains a texture feature vector.

[0019] Further, when the video feature data is derived according to the motion complexity feature vector, the scene switching feature vector and the texture feature vector, it includes:

[0020] The video feature data is obtained by the following formula:

[0021] ;

[0022] in, Represents video feature data, represents the motion complexity feature vector, represents the scene switching feature vector, represents the texture feature vector, represents the weight of the motion complexity feature vector, represents the weight of the scene switching feature vector, represents the weight of the texture feature vector, and .

[0023] Further, the video decoding module is configured to establish an initial deep learning model, and when decoding the plurality of video feature data based on the initial deep learning model to obtain the video adjustment parameters, it includes:

[0024] The video decoding module collects a plurality of the video feature data and performs standard processing to obtain standard data, and the standard data is obtained by the following formula:

[0025] ;

[0026] in, Indicates standard data, Represents video feature data, Represents the mean of multiple video feature data, Represents the standard deviation of multiple video feature data;

[0027] The standard data is divided into a training set and a test set. The initial deep learning network model adopts a neural network model. Cross-validation combined with grid search is used to find model parameters of the neural network model, and a neural network model is established. The training set is used to fit the neural network model. The test set is substituted into the neural network model and the accuracy of the video adjustment parameters is calculated. When the accuracy of the video adjustment parameters reaches a preset accuracy threshold, the video decoding module decodes the standard data based on the neural network model to obtain the video adjustment parameters, and the video adjustment parameters include a corrected bit rate, a corrected quantization level, and a corrected frame rate.

[0028] Furthermore, the video decoding module decodes the standard data based on the neural network model to obtain video adjustment parameters, and the video adjustment parameters include a modified bit rate, a modified quantization level, and a modified frame rate, including:

[0029] When the video decoding module obtains the modified bit rate, the modified quantization level and the modified frame rate by decoding the neural network model, the neural network model outputs continuous values ​​of bit rate, continuous values ​​of quantization level and continuous values ​​of frame rate, and the continuous values ​​of bit rate, quantization level and frame rate are rounded to obtain the modified bit rate, the modified quantization level and the modified frame rate, and the rounding calculation is obtained by the following formula:

[0030] ;

[0031] ;

[0032] ;

[0033] in, Indicates the rounded bit rate, Indicates the continuous value of bit rate, Indicates the rounding quantization level, Represents a continuous value of the quantization level, Indicates the frame rate. represents a continuous value of the frame rate, the rounded bit rate is the modified bit rate, the rounded quantization level is the modified quantization level, and the rounded frame rate is the modified frame rate.

[0034] Further, when the encoding adjustment module is configured to adjust the video adjustment parameters according to the network status of the input video, the scene of the input video and the frame rate of the input video to obtain the first output video, it includes:

[0035] The encoding adjustment module determines whether to adjust the modified bit rate according to the network status of the input video to obtain a first report, determines whether to adjust the modified quantization level according to the scene of the input video to obtain a second report, determines whether to adjust the modified frame rate according to the frame rate of the input video to obtain a third report, and outputs the first report, the second report and the third report as a first output video.

[0036] Further, the video tracing module is configured to extract edge image features of the first output video, calculate edge feature values ​​of the first output video, and determine whether the video quality of the first output video meets the standard according to the edge feature values, including:

[0037] The video tracing module extracts the horizontal gradient and the vertical gradient of the first output video, and the edge feature value is obtained by the following formula:

[0038]

[0039] in, represents the edge eigenvalue, represents the horizontal gradient, represents the vertical gradient, and represents the weight coefficient, and ;

[0040] when If the value is greater than or equal to a preset characteristic value, it is determined that the video quality of the first output video meets the standard;

[0041] when If the value is less than a preset characteristic value, it is determined that the video quality of the first output video does not meet the standard.

[0042] Furthermore, if the target is not met, the initial deep learning model is optimized using a gradient descent algorithm to obtain a target deep learning model, including:

[0043] The video tracing module substitutes the test set into the neural network model to obtain a predicted value of the video adjustment parameter, and determines a cross entropy loss according to the predicted value and the combined value of the video adjustment parameter. The cross entropy loss is obtained by the following formula:

[0044]

[0045]

[0046] in, Represents the stapled value, represents the i-th predicted value, represents the cross entropy loss, represents the rounded bit rate, represents the rounding quantization level, represents the rounded frame rate;

[0047] The target deep learning model is obtained by controlling the learning of the neural network model according to the cross entropy loss.

[0048] Further, the video evaluation module is configured to delimit the second output video into a grayscale frame area, extract the video frame of the grayscale frame area, determine the grayscale image brightness value of each pixel according to the video frame, determine the contrast according to all the grayscale image brightness values, and judge whether the video quality of the second output video meets the standard according to the contrast, including:

[0049] The video evaluation module converts the second output video into a grayscale image, divides a 2x2 matrix area on the grayscale image, sets the matrix area as a grayscale frame area, counts each pixel of the video frame and records the number of all pixels, determines the contrast according to the variance of the grayscale image brightness values ​​of all pixels, and sets a contrast threshold;

[0050] When the contrast is greater than the contrast threshold, it is determined that the video quality of the second output video is qualified;

[0051] When the contrast is less than or equal to the contrast threshold, it is determined that the video quality of the second output video is unqualified.

[0052] Further, if the standard is not met, the third output video is obtained by adjusting the parameters according to the historical input video, including:

[0053] The video evaluation module counts the video adjustment parameters of each input video and establishes a video adjustment parameter set. When the video quality of the second output video is unqualified, the second output video adjustment parameters are determined according to the average of the adjustment parameters of the historical input videos in the video adjustment parameter set, and the third output video is obtained according to the second output video adjustment parameters.

[0054] Compared with the prior art, the beneficial effects of the present invention are: by analyzing the motion complexity, scene switching and texture details of the input video to generate corresponding feature vectors, and decoding and dynamic adjustment based on the deep learning model, adaptive optimization of different types of input videos is achieved. Whether in complex motion states, fast scene switching and detailed texture content, the system can ensure that the video quality is improved by dynamically adjusting the encoding, and optimize the deep learning model through the gradient descent algorithm to improve the encoding accuracy, so as to adapt to the needs of different network environments and playback devices, ensure the smooth playback of the first output video, the second output video and the third output video and reduce the impact of network bandwidth limitations on the video quality, and evaluate through the brightness value and contrast of the grayscale image to ensure that the video quality of the first output video, the second output video and the third output video meets the expected standard, thereby providing efficient video quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the detailed description of the preferred embodiments below. The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present invention. Moreover, the same reference symbols are used throughout the accompanying drawings to represent the same components. In the accompanying drawings:

[0056] Figure 1A functional block diagram of a deep learning-based video content adaptive optimization system provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0057] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments described herein. On the contrary, these embodiments are provided in order to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art. It should be noted that, in the absence of conflict, the embodiments of the present invention and the features described in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0058] See also Figure 1 As shown, this embodiment provides a video content adaptive optimization system based on deep learning, including: a video extraction module, a video decoding module, a coding adjustment module, a video backtracking module and a video evaluation module, the video extraction module is configured to use motion vector analysis to obtain a motion complexity feature vector for input video, use a scene detection algorithm to identify the scene switching of the input video to obtain a scene switching feature vector, use texture detail analysis to obtain a texture feature vector, and obtain video feature data based on the motion complexity feature vector, the scene switching feature vector and the texture feature vector. The video decoding module is configured to establish an initial deep learning model, decode multiple video feature data based on the initial deep learning model to obtain video adjustment parameters, the coding adjustment module is configured to adjust the video adjustment parameters according to the network status of the input video, the scene of the input video and the frame rate of the input video to obtain a first output video, and the video backtracking module is configured to obtain a first output video according to the network status of the input video, the scene of the input video and the frame rate of the input video. The method is configured to extract edge image features of a first output video, calculate edge feature values ​​of the first output video, and determine whether the video quality of the first output video meets the standards based on the edge feature values. If not, the initial deep learning model is optimized using a gradient descent algorithm to obtain a target deep learning model, the first output video is processed to obtain first video feature data, and the first video feature data is substituted into the target deep learning model to obtain first output video adjustment parameters, and the second output video is obtained based on the first output video adjustment parameters. The video evaluation module is configured to demarcate a grayscale frame area for the second output video, extract video frames of the grayscale frame area, determine the grayscale image brightness value of each pixel based on the video frame, determine the contrast based on all grayscale image brightness values, and determine whether the video quality of the second output video meets the standards based on the contrast. If not, the third output video is obtained based on the historical input video adjustment parameters.

[0059] Specifically, the video extraction module is responsible for real-time monitoring and analysis of the input video. By analyzing the motion vector of the input video, the motion complexity feature vector obtained reflects the moving direction of the object in the input video. The scene detection algorithm is used to identify the scene switching of the input video to obtain the scene switching feature vector, which helps to optimize the subsequent encoding strategy of switching scenes and reduce the visual quality loss caused by scene mutations. The texture complexity of the input video screen is analyzed by texture detail technology to obtain the texture feature vector. The motion complexity feature vector, scene switching feature vector and texture feature vector constitute a comprehensive description of the input video content, providing accurate video feature data for the subsequent optimization of the deep learning model. The video decoding module is responsible for establishing the initial deep learning model, parsing the video feature data and obtaining the video adjustment parameters. According to the training set, test set and deep learning algorithm, an initial deep learning model that can identify the multi-dimensional features of the input video is established. The initial deep learning model can provide customized video adjustment parameters for different input videos. The encoding adjustment module is responsible for dynamically optimizing the video adjustment parameters. It adjusts the video adjustment parameters according to the network status, the picture scene of the input video and the frame rate of the input video to generate a first output video after preliminary optimization. The video backtracking module ensures that the first output video meets the video quality by analyzing the first output video. It adopts edge image feature extraction technology to reflect the video quality of the first output video. If the quality of the first output video does not meet the standard, the initial deep learning model is optimized by the gradient descent algorithm to generate a new target deep learning model. The target deep learning model improves the video processing capability. The first output video adjustment parameters are obtained through deep processing of the target deep learning model and the second output video is generated, which realizes feedback optimization and improves the video quality of the second output video. The video evaluation module is responsible for comprehensively evaluating the second output video. It extracts the grayscale frame area in the second output video, calculates the grayscale image brightness value of the pixel points in the grayscale frame area to determine the contrast, and evaluates the video quality of the second output video. If it still does not meet the standard, the historical input video adjustment parameters are referred to generate the third output video to ensure the video quality of the third output video.

[0060] It can be understood that, combined with the analysis of deep learning models and multi-dimensional feature vectors, the system can derive video adjustment parameters based on different types of input videos to dynamically adjust the input video and perform corresponding video quality detection, thereby improving the accuracy and adaptability of optimizing the input video content and ensuring that the video quality of the first output video, the second output video, and the third output video meet the standards.

[0061] In some embodiments of the present application, when a motion complexity feature vector is obtained by analyzing an input video using a motion vector, a scene switching feature vector is obtained by identifying a scene switching feature vector of the input video using a scene detection algorithm, and a texture feature vector is obtained by analyzing a texture detail, the following steps are included:

[0062] The video extraction module extracts the video pixels of the input video and calculates the motion complexity feature vector based on the video pixels. The motion complexity feature vector is obtained by the following formula:

[0063] ;

[0064] in, represents the motion complexity feature vector, Indicates the total number of video pixels. Represents the optical flow vector of the i-th video pixel, μ represents the mean of the optical flow vector. The scene detection algorithm obtains binary variable features by comparing the scene differences between different frames of the input video. If there is a scene difference, the binary variable feature is equal to 1, and if there is no scene difference, the binary variable feature value is equal to 0, and the scene switching feature vector is obtained. The texture detail analysis captures the local texture features of the input video by generating binary values ​​from the grayscale values ​​of the video pixels, and obtains the texture feature vector.

[0065] It can be understood that the motion complexity feature vector reflects the moving direction of the object in the input video. The scene detection algorithm is used to identify the scene switching of the input video to obtain the scene switching feature vector. The scene detection algorithm obtains the binary variable feature by comparing the scene differences between different frames of the input video. The scene differences include the shape and number of objects in the input video. When the shape and number of objects in the input video change, the binary variable feature is equal to 1. When the shape and number of objects in the input video do not change, the binary variable feature value is equal to 0. The scene switching feature vector is obtained. For example, if the input video time is 5 seconds and 20 frames per second, the input video will have 100 frames. If scene switching is detected in the 1st frame, the 3rd frame and the 100th frame, the scene switching feature vector is Texture detail analysis captures the local texture features of the input video by converting the grayscale values ​​of the video pixels into binary values. For example, a video pixel of the input video is in a 3x3 pixel matrix, and the pixel matrix is ​​as follows:

[0066]

[0067] The grayscale value of the video pixel at the center of the pixel matrix is ​​70. The grayscale value of the video pixel at the center of the matrix is ​​compared with the 8 grayscale values ​​around it from the upper left corner in clockwise order. If the grayscale value is greater than or equal to 70, the output is 1, otherwise it is 0. The binary number obtained is 00001110, which is converted to decimal to get 14, which is its local texture feature. Assuming that there are 3 video pixels, the decimal conversion is 35, 20, and 16, then the texture feature vector is By comprehensively analyzing the motion complexity, scene switching, and texture details of the input video, a data foundation is laid for the subsequent calculation of video feature data.

[0068] In some embodiments of the present application, when video feature data is derived according to the motion complexity feature vector, the scene switching feature vector and the texture feature vector, the video feature data is derived by the following formula:

[0069] ;

[0070] in, Represents video feature data, represents the motion complexity feature vector, represents the scene switching feature vector, represents the texture feature vector, represents the weight of the motion complexity feature vector, represents the weight of the scene switching feature vector, represents the weight of the texture feature vector, and .

[0071] It is understandable that the video feature data comprehensively considers the motion complexity, scene changes and texture features of the input video. When different types of feature vectors are fused, their corresponding weights can reasonably allocate the weights of the feature vectors, thereby improving the adaptability and accuracy of the adaptive optimization system and providing accurate data support for subsequent decoding.

[0072] In some embodiments of the present application, the video decoding module is configured to establish an initial deep learning model, and when decoding multiple video feature data based on the initial deep learning model to obtain video adjustment parameters, the video decoding module collects multiple video feature data and performs standard processing to obtain standard data, and the standard data is obtained by the following formula:

[0073] ;

[0074] in, Indicates standard data, Represents video feature data, Represents the mean of multiple video feature data, Represents the standard deviation of multiple video feature data, divides the standard data into a training set and a test set, uses a neural network model as the initial deep learning network model, uses cross-validation combined with grid search to find the model parameters of the neural network model, establishes the neural network model, uses the training set to fit the neural network model, substitutes the test set into the neural network model and calculates the accuracy of the video adjustment parameters. When the accuracy of the video adjustment parameters reaches a preset accuracy threshold, the video decoding module decodes the standard data based on the neural network model to obtain the video adjustment parameters. The video adjustment parameters include a corrected bit rate, a corrected quantization level and a corrected frame rate.

[0075] Specifically, standard processing improves the uniformity of standard data distribution and prevents numerical problems from occurring during the calculation process. If the range of video feature data is large, data overflow or precision loss may occur. The standard data obtained by standardization will be limited to a limited interval, reducing the risk of data redundancy and precision loss. In the video decoding process, in order to improve the accuracy of the video adjustment parameters, the neural network model of the deep learning model is used to process the standard data. The standard data is divided into a training set and a test set, the neural network model is initialized, and the hyperparameters of the neural network model are collected by combining cross-validation and grid search technology. The parameters of the neural network model are adjusted to improve the decoding accuracy. During the training process, the training set is used to fit the neural network model, and then the test set is substituted into the trained neural network model for verification, and the performance of the neural network model is evaluated by calculating the accuracy of the video adjustment parameters. If the accuracy of the video adjustment parameters is greater than or equal to the preset accuracy of 70%, it is considered that the neural network model has good decoding prediction ability. Otherwise, cross-validation and grid search technology are continued to collect the hyperparameters of the neural network model, and the parameters of the neural network model are optimized so that the accuracy of the video adjustment parameters is greater than or equal to the preset accuracy.

[0076] It is understandable that when the neural network model reaches the preset accuracy, the neural network model is used to decode the standard data and output the predicted corrected bit rate, corrected quantization level and corrected frame rate. This improves the system's ability to adaptively optimize input video, ensures the smooth playback of subsequent output videos, and improves the video quality of the output videos.

[0077] In some embodiments of the present application, the video decoding module decodes the standard data based on the neural network model to obtain the video adjustment parameters, and the video adjustment parameters include the modified bit rate, the modified quantization level and the modified frame rate, including: when the video decoding module decodes the modified bit rate, the modified quantization level and the modified frame rate through the neural network model, the neural network model outputs the bit rate continuous value, the quantization level continuous value and the frame rate continuous value, and the bit rate continuous value, the quantization level continuous value and the frame rate continuous value are rounded to obtain the modified bit rate, the modified quantization level and the modified frame rate, and the rounding calculation is obtained by the following formula:

[0078] ;

[0079] ;

[0080] ;

[0081] in, Indicates the rounded bit rate, Indicates the continuous value of bit rate, Indicates the rounding quantization level, Represents a continuous value of the quantization level, Indicates the frame rate. Indicates the continuous value of frame rate. The integer bit rate is the corrected bit rate, the integer quantization level is the corrected quantization level, and the integer frame rate is the corrected frame rate.

[0082] It is understandable that the bit rate, quantization level and frame rate obtained by the video decoding module through the neural network model are continuous values. However, the bit rate increases with a fixed step size, the quantization level is also discrete, and the frame rate is an integer frame per second. Only when the continuous value is a discrete integer can the input video be effectively optimized. By rounding, it can ensure that the corrected bit rate, corrected quantization level and corrected frame rate meet the discrete limit, improve the optimization effect of the input video, and ensure the video quality and playback smoothness of the output video.

[0083] In some embodiments of the present application, when the encoding adjustment module is configured to adjust the video adjustment parameters according to the network status of the input video, the scene of the input video and the frame rate of the input video to obtain a first output video, it includes: the encoding adjustment module determines whether to adjust the corrected bit rate according to the network status of the input video to obtain a first report, determines whether to adjust the corrected quantization level according to the scene of the input video to obtain a second report, determines whether to adjust the corrected frame rate according to the frame rate of the input video to obtain a third report, and outputs the first output video according to the first report, the second report and the third report.

[0084] Specifically, the encoding adjustment module has the ability to dynamically adjust the modified bit rate, modified quantization level and modified frame rate. When the network delay of the input video transmitted to the system is less than 30 milliseconds, the network status of the input video is considered to be good, and the integer value of the modified bit rate is increased to further improve the video quality of the output video. When the network delay of the input video transmitted to the system is greater than or equal to 30 milliseconds, the network status of the input video is considered to be poor, and the integer value of the modified bit rate is maintained to enhance the video quality of the output video while reducing the network burden. The first report is obtained based on the result of maintaining the modified bit rate or increasing the integer value of the modified bit rate. When the scene of the input video is a motion scene, the integer value of the modified quantization level is reduced to retain details. When the scene of the input video is a static scene, the integer value of the modified quantization level is increased to improve the compression rate of the input video. The second report is obtained based on the result of reducing the integer value of the modified quantization level or increasing the integer value of the modified quantization level. When the frame rate of the input video is greater than or equal to 60FPS, the frame rate of the input video is considered to be good, and the integer value of the corrected frame rate is maintained. When the frame rate of the input video is less than 60FPS, the frame rate of the input video is considered to be poor, and the integer value of the corrected frame rate is increased. The third report is obtained based on the result of maintaining the integer value of the corrected frame rate or increasing the integer value of the corrected frame rate.

[0085] It is understandable that the encoding adjustment module adjusts the bit rate, quantization level and frame rate of the input video according to the first report, the second report and the third report to obtain the first output video. By monitoring the video quality and network status of the input video, the modified bit rate, modified quantization level and modified frame rate can be predicted and adjusted in real time, thereby changing the bit rate, quantization level and frame rate of the input video, optimizing the video quality of the input video to obtain the output video, which not only ensures the video quality of the output video, but also effectively utilizes network resources.

[0086] In some embodiments of the present application, the video tracing module is configured to extract edge image features of the first output video, calculate edge feature values ​​of the first output video, and determine whether the video quality of the first output video meets the standard according to the edge feature values, including: the video tracing module extracts the horizontal gradient and the vertical gradient of the first output video, and the edge feature value is obtained by the following formula:

[0087]

[0088] in, represents the edge eigenvalue, represents the horizontal gradient, represents the vertical gradient, and represents the weight coefficient, and ,when is greater than or equal to the preset characteristic value, it is determined that the video quality of the first output video meets the standard. If the value is less than a preset characteristic value, it is determined that the video quality of the first output video does not meet the standard.

[0089] It can be understood that the video tracing module is used to evaluate the video quality of the first output video. Each frame of the first output video has its corresponding horizontal gradient matrix. The horizontal gradient of the first output video is the time series of the horizontal gradient matrices of multiple frames. Through the edge image feature technology, the horizontal gradient of the first output video can be extracted. The vertical gradient extraction is similar. By extracting the horizontal gradient and the vertical gradient to calculate the edge feature value, the video quality evaluation of the first output video is realized. The preset feature value is 30. When the edge feature value is less than 30, it is considered that the video quality of the first output video does not meet the standard. Otherwise, it is judged to meet the standard. The integrity and clarity of the first output video image are quantified by the edge image feature technology, providing an intuitive quality evaluation standard without human intervention, reducing subjective errors caused by human factors, and flexibly adjusting the weight coefficient and the preset threshold. It can adapt to the quality requirements of different input videos for the output video, and enhance the applicability and flexibility of the system.

[0090] In some embodiments of the present application, if the target is not met, the initial deep learning model is optimized by a gradient descent algorithm. When the target deep learning model is obtained, the video backtracking module substitutes the test set into the neural network model to obtain the predicted value of the video adjustment parameter, and determines the cross entropy loss according to the combined value of the predicted value and the video adjustment parameter. The cross entropy loss is obtained by the following formula:

[0091]

[0092]

[0093] in, Represents the stapled value, represents the i-th predicted value, represents the cross entropy loss, Indicates the rounded bit rate, Indicates the rounding quantization level, Indicates the integer frame rate, and the target deep learning model is obtained by controlling the learning of the neural network model according to the cross entropy loss.

[0094] Specifically, the video quality of the first output video does not meet the standard, and the video backtracking module inputs the test set into the neural network model to generate the predicted value of the video adjustment parameter. The test set is a sample in the standard data. The sample is processed by the neural network model to obtain the predicted value of the video adjustment parameter of each input video. In order to accurately evaluate the performance of the model, the predicted value needs to be compared with the bound value, which integrates each parameter in the video adjustment parameter. By comparing the predicted value with the bound value, the cross entropy loss is calculated. The cross entropy loss is a loss function in deep learning. The cross entropy loss reflects the gap between the predicted value and the bound value. After calculating the cross entropy loss, the system uses the gradient descent algorithm to optimize the neural network model. Gradient descent is an optimization method that iteratively updates model parameters to reduce cross entropy loss. In each iteration, the gradient descent algorithm calculates the gradient of the loss function relative to the neural network model parameters. The gradient descent algorithm adjusts the parameters of the neural network model through the following formula.

[0095] ;

[0096] in, Represents the current model parameters of the neural network model, represents the learning rate, Represents the gradient of the loss function with respect to the current model parameters.

[0097] It is understandable that through multiple iterations, the parameters of the current neural network model will be gradually optimized. When the accuracy of the predicted video adjustment parameters is greater than or equal to the preset accuracy of 80%, it is considered that the target deep learning model is obtained. The first output video is used to obtain the first video feature data, and its processing method is consistent with the processing method of the input video, which will not be repeated here. The obtained first video feature data is substituted into the target deep learning model to obtain the first output video adjustment parameters. The first output video adjustment parameters further adjust the bit rate, quantization level and frame rate of the first output video to obtain the second output video. The adjustment process is consistent with the process of obtaining the first output video from the input video, which will not be repeated here. The second output video obtained by processing the first output video further improves the video quality, meets the video quality requirements of different types of input videos, and improves the system's optimization processing capabilities and adaptability to the input video.

[0098] In some embodiments of the present application, the video evaluation module is configured to demarcate the second output video into a grayscale frame area, extract the video frame of the grayscale frame area, determine the grayscale image brightness value of each pixel according to the video frame, determine the contrast according to all the grayscale image brightness values, and judge whether the video quality of the second output video meets the standard according to the contrast, including: the video evaluation module converts the second output video into a grayscale image, divides a 2x2 matrix area on the grayscale image, the matrix area is set as the grayscale frame area, counts each pixel of the video frame and records the number of all pixels, determines the contrast according to the variance of the grayscale image brightness values ​​of all pixels, sets a contrast threshold, when the contrast is greater than the contrast threshold, it is determined that the video quality of the second output video is qualified, and when the contrast is less than or equal to the contrast threshold, it is determined that the video quality of the second output video is unqualified.

[0099] Specifically, the video evaluation module converts the second output video into a grayscale image. The grayscale image converts the color information of each pixel of the second output video frame into a brightness value. By converting it into a grayscale image, color interference can be removed and the focus can be placed on brightness changes. A 2x2 matrix area is divided on the grayscale image, and the area is set as a grayscale frame area. The matrix area contains a certain number of pixels, and the grayscale image brightness value of each pixel in the matrix area is counted. The grayscale image brightness value reflects the brightness distribution. A higher grayscale image brightness value indicates that the video image has better contrast and clarity. The contrast is calculated based on the variance of the grayscale image brightness values ​​of all pixels. A lower contrast indicates that the video image has a certain degree of blur or distortion. A contrast threshold is set. When the calculated contrast is greater than the set contrast threshold, the system determines that the video quality of the second output video is qualified. If the contrast is less than or equal to the contrast threshold, the video quality of the second output video is determined to be unqualified.

[0100] It is understandable that the contrast threshold is an artificially set threshold for image detection, and the contrast threshold can be adjusted appropriately. By extracting the grayscale image and calculating the contrast, the consistency and objectivity of the evaluation are guaranteed, and the video quality effect of the second output video is ensured.

[0101] In some embodiments of the present application, if the standard is not met, a third output video is obtained based on historical input video adjustment parameters, including: a video evaluation module counts the video adjustment parameters of each input video and establishes a video adjustment parameter set, when the video quality of the second output video is unqualified, the adjustment parameter mean of the historical input video in the video adjustment parameter set is determined as the second output video adjustment parameter, and the third output video is obtained based on the second output video adjustment parameter.

[0102] Specifically, the video evaluation module will count the video adjustment parameters of each input video and summarize these parameters to form a video adjustment parameter set. Different types of input videos will obtain different video adjustment parameters. The video evaluation module records and stores the video adjustment parameters to provide a reference basis for the subsequent video quality evaluation and optimization of the second output video. When the second output video is judged to have unqualified video quality, the video evaluation module will extract the adjustment parameters of all historical input videos from the video adjustment parameter set, calculate the mean of the adjustment parameters of the historical input videos and set them as the adjustment parameters of the second output video. The mean reflects the average adjustment level of multiple input videos under specific conditions, so it can avoid the deviation of the adjustment parameters of a single output video and provide a stable and reliable reference value for adjusting the second output video. The system adjusts the parameters of the second output video, including the bit rate, quantization level and frame rate, according to the mean of the adjustment parameters of the historical input videos to generate a third output video with good video quality.

[0103] It can be understood that by introducing the mean of adjustment parameters of historical input videos, the system can accurately optimize the video quality of the second output video, meet the requirements of different types of input videos for different video qualities, avoid blindly adjusting the video quality of the second output video, and improve the scientificity and rigor of system optimization, so that the system can maintain good adaptability while meeting the requirements of optimizing videos.

[0104] In summary, the beneficial effects of the present invention are: by analyzing the motion complexity, scene switching and texture details of the input video to generate corresponding feature vectors, and decoding and dynamic adjustment based on the deep learning model, adaptive optimization of different types of input videos is achieved. Whether in complex motion states, fast scene switching and detailed texture content, the system can ensure that the video quality is improved by dynamically adjusting the encoding, and optimize the deep learning model through the gradient descent algorithm to improve the encoding accuracy, so as to adapt to the needs of different network environments and playback devices, ensure the smooth playback of the first output video, the second output video and the third output video and reduce the impact of network bandwidth limitations on the video quality, and evaluate through the brightness value and contrast of the grayscale image to ensure that the video quality of the first output video, the second output video and the third output video meets the expected standard, thereby providing efficient video quality.

[0105] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems or computer program products. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0106] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0107] These computer program instructions may also be stored in a computer-readable storage device that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable storage device produce a product including an instruction device that implements the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0108] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0109] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the relevant field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.

Claims

1. A video content adaptive optimization system based on deep learning, characterized in that: include: Video extraction module, video decoding module, encoding adjustment module, video backtracking module and video evaluation module; The video extraction module is configured to use motion vector analysis to obtain a motion complexity feature vector from the input video, use a scene detection algorithm to identify scene switching of the input video to obtain a scene switching feature vector, use texture detail analysis to obtain a texture feature vector, and obtain video feature data based on the motion complexity feature vector, the scene switching feature vector and the texture feature vector; The video decoding module is configured to establish an initial deep learning model, and decode the plurality of video feature data based on the initial deep learning model to obtain video adjustment parameters; The encoding adjustment module is configured to adjust the video adjustment parameter according to the network status of the input video, the scene of the input video and the frame rate of the input video to obtain a first output video; The video tracing module is configured to extract edge image features of the first output video, calculate edge feature values ​​of the first output video, and determine whether the video quality of the first output video meets the standard according to the edge feature values; If the target is not met, the initial deep learning model is optimized by a gradient descent algorithm to obtain a target deep learning model, the first output video is processed to obtain first video feature data, and the first video feature data is substituted into the target deep learning model to obtain a first output video adjustment parameter, and a second output video is obtained based on the first output video adjustment parameter; The video evaluation module is configured to delimit the second output video into a grayscale frame area, extract the video frame of the grayscale frame area, determine the grayscale image brightness value of each pixel based on the video frame, determine the contrast based on all the grayscale image brightness values, and determine whether the video quality of the second output video meets the standards based on the contrast. If not, adjust the parameters based on the historical input video to obtain the third output video.

2. The video content adaptive optimization system based on deep learning according to claim 1, characterized in that: When the motion complexity feature vector is obtained by analyzing the input video using a motion vector, the scene switching of the input video is identified using a scene detection algorithm to obtain a scene switching feature vector, and the texture feature vector is obtained by analyzing the texture details, the method includes: The video extraction module extracts video pixels of the input video, and calculates the motion complexity feature vector according to the video pixels. The motion complexity feature vector is obtained by the following formula: ; in, represents the motion complexity feature vector, Indicates the total number of video pixels. represents the optical flow vector of the i-th video pixel, μ represents the mean of the optical flow vector; The scene detection algorithm obtains a binary variable feature value by comparing the scene differences between different frames of the input video, and if there is a scene difference, the binary variable feature value is equal to 1, and if there is no scene difference, the binary variable feature value is equal to 0, thereby obtaining a scene switching feature vector; The texture detail analysis captures the local texture features of the input video by generating binary values ​​from the grayscale values ​​of the video pixels, and obtains a texture feature vector.

3. The video content adaptive optimization system based on deep learning according to claim 2, characterized in that: When video feature data is derived according to the motion complexity feature vector, the scene switching feature vector and the texture feature vector, the method includes: The video feature data is obtained by the following formula: ; in, Represents video feature data, represents the motion complexity feature vector, represents the scene switching feature vector, represents the texture feature vector, represents the weight of the motion complexity feature vector, represents the weight of the scene switching feature vector, represents the weight of the texture feature vector, and .

4. The video content adaptive optimization system based on deep learning according to claim 3 is characterized in that: The video decoding module is configured to establish an initial deep learning model, and when decoding the plurality of video feature data based on the initial deep learning model to obtain the video adjustment parameter, it includes: The video decoding module collects a plurality of the video feature data and performs standard processing to obtain standard data, and the standard data is obtained by the following formula: ; in, Indicates standard data, Represents video feature data, Represents the mean of multiple video feature data, Represents the standard deviation of multiple video feature data; The standard data is divided into a training set and a test set. The initial deep learning model adopts a neural network model. Cross-validation combined with grid search is used to find model parameters of the neural network model, and a neural network model is established. The training set is used to fit the neural network model. The test set is substituted into the neural network model and the accuracy of the video adjustment parameters is calculated. When the accuracy of the video adjustment parameters reaches a preset accuracy threshold, the video decoding module decodes the standard data based on the neural network model to obtain the video adjustment parameters, and the video adjustment parameters include a corrected bit rate, a corrected quantization level, and a corrected frame rate.

5. The video content adaptive optimization system based on deep learning according to claim 4 is characterized in that: The video decoding module decodes the standard data based on the neural network model to obtain the video adjustment parameters, and when the video adjustment parameters include a modified bit rate, a modified quantization level, and a modified frame rate, the video decoding module comprises: When the video decoding module obtains the modified bit rate, the modified quantization level and the modified frame rate by decoding the neural network model, the neural network model outputs continuous values ​​of bit rate, continuous values ​​of quantization level and continuous values ​​of frame rate, and the continuous values ​​of bit rate, quantization level and frame rate are rounded to obtain the modified bit rate, the modified quantization level and the modified frame rate, and the rounding calculation is obtained by the following formula: ; ; ; in, Indicates the rounded bit rate, Indicates the continuous value of bit rate, Indicates the rounding quantization level, Represents a continuous value of the quantization level, Indicates the frame rate. represents a continuous value of the frame rate, the rounded bit rate is the modified bit rate, the rounded quantization level is the modified quantization level, and the rounded frame rate is the modified frame rate.

6. The video content adaptive optimization system based on deep learning according to claim 5, characterized in that: When the encoding adjustment module is configured to adjust the video adjustment parameter according to the network status of the input video, the scene of the input video and the frame rate of the input video to obtain the first output video, it includes: The encoding adjustment module determines whether to adjust the modified bit rate according to the network status of the input video to obtain a first report, determines whether to adjust the modified quantization level according to the scene of the input video to obtain a second report, determines whether to adjust the modified frame rate according to the frame rate of the input video to obtain a third report, and outputs the first report, the second report and the third report as a first output video.

7. The video content adaptive optimization system based on deep learning according to claim 6, characterized in that: The video tracing module is configured to extract edge image features of the first output video, calculate edge feature values ​​of the first output video, and determine whether the video quality of the first output video meets the standard according to the edge feature values, including: The video tracing module extracts the horizontal gradient and the vertical gradient of the first output video, and the edge feature value is obtained by the following formula: in, represents the edge eigenvalue, represents the horizontal gradient, represents the vertical gradient, and represents the weight coefficient, and ; when is greater than or equal to a preset characteristic value, it is determined that the video quality of the first output video meets the standard; when If the value is less than a preset characteristic value, it is determined that the video quality of the first output video does not meet the standard.

8. The video content adaptive optimization system based on deep learning according to claim 7, characterized in that: If the target is not met, the initial deep learning model is optimized using a gradient descent algorithm to obtain a target deep learning model, including: The video tracing module substitutes the test set into the neural network model to obtain a predicted value of the video adjustment parameter, and determines a cross entropy loss according to the predicted value and the combined value of the video adjustment parameter. The cross entropy loss is obtained by the following formula: in, Represents the stapled value, represents the i-th predicted value, represents the cross entropy loss, represents the rounded bit rate, represents the rounding quantization level, represents the rounded frame rate, represents the number of standard data in the test set; The target deep learning model is obtained by controlling the learning of the neural network model according to the cross entropy loss.

9. The video content adaptive optimization system based on deep learning according to claim 8, characterized in that: The video evaluation module is configured to delimit the second output video into a grayscale frame area, extract a video frame of the grayscale frame area, determine a grayscale image brightness value of each pixel according to the video frame, determine a contrast according to all grayscale image brightness values, and determine whether the video quality of the second output video meets the standard according to the contrast, including: The video evaluation module converts the second output video into a grayscale image, divides a 2x2 matrix area on the grayscale image, sets the matrix area as a grayscale frame area, counts each pixel of the video frame and records the number of all pixels, determines the contrast according to the variance of the grayscale image brightness values ​​of all pixels, and sets a contrast threshold; When the contrast is greater than the contrast threshold, it is determined that the video quality of the second output video is qualified; When the contrast is less than or equal to the contrast threshold, it is determined that the video quality of the second output video is unqualified.

10. The video content adaptive optimization system based on deep learning according to claim 9, characterized in that: If the standard is not met, the parameters are adjusted according to the historical input video to obtain the third output video, including: The video evaluation module counts the video adjustment parameters of each input video and establishes a video adjustment parameter set. When the video quality of the second output video is unqualified, the second output video adjustment parameters are determined according to the average of the adjustment parameters of the historical input videos in the video adjustment parameter set, and the third output video is obtained according to the second output video adjustment parameters.

Citation Information

Patent Citations

  • Ultra-high-definition video compression damage grade evaluation method based on deep learning network

    CN116524387A

  • Low-delay video transmission method and system

    CN118945364A