Video Optimization Processing Method, System, Device and Storage Medium
By constructing a video quality scoring model and a video scene classification model, and automatically selecting a suitable super-scoring model for video optimization processing, the problem of ineffective identification of low-definition videos and unstable processing effects in the existing technology is solved, and a more efficient and stable video optimization effect is achieved.
Patent Information
- Application Number
- CN202410449985.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-15
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2044-04-15
AI Technical Summary
The prior art cannot effectively and accurately identify low-definition videos in video optimization processing, and due to the diverse video content scenes, the performance is unstable when processing using a single super-resolution model and its practicality is average.
By constructing a video quality scoring model, identify videos with high resolution but low actual image quality, and identify the scene classification of the video through the video scene classification model, and automatically select the appropriate target super-score model from the super-score model database for optimization processing.
It improves the performance of video optimization processing, improves the user's video viewing experience, and enhances the stability and practicality of processing effects.
Smart Images

Figure CN118433446B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of video processing, and in particular to a video optimization processing method, system, device, and storage medium. Background Art
[0002] With the rapid development of Internet technology, video has become the main medium for people to obtain and share information. The full promotion and popularization of 5G technology have made network bandwidth no longer a bottleneck, and users can smoothly play high-definition videos on their devices. Correspondingly, users' requirements for video quality are also getting higher and higher, and how to provide users with high-quality and clear video services has become a hot topic in the technical research of the video industry.
[0003] In related technologies, there are applications for optimizing video processing. For example, the resolution of the video is improved through relevant technical means to improve the clarity of the video. The common implementation method is mainly to obtain the resolution information of the video, and for videos with a resolution less than a preset threshold, super-resolution processing is performed through a general super-resolution model to improve the image quality of low-resolution videos. However, this implementation method cannot effectively and accurately identify low-definition videos; moreover, due to the diverse scenes of video content and different video texture details in different content scenes, there may be a situation where the processing effect of some video content is better while the processing effect of some other video content is worse when using the super-resolution model, resulting in unstable performance and general practicality of video optimization.
[0004] Therefore, the problems existing in the prior art still need to be solved and optimized urgently. Summary of the Invention
[0005] The purpose of this application is to solve at least to some extent one of the technical problems existing in the related technologies.
[0006] To this end, an object of an embodiment of this application is to provide a video optimization processing method, system, device, and storage medium.
[0007] To achieve the above technical purpose, the technical solutions adopted in the embodiments of this application include:
[0008] On the one hand, an embodiment of this application provides a video optimization processing method, and the method includes:
[0009] Obtain a target video segment;
[0010] Perform frame extraction on the target video segment to obtain a set of video frames, and determine the first resolution corresponding to the target video segment according to the set of video frames;
[0011] If the first resolution is greater than or equal to the first preset threshold, input the video frame set into the video picture quality scoring model to obtain the first picture quality score corresponding to the target video segment;
[0012] If the first picture quality score is less than the second preset threshold, input the video frame set into the video scene classification model to obtain the first scene prediction result corresponding to the target video segment;
[0013] Calculate the difference between the first resolution and the predetermined target resolution;
[0014] According to the first scene prediction result and the difference, determine the target super-resolution model from the super-resolution model database, and perform optimization processing on the target video segment based on the target super-resolution model; wherein, the super-resolution model database includes multiple super-resolution models, each super-resolution model includes a scene type parameter and an optimization magnification parameter, the first scene prediction result is used to determine the scene type parameter corresponding to the target super-resolution model, and the difference is used to determine the optimization magnification parameter corresponding to the target super-resolution model.
[0015] In addition, according to a video optimization processing system of the above embodiments of the present application, it may further have the following additional technical features:
[0016] Further, in an embodiment of the present application, the process of performing frame extraction on the target video segment to obtain a video frame set includes:
[0017] Perform frame extraction on the target video segment to obtain the original video frames;
[0018] Crop the image content of a predetermined size from the central area of the original video frame to obtain the first video frame;
[0019] Perform pixel normalization processing on the first video frame to obtain the second video frame;
[0020] Integrate the second video frames to obtain the video frame set.
[0021] Further, in an embodiment of the present application, the method further includes:
[0022] If the first resolution is less than the first preset threshold, input the video frame set into the video scene classification model to obtain the first scene prediction result corresponding to the target video segment;
[0023] Calculate the difference between the first resolution and the predetermined target resolution;
[0024] Based on the first scenario prediction result and the difference value, determine a target super-resolution model from the super-resolution model database, and perform optimization processing on the target video segment based on the target super-resolution model; wherein, the super-resolution model database includes multiple super-resolution models, each super-resolution model includes a scene type parameter and an optimization magnification parameter, the first scenario prediction result is used to determine the scene type parameter corresponding to the target super-resolution model, and the difference value is used to determine the optimization magnification parameter corresponding to the target super-resolution model.
[0025] Further, in an embodiment of the present application, the video image quality scoring model is trained through the following steps:
[0026] Obtain a sample video segment and a first label value corresponding to the sample video segment; the first label value is used to represent the true result of the image quality score corresponding to the sample video segment;
[0027] Perform frame extraction on the sample video segment to obtain a set of sample frames;
[0028] Input the set of sample frames into the video image quality scoring model to obtain a second image quality score corresponding to the sample video segment;
[0029] Determine a first loss value according to the first label value and the second image quality score;
[0030] Update the parameters of the video image quality scoring model according to the first loss value to obtain a trained video image quality scoring model.
[0031] Further, in an embodiment of the present application, the video scene classification model is trained through the following steps:
[0032] Obtain a sample video segment and a second label value corresponding to the sample video segment; the second label value is used to represent the true result of the scene type corresponding to the sample video segment;
[0033] Perform frame extraction on the sample video segment to obtain a set of sample frames;
[0034] Input the set of sample frames into the video scene classification model to obtain a second scenario prediction result corresponding to the sample video segment;
[0035] Determine a second loss value according to the second label value and the second scenario prediction result;
[0036] Update the parameters of the video scene classification model according to the second loss value to obtain a trained video scene classification model.
[0037] Further, in an embodiment of the present application, determining the second loss value according to the second tag value and the second scenario prediction result includes:
[0038] Determining the second loss value through a cross-entropy loss function according to the second tag value and the second scenario prediction result.
[0039] Further, in an embodiment of the present application, the super-resolution model includes at least one of SRCNN, ESPCN, VDSR, and SRGAN.
[0040] On the other hand, an embodiment of the present application provides a video optimization processing system, and the system includes:
[0041] An acquisition unit, configured to acquire a target video segment;
[0042] A frame extraction unit, configured to perform frame extraction on the target video segment to obtain a video frame set, and determine the first resolution corresponding to the target video segment according to the video frame set;
[0043] A first prediction unit, configured to input the video frame set into a video image quality scoring model if the first resolution is greater than or equal to a first preset threshold, to obtain a first image quality score corresponding to the target video segment;
[0044] A second prediction unit, configured to input the video frame set into a video scene classification model if the first image quality score is less than a second preset threshold, to obtain a first scenario prediction result corresponding to the target video segment;
[0045] A calculation unit, configured to calculate the difference between the first resolution and a predetermined target resolution;
[0046] A processing unit, configured to determine a target super-resolution model from a super-resolution model database according to the first scenario prediction result and the difference, and perform optimization processing on the target video segment based on the target super-resolution model; wherein, the super-resolution model database includes multiple super-resolution models, each super-resolution model includes a scene type parameter and an optimization magnification parameter, the first scenario prediction result is used to determine the scene type parameter corresponding to the target super-resolution model, and the difference is used to determine the optimization magnification parameter corresponding to the target super-resolution model.
[0047] On the other hand, an embodiment of the present application provides a computer device, including:
[0048] At least one processor;
[0049] At least one memory, configured to store at least one program;
[0050] When the at least one program is executed by the at least one processor, the at least one processor implements the above-mentioned video optimization processing method.
[0051] On the other hand, an embodiment of the present application also provides a computer-readable storage medium, in which a program executable by a processor is stored, and the program executable by the processor is used to implement the above-mentioned video optimization processing method when executed by the processor.
[0052] The advantages and beneficial effects of the present application will be partially given in the following description, partially will become obvious from the following description, or will be learned through the practice of the present application:
[0053] A video optimization processing method disclosed in an embodiment of the present application includes: obtaining a target video segment; performing frame extraction on the target video segment to obtain a video frame set, and determining a first resolution corresponding to the target video segment according to the video frame set; if the first resolution is greater than or equal to a first preset threshold, inputting the video frame set into a video image quality scoring model to obtain a first image quality score corresponding to the target video segment; if the first image quality score is less than a second preset threshold, inputting the video frame set into a video scene classification model to obtain a first scene prediction result corresponding to the target video segment; calculating a difference between the first resolution and a predetermined target resolution; determining a target super-resolution model from a super-resolution model database according to the first scene prediction result and the difference, and performing optimization processing on the target video segment based on the target super-resolution model. This method discriminates low-quality videos with high resolution but low actual image quality by constructing a video image quality scoring model, and then identifies the scene classification of the video to be processed by constructing a video scene classification model, and automatically selects a suitable target super-resolution model from the super-resolution model database to perform super-resolution processing on the video to be processed, so as to improve the performance of video optimization processing and improve the user's video viewing experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following introduces the relevant technical solution drawings in the embodiments of the present application or the prior art. It should be understood that the drawings introduced below are only for conveniently and clearly presenting some embodiments of the technical solutions in the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.
[0055] Figure 1 It is a schematic diagram of the implementation environment of a video optimization processing method provided in an embodiment of the present application;
[0056] Figure 2 It is a schematic flowchart of a video optimization processing method provided in an embodiment of the present application;
[0057] Figure 3 This is a schematic structural diagram of a video optimization processing system provided in an embodiment of the present application;
[0058] Figure 4 This is a schematic structural diagram of a computer device provided in an embodiment of the present application. Detailed implementation manners
[0059] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the embodiments of the present application. They are only examples of devices and methods consistent with some aspects of the embodiments of the present application detailed in the appended claims.
[0060] It can be understood that the terms "first", "second", etc. used in the present application can be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present application, the first information can also be referred to as the second information, and similarly, the second information can also be referred to as the first information. Depending on the context, the words "if", "when" as used herein can be interpreted as "when...", "while...", or "in response to determining".
[0061] The terms "at least one", "multiple", "each", "any one", etc. used in the present application, at least one includes one, two or more than two, multiple includes two or more than two, each refers to each of the corresponding multiple, and any one refers to any one of the multiple.
[0062] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0063] First, several nouns involved in the present application are analyzed:
[0064] MOS (Mean Opinion Score): The mean opinion score is a standard used to evaluate the quality of audio or video. The MOS score is based on expert evaluations and reflects the opinions of viewers or listeners on the quality of audio or video under certain conditions. The MOS score usually ranges from 1 to 5, where 5 represents the best quality and 1 represents the worst quality.
[0065] Artificial Intelligence (AI): It is a new technical science that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence; artificial intelligence is a branch of computer science. Artificial intelligence attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. The research in this field includes robots, speech recognition, image recognition, natural language processing, and expert systems, etc. Artificial intelligence can simulate the information process of human consciousness and thinking. Artificial intelligence also uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results in terms of theories, methods, technologies, and application systems.
[0066] Machine Learning (ML): It is an interdisciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications cover all fields of artificial intelligence. Machine learning usually includes technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.
[0067] Deep Learning (DL): It is a new research direction in the field of machine learning. It is introduced into machine learning to make it closer to the original goal - artificial intelligence. Deep learning is to learn the internal laws and representation levels of sample data, and the information obtained during these learning processes is very helpful for the interpretation of data such as text, images, and sounds. Its ultimate goal is to enable machines to have the ability of analytical learning like humans and be able to recognize data such as text, images, and sounds.
[0068] With the rapid development of Internet technology, video has become the main media for people to obtain and share information. The full promotion and popularization of 5G technology have made network bandwidth no longer a bottleneck, and users can smoothly play high-definition videos on their devices. Correspondingly, users' requirements for video quality are getting higher and higher, and how to provide users with high-quality and clear video services has become a hot topic in the technical research of the video industry.
[0069] In the related art, there are applications for optimizing video processing. For example, the resolution of the video is improved by related technical means to improve the clarity of the video. The common implementation method is mainly to obtain the resolution information of the video. For videos with a resolution less than a preset threshold, super-resolution processing is performed through a general super-resolution model to improve the picture quality of low-resolution videos. However, this implementation method cannot effectively and accurately identify low-definition videos. Moreover, due to the variety of video content scenes, the video texture details of different content scenes are different. For example, video scenes of the animation category have obvious texture features such as lines and color blocks, video scenes of the portrait category have texture features such as facial features and hair, and video scenes of the landscape category have texture features such as leaves, buildings, and lights. There are significant differences in the texture features of different scenes, and the texture restoration targets are different. When using a single super-resolution model for processing, there may be a situation where the processing effect of some video content is better, while the processing effect of some other video content is worse, resulting in unstable performance of video optimization and general practicality.
[0070] In view of this, in the embodiments of the present application, a video optimization processing method is provided. This method constructs a video picture quality scoring model to identify inferior videos with a relatively high resolution but actually low picture quality, and then constructs a video scene classification model to identify the scene classification to which the video to be processed belongs, and automatically selects a suitable target super-resolution model from the super-resolution model database to perform super-resolution processing on the video to be processed, so as to improve the performance of video optimization processing and improve the user's video viewing experience.
[0071] Next, first introduce the implementation environment involved in the video optimization processing method provided in the embodiments of the present application. Refer to Figure 1 , Figure 1 FIG. shows a schematic diagram of the implementation environment of a video optimization processing method. The main software and hardware of this implementation environment mainly include a terminal device 110 and a server 120, and the terminal device 110 is communicatively connected to the server 120. Among them, this video optimization processing method can be configured on the side of the terminal device 110 or on the side of the server 120. For example, when this video optimization processing method is configured on the side of the terminal device 110, the application of video optimization processing can be realized independently relying on the terminal device 110; when this video optimization processing method is configured on the side of the server 120, the application of video optimization processing can be realized through the interaction between the terminal device 110 and the server 120.
[0072] Specifically, the terminal device 110 in the present application may include, but is not limited to, any one or more of a smart watch, a smart phone, a computer, a personal digital assistant (PDA), a smart voice interaction device, a smart home appliance, or a vehicle-mounted terminal. The server 120 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It may also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. A communication connection may be established between the terminal device 110 and the server 120 through a wireless network or a wired network. The wireless network or wired network uses standard communication technologies and / or protocols. The network may be set to the Internet or any other network, such as any combination including, but not limited to, a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile, wired or wireless network, a private network, or a virtual private network.
[0073] Of course, it can be understood that Figure 1 the implementation environment in Figure 1 is only an optional application scenario of the video optimization processing method provided in the embodiments of the present application. The actual application is not fixed to the Figure 1 software and hardware environment shown. The video optimization processing method provided in the embodiments of the present application will be described in detail below with reference to the
[0074] Please refer to Figure 2 Figure 2 which is a schematic flowchart of a video optimization processing method provided in the embodiments of the present application. Referring to Figure 2 the video optimization processing method provided in the present application includes, but is not limited to:
[0075] Step 210, obtain a target video segment;
[0076] In this step, when optimizing the video, the target video segment to be optimized can be obtained. In the embodiments of the present application, the target video segment can come from various different data sources, such as real-time videos captured by a camera, video files stored on a device or server, videos downloaded from the Internet, and so on. Moreover, the present application places no restrictions on the length, quality, or content of the target video segment, and a suitable target video segment can be selected for recognition according to specific requirements and purposes. For example, in some embodiments, a relatively long video segment can be obtained, and the embodiments of the present application can appropriately crop it and use the obtained partial segment as the target video segment.
[0077] Step 220: Perform frame extraction on the target video segment to obtain a set of video frames, and determine the first resolution corresponding to the target video segment according to the set of video frames;
[0078] In this step, after obtaining the target video segment, frame extraction can be performed on it, and these obtained video frames are recorded as the set of video frames. In the embodiments of the present application, when performing frame extraction, it can be equidistant extraction. For example, frame extraction can be performed at a frequency of every 5 frames, or it can be non-equidistant extraction. For example, a sequence of key frames can be extracted to form the set of video frames. The present application places no restrictions on the extraction frequency. In some embodiments, it can be determined according to the total number of frames of the target video segment. If the total number of frames of the target video segment is small, the extraction frequency can be appropriately increased; conversely, the extraction frequency can be appropriately decreased. In this way, it can not only ensure that there are enough video frames to form the set of video frames but also reduce the data processing volume as much as possible.
[0079] In this step, after obtaining the set of video frames, the resolution of the target video segment can be determined according to the set of video frames, denoted as the first resolution. Specifically, in the embodiments of the present application, the resolution of each video frame in the set of video frames can be calculated and then averaged to obtain the first resolution.
[0080] Step 230: If the first resolution is greater than or equal to the first preset threshold, input the set of video frames into the video image quality scoring model to obtain the first image quality score corresponding to the target video segment;
[0081] In the embodiments of the present application, a threshold value of the resolution can be preset, denoted as the first preset threshold. If the resolution corresponding to a certain target video segment is greater than or equal to the first preset threshold, then it can be considered that from the perspective of resolution, this target video segment is very likely to be a video segment with good image quality; relatively, if the resolution corresponding to a certain target video segment is less than the first preset threshold, it can be considered that the image quality of this target video segment is poor and needs to be optimized. It can be understood that the size of the first preset threshold in the embodiments of the present application can be flexibly set according to needs, and the present application does not limit its specific value.
[0082] In this step, after obtaining the first resolution, it can be compared with the first preset threshold. If the first resolution is greater than or equal to the first preset threshold, it means that from the perspective of resolution only, the image quality of this target video segment is good and does not need to be optimized. However, as mentioned in the background part, the resolution index is only a parameter of one dimension of the video image quality. Some low-definition videos can be forcibly set to a high resolution during the encoding stage, which may bypass the optimization process and cause low-quality videos to flow out, affecting the user's viewing experience. Therefore, in the embodiments of the present application, when the first resolution is greater than or equal to the first preset threshold, the target video segment is not directly recognized as a video segment with good image quality, but its image quality is further detected. Specifically, in this step, the video frame set corresponding to the target video segment can be input into the video image quality scoring model, and the video image quality scoring model is used to detect the image quality situation of the target video segment to obtain a scoring result, denoted as the first image quality score. It can be understood that in the embodiments of the present application, there is no limit to the interval of the scoring data output by the video image quality scoring model. For example, in some embodiments, the image quality scoring result output by the video image quality scoring model can be a positive integer. For example, it can output an integer from 1 to 5, and the larger the value, the higher the image quality of the input video segment considered by the model; in other embodiments, the image quality scoring result output by the video image quality scoring model can be any positive value, such as the range can be between 1 and 100. Similarly, the larger the value, the higher the image quality of the input video segment considered by the model. The present application does not limit this.
[0083] Step 240: If the first image quality score is less than the second preset threshold, input the video frame set into the video scene classification model to obtain the first scene prediction result corresponding to the target video segment;
[0084] In an embodiment of the present application, a threshold for image quality scoring may also be preset in advance, denoted as the second preset threshold. If the image quality score corresponding to a certain target video segment is greater than or equal to the second preset threshold, it can be considered that the image quality of the target video segment is good and no optimization processing is required. On the contrary, if the image quality score corresponding to a certain target video segment is less than the second preset threshold, it can be considered that the image quality of the target video segment is poor and optimization processing is required. Similarly, it can be understood that the size of the second preset threshold in the embodiment of the present application can be flexibly set according to needs, and the present application does not limit its specific value.
[0085] In an embodiment of the present application, after obtaining the first image quality score corresponding to the target video segment, it can be compared with the second preset threshold. If it is greater than or equal to the second preset threshold, it can be considered that the target video segment does not need to be optimized, and the subsequent content can be processed by skipping the current target video segment. If the first image quality score is less than the second preset threshold, it can be considered that the current target video segment needs to be optimized.
[0086] In this step, when the first image quality score is less than the second preset threshold, the target video segment is optimized. First, the video frame set is input into the trained video scene classification model. In the embodiment of the present application, the video scene classification model is used to predict the scene category corresponding to the content in the video segment. The scene category may include portrait, scenery, animation, etc., and the present application does not limit this. By predicting through the video scene classification model, the scene prediction result corresponding to the target video segment can be obtained, which is denoted as the first scene prediction result in the embodiment of the present application.
[0087] Step 250: Calculate the difference between the first resolution and the predetermined target resolution;
[0088] In this step, when it is determined that the target video segment needs to be optimized, the difference between the first resolution and the predetermined target resolution can also be calculated. Here, the target resolution refers to the resolution corresponding to the set optimization target, and its value can be relatively large to facilitate determining how much each target video segment differs from the optimization target in terms of resolution.
[0089] Step 260: Determine a target super-resolution model from the super-resolution model database according to the first scene prediction result and the difference, and optimize the target video segment based on the target super-resolution model; wherein, the super-resolution model database includes multiple super-resolution models, each super-resolution model includes a scene type parameter and an optimization magnification parameter, the first scene prediction result is used to determine the scene type parameter corresponding to the target super-resolution model, and the difference is used to determine the optimization magnification parameter corresponding to the target super-resolution model.
[0090] In this step, based on the first scene prediction result and the obtained resolution difference, a super-resolution model used for this optimization process can be determined from a pre-established super-resolution model database, denoted as the target super-resolution model. Specifically, in the embodiments of the present application, the super-resolution model database may include multiple super-resolution models. A super-resolution model is a technology used to enhance low-resolution images or videos to high-resolution. Common super-resolution models may include SRCNN, ESPCN, VDSR, SRGAN, etc. These models utilize deep learning technology to learn the mapping relationship between low-resolution images and high-resolution images, thereby achieving the optimization process of images or videos. In the field of video processing, these super-resolution models can be applied to improve the clarity and quality of videos and enhance the viewing experience of users. In the embodiments of the present application, for each super-resolution model, it may include a scene type parameter and an optimization magnification parameter. Among them, the scene type parameter is used to characterize which scene type of video data the super-resolution model is applicable to process, and the optimization magnification parameter is used to characterize the multiple of the resolution enhancement of the super-resolution model. It can be understood that a super-resolution model can be trained on a predetermined data set, and through these data sets, its corresponding scene type parameter and optimization magnification parameter can be determined. For example, if a certain super-resolution model is trained on a portrait data set and the training goal is to double the resolution, then it can be determined that its corresponding scene type parameter is portrait and the optimization magnification parameter is 1. The present application does not limit the scene type parameters and optimization magnification parameters corresponding to each super-resolution model.
[0091] It can be understood that in the embodiments of the present application, when determining the target super-resolution model, the corresponding scene type parameter can be determined based on the first scene prediction result. For example, assuming that the first scene prediction result corresponding to the target video segment is a landscape video, then a super-resolution model with a corresponding scene type parameter of landscape can be selected for optimization processing; similarly, the corresponding optimization magnification parameter can also be determined according to the resolution difference. If the difference is large, it indicates that the resolution of the current target video segment is low, and a larger optimization magnification parameter can be determined; relatively, if the difference is small, it indicates that the resolution of the current target video segment is high, and a smaller optimization magnification parameter can be determined, that is, this difference can be positively correlated with the optimization magnification parameter. The present application does not limit the specific functional relationship between the two. After determining the appropriate scene type parameter and optimization magnification parameter, the target super-resolution model can be determined, and thus the target super-resolution model can be used to optimize the target video segment.
[0092] It can be understood that in the video optimization processing method provided in the embodiments of the present application, a video quality scoring model is constructed to identify low-quality videos with high resolution but low actual picture quality. Then, a video scene classification model is constructed to identify the scene classification to which the video to be processed belongs, and an appropriate target super-resolution model is automatically selected from the super-resolution model database to perform super-resolution processing on the video to be processed, so as to improve the performance of video optimization processing and improve the user's video viewing experience.
[0093] In some embodiments, the frame extraction process for the target video segment to obtain a video frame set includes:
[0094] Perform frame extraction on the target video segment to obtain original video frames;
[0095] Crop an image content of a predetermined size from the central region of the original video frame to obtain a first video frame;
[0096] Perform pixel normalization processing on the first video frame to obtain a second video frame;
[0097] Integrate the second video frames to obtain the video frame set.
[0098] In the embodiments of the present application, when obtaining a video frame set by performing frame extraction on a target video segment, the target video segment can be first subjected to frame extraction, and the obtained video frames are recorded as original video frames. Then, an image content of a predetermined size can be cropped from the central region of the original video frame. For example, an image in a 224x224 region can be cut out, and the obtained video frame is recorded as the first video frame. Next, pixel normalization processing can be performed on the first video frame. For example, the pixel value range of the first video frame can be normalized from 0 to 255 to between 0 and 1.0, and the obtained video frame is recorded as the second video frame. Integrating the obtained second video frames can obtain the video frame set corresponding to the target video segment.
[0099] In some embodiments, the method further includes:
[0100] If the first resolution is less than the first preset threshold, input the video frame set into the video scene classification model to obtain a first scene prediction result corresponding to the target video segment;
[0101] Calculate the difference between the first resolution and a predetermined target resolution;
[0102] Determine a target super-resolution model from a super-resolution model database according to the first scenario prediction result and the difference value, and perform optimization processing on the target video segment based on the target super-resolution model; wherein, the super-resolution model database includes multiple super-resolution models, each super-resolution model includes a scenario type parameter and an optimization magnification parameter, the first scenario prediction result is used to determine the scenario type parameter corresponding to the target super-resolution model, and the difference value is used to determine the optimization magnification parameter corresponding to the target super-resolution model.
[0103] In the embodiments of the present application, when comparing the first resolution with the first preset threshold, if the first resolution is less than the first preset threshold, the target video segment can also be optimized. The specific processing flow is similar to that of the foregoing embodiments and will not be elaborated herein.
[0104] It should be noted that in the embodiments of the present application, through the two-stage detection of the resolution and the video image quality scoring model, the target video segment with poor image quality can be effectively determined. And from a practical perspective, the efficiency of detecting the resolution is faster than that of using the model to predict the image quality, and the data processing volume is lower. Therefore, in the embodiments of the present application, a large number of low-definition videos with poor image quality can be screened out first through the first preset threshold corresponding to the resolution, which can improve the efficiency of video optimization processing. Then, further screening through the video image quality scoring model can select videos with higher resolution but poor actual image quality, which can improve the practicality of video optimization processing, reduce the probability of low-definition videos flowing out, and is beneficial to improving the user's viewing experience.
[0105] In some embodiments, the video image quality scoring model is trained through the following steps:
[0106] Obtain a sample video segment and a first label value corresponding to the sample video segment; the first label value is used to represent the true result of the image quality score corresponding to the sample video segment;
[0107] Perform frame extraction on the sample video segment to obtain a sample frame set;
[0108] Input the sample frame set into the video image quality scoring model to obtain a second image quality score corresponding to the sample video segment;
[0109] Determine a first loss value according to the first label value and the second image quality score;
[0110] Update the parameters of the video image quality scoring model according to the first loss value to obtain a trained video image quality scoring model.
[0111] In the embodiments of the present application, a no-reference video image quality scoring model can be used to implement image quality scoring prediction. For example, ResNet18 can be used as the backbone network. After the output layer of ResNet18, a Relu layer and an FC layer are stacked, and then the image quality score of the predicted video is output through Sigmoid. In terms of labels, the MOS score obtained by manual scoring can be used.
[0112] Specifically, when training the video image quality scoring model, a batch of sample video segments and the corresponding first label values of the sample video segments can be obtained. Here, the first label values are used to represent the true results of the image quality scores corresponding to the sample video segments. Then, the sample video segments can be frame-extracted to obtain a set of sample frames, and the set of sample frames is input into the video image quality scoring model. The image quality score of the sample video segment is predicted through the video image quality scoring model to obtain a prediction result, which is denoted as the second image quality score in the embodiments of the present application. After obtaining the second image quality score, the prediction accuracy of the video image quality scoring model can be evaluated according to the first label value and the second image quality score, so as to facilitate the iterative update of the parameters of the video image quality scoring model. Specifically, the loss value of the training can be determined according to the first label value and the second image quality score, which is denoted as the first loss value in the embodiments of the present application. After obtaining the first loss value, the video image quality scoring model can be trained by backpropagation to update its internal relevant parameters.
[0113] Specifically, for various models in the field of artificial intelligence, its prediction accuracy can be measured by a loss function. The loss function is defined on a single training data and is used to measure the prediction error of a training data. Specifically, the loss value of the training data is determined by the label of the single training data and the prediction result of the model for the training data. During actual training, a training data set includes many training data, so generally a cost function is used to measure the overall error of the training data set. The cost function is defined on the entire training data set and is used to calculate the average value of the prediction errors of all training data, which can better measure the prediction effect of the model. For a general machine learning model, based on the aforementioned cost function, plus a regularization term that measures the model complexity, it can be used as the training objective function. Based on this objective function, the loss value of the entire training data set can be obtained. There are many common types of loss functions. For example, the 0-1 loss function, the square loss function, the absolute loss function, the logarithmic loss function, the cross-entropy loss function, etc. can all be used as the loss function of the machine learning model, which will not be elaborated one by one here.
[0114] In some embodiments, the video scene classification model is trained through the following steps:
[0115] Obtain a sample video segment and a second label value corresponding to the sample video segment; the second label value is used to characterize the true result of the scene type corresponding to the sample video segment;
[0116] Perform frame extraction on the sample video segment to obtain a set of sample frames;
[0117] Input the set of sample frames into the video scene classification model to obtain a second scene prediction result corresponding to the sample video segment;
[0118] Determine a second loss value according to the second label value and the second scene prediction result;
[0119] Update the parameters of the video scene classification model according to the second loss value to obtain a trained video scene classification model.
[0120] In an embodiment of the present application, similarly, when training a video scene classification model, a batch of sample video segments and second label values corresponding to the sample video segments can also be obtained. Here, the second label value is used to characterize the true result of the scene type corresponding to the sample video segment. Then, frame extraction can be performed on the sample video segment to obtain a set of sample frames. The set of sample frames is input into the video scene classification model, and the scene category of the sample video segment is predicted by the video scene classification model to obtain a prediction result, which is denoted as the second scene prediction result in an embodiment of the present application. After obtaining the second scene prediction result, the prediction accuracy of the video scene classification model can be evaluated according to the second label value and the second scene prediction result, so as to facilitate iterative updating of the parameters of the video scene classification model. Specifically, the loss value of training can be determined according to the second label value and the second scene prediction result, which is denoted as the second loss value in an embodiment of the present application. After obtaining the second loss value, the video scene classification model can be trained by backpropagation to update its internal relevant parameters.
[0121] It should be noted that in an embodiment of the present application, the obtained set of sample frames can be input into the video scene classification model frame by frame to obtain frame classification confidence results. After calculating the frame classification confidence results of all frames in the set, the average value is taken as the classification confidence result of the sample video segment. When judging the classification confidence result, it is detected whether the maximum confidence is greater than the system preset threshold. If it is greater than the threshold, the classification to which the maximum confidence belongs is used as the second scene prediction result corresponding to the sample video segment. If it is less than the threshold, the "general" classification is used as the second scene prediction result corresponding to the sample video segment. When calculating the second loss value, a cross-entropy loss function can be used, and the present application does not limit this.
[0122] In the embodiments of the present application, when training the super-resolution model, a network structure of an appropriate size can be selected from the network structure library according to the applicable classification. For example, SRVGGNet with a smaller number of parameters can be selected for the anime scene, RRDBNet can be selected for the portrait scene, and the number of RRDB blocks can be increased on the basis of RRDBNet for the natural scenery scene to expand the model scale.
[0123] During training, high-definition video data of the corresponding classification can be collected, and videos with high picture quality scores and resolutions not lower than 1080P can be screened out according to the video picture quality evaluation results. Frames can be extracted from each video at a frequency of one frame per second and encoded into WebP files using the WebP lossless compression algorithm to reduce the file size without losing picture quality, so as to solve the file I / O bottleneck problem during multi-card training.
[0124] In the construction of training samples, real-world degradation is simulated by randomly cropping, randomly scaling, randomly compressing high-definition pictures, and adding random Gaussian blur, Gaussian noise, Poisson noise, etc. to generate low-resolution pictures. The training process adopts two-stage training. In the first stage, the MSE loss is used to guide the model to quickly fit and achieve numerical stability. In the second stage of training, the weights obtained from the first stage of training are used for initialization, the charbonnier loss is used instead of the MSE loss, and the perceptual loss based on the pre-trained VGG19 model and the GAN loss are used to measure the overall loss. The two-stage loss function and weight design are as follows:
[0125] L total = 1.0 * Lcharbonnier + 1.0L perceptual + 0.05 * L GAN
[0126] In the formula, L charbonnier is the charbonnier loss, L perceptual is the perceptual loss, and L GAN is the adversarial loss.
[0127] Referring to Figure 3 the video optimization processing system proposed in the embodiments of the present application includes:
[0128] An acquisition unit 310, configured to acquire a target video segment;
[0129] A frame extraction unit 320, configured to perform frame extraction on the target video segment to obtain a video frame set, and determine a first resolution corresponding to the target video segment according to the video frame set;
[0130] A first prediction unit 330, configured to input the video frame set into a video picture quality scoring model if the first resolution is greater than or equal to a first preset threshold, so as to obtain a first picture quality score corresponding to the target video segment;
[0131] A second prediction unit 340, configured to input the video frame set into a video scene classification model if the first picture quality score is less than a second preset threshold, so as to obtain a first scene prediction result corresponding to the target video segment;
[0132] A calculation unit 350, configured to calculate a difference between the first resolution and a predetermined target resolution;
[0133] A processing unit 360, configured to determine a target super-resolution model from a super-resolution model database according to the first scene prediction result and the difference, and perform optimization processing on the target video segment based on the target super-resolution model; wherein, the super-resolution model database includes a plurality of super-resolution models, each super-resolution model includes a scene type parameter and an optimization magnification parameter, the first scene prediction result is used to determine the scene type parameter corresponding to the target super-resolution model, and the difference is used to determine the optimization magnification parameter corresponding to the target super-resolution model.
[0134] It can be understood that the content in the above method embodiments is applicable to the system embodiments of the present application. The functions specifically implemented by the system embodiments of the present application are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.
[0135] Referring to Figure 4 , an embodiment of the present application provides a computer device, including:
[0136] At least one processor 410;
[0137] At least one memory 420, configured to store at least one program;
[0138] When at least one program is executed by at least one processor 410, at least one processor 410 is caused to implement Figure 2 A video optimization processing method as shown.
[0139] Similarly, the content in the above method embodiments is applicable to the computer device embodiments of the present application. The functions specifically implemented by the computer device embodiments of the present application are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.
[0140] An embodiment of the present application further provides a computer-readable storage medium, in which a program executable by a processor 410 is stored, and the program executable by the processor 410 is used to execute the above video optimization processing method when executed by the processor 410.
[0141] The embodiments of the present application also disclose a computer-readable storage medium, in which there is a program executable by a processor. When the program executable by the processor is executed by the processor, it is used to implement an embodiment of a video optimization processing method as shown in Figure 2 the following.
[0142] It can be understood that the content in the embodiment of a video optimization processing method as shown in Figure 2 is applicable to the embodiment of this computer-readable storage medium. The functions specifically implemented by the embodiment of this computer-readable storage medium are the same as those in the embodiment of a video optimization processing method as shown in Figure 2 the following, and the beneficial effects achieved are also the same as those in the embodiment of a video optimization processing method as shown in Figure 2 the following.
[0143] In some alternative embodiments, the functions / operations mentioned in the block diagrams may not occur in the order mentioned in the operation diagrams. For example, depending on the functions / operations involved, two consecutive blocks shown may actually be executed substantially simultaneously or the blocks can sometimes be executed in the reverse order. In addition, the embodiments presented and described in the flowcharts of the present application are provided by way of example for the purpose of providing a more comprehensive understanding of the technology. The disclosed method is not limited to the operations and logical flows presented herein. Alternative embodiments are contemplated, in which the order of various operations is changed and the sub-operations described as part of a larger operation are executed independently.
[0144] In addition, although the present application is described in the context of functional modules, it should be understood that, unless otherwise stated to the contrary, one or more of the functions and / or features may be integrated in a single physical system and / or software module, or one or more functions and / or features may be implemented in separate physical systems or software modules. It can also be understood that a detailed discussion of the actual implementation of each module is not necessary for understanding the present application. Rather, considering the attributes, functions, and internal relationships of the various functional modules in the system disclosed herein, the actual implementation of the module will be understood within the ordinary skill of an engineer. Therefore, those skilled in the art can implement the present application as set forth in the claims without undue experimentation. It can also be understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present application, which is determined by the full scope of the appended claims and their equivalents.
[0145] If a function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.
[0146] The logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a predefined sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, system, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, system, or device), or in combination with these instruction execution systems, systems, or devices. For the purposes of this specification, a "computer-readable medium" can be any system that can contain, store, communicate, propagate, or transport a program for use by or in combination with an instruction execution system, system, or device.
[0147] More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection part (electronic system) having one or more wirings, a portable computer disk cartridge (magnetic system), random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), an optical fiber system, and a portable compact disc read-only memory (CDROM). Additionally, a computer-readable medium can even be paper or other suitable media on which a program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other suitable processing as necessary, and then stored in a computer memory.
[0148] It should be understood that various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0149] In the foregoing description of this specification, descriptions with reference to the terms "one embodiment / example", "another embodiment / example", or "certain embodiments / examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0150] Although the embodiments of the present application have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present application, and the scope of the present application is defined by the claims and their equivalents.
[0151] The above has specifically described the preferred embodiments of the present application, but the present application is not limited to the embodiments. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the present application, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present application.
[0152] In the description of this specification, descriptions with reference to the terms "one embodiment", "another embodiment", or "certain embodiments", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0153] Although the embodiments of the present application have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present application, and the scope of the present application is defined by the claims and their equivalents.
Claims
1. A video optimization processing method, characterized in that: The method comprises: Get the target video clip; Performing frame extraction processing on the target video segment to obtain a video frame set, and determining a first resolution corresponding to the target video segment according to the video frame set; If the first resolution is greater than or equal to a first preset threshold, inputting the video frame set into a video quality scoring model to obtain a first quality score corresponding to the target video segment; If the first image quality score is less than a second preset threshold, inputting the video frame set into a video scene classification model to obtain a first scene prediction result corresponding to the target video clip; calculating a difference between the first resolution and a predetermined target resolution; According to the first scene prediction result and the difference, a target super-resolution model is determined from a super-resolution model database, and the target video clip is optimized based on the target super-resolution model; wherein the super-resolution model database includes multiple super-resolution models, each of the super-resolution models includes a scene type parameter and an optimization ratio parameter, the first scene prediction result is used to determine the scene type parameter corresponding to the target super-resolution model, and the difference is used to determine the optimization ratio parameter corresponding to the target super-resolution model.
2. A video optimization processing method according to claim 1, characterized in that: The step of performing frame extraction processing on the target video segment to obtain a video frame set includes: Performing frame extraction processing on the target video segment to obtain original video frames; Cropping an image content of a predetermined size from a central area of the original video frame to obtain a first video frame; Performing pixel normalization processing on the first video frame to obtain a second video frame; The second video frames are integrated to obtain the video frame set.
3. A video optimization processing method according to claim 1, characterized in that: The method further comprises: If the first resolution is less than the first preset threshold, the video frame set is input into a video scene classification model to obtain a first scene prediction result corresponding to the target video segment.
4. A video optimization processing method according to claim 1, characterized in that: The video quality scoring model is trained by the following steps: Obtaining a sample video segment and a first label value corresponding to the sample video segment; the first label value is used to represent a true result of a picture quality score corresponding to the sample video segment; Performing frame extraction processing on the sample video clip to obtain a sample frame set; Inputting the sample frame set into the video quality scoring model to obtain a second quality score corresponding to the sample video segment; Determining a first loss value according to the first label value and the second image quality score; According to the first loss value, the parameters of the video quality scoring model are updated to obtain a trained video quality scoring model.
5. A video optimization processing method according to claim 1, characterized in that: The video scene classification model is trained by the following steps: Acquire a sample video clip and a second label value corresponding to the sample video clip; the second label value is used to represent a true result of the scene type corresponding to the sample video clip; Performing frame extraction processing on the sample video clip to obtain a sample frame set; Inputting the sample frame set into the video scene classification model to obtain a second scene prediction result corresponding to the sample video clip; Determine a second loss value according to the second label value and the second scenario prediction result; According to the second loss value, the parameters of the video scene classification model are updated to obtain a trained video scene classification model.
6. A video optimization processing method according to claim 5, characterized in that: The determining a second loss value according to the second label value and the second scenario prediction result includes: According to the second label value and the second scene prediction result, a second loss value is determined by a cross entropy loss function.
7. A video optimization processing method according to claim 1, characterized in that: The super-resolution model includes at least one of SRCNN, ESPCN, VDSR, and SRGAN.
8. A video optimization processing system, characterized in that: The system comprises: An acquisition unit, used for acquiring a target video segment; a frame extraction unit, configured to perform frame extraction processing on the target video segment to obtain a video frame set, and determine a first resolution corresponding to the target video segment according to the video frame set; A first prediction unit, configured to input the video frame set into a video quality scoring model to obtain a first quality score corresponding to the target video segment if the first resolution is greater than or equal to a first preset threshold; A second prediction unit, configured to input the video frame set into a video scene classification model to obtain a first scene prediction result corresponding to the target video segment if the first image quality score is less than a second preset threshold; a calculation unit, configured to calculate a difference between the first resolution and a predetermined target resolution; A processing unit is used to determine a target super-resolution model from a super-resolution model database according to the first scene prediction result and the difference, and optimize the target video clip based on the target super-resolution model; wherein the super-resolution model database includes multiple super-resolution models, each of the super-resolution models includes a scene type parameter and an optimization ratio parameter, the first scene prediction result is used to determine the scene type parameter corresponding to the target super-resolution model, and the difference is used to determine the optimization ratio parameter corresponding to the target super-resolution model.
9. A computer device, characterized in that: include: at least one processor; at least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements a video optimization processing method as described in any one of claims 1-7.
10. A computer-readable storage medium storing a program executable by a processor, characterized in that: The processor-executable program is used to implement a video optimization processing method as described in any one of claims 1-7 when executed by the processor.
Citation Information
Patent Citations
Super-resolution reconstruction method, device and equipment apparatus and storage medium
CN111340711A
Video real-time super-resolution processing method and device, terminal and storage medium
CN115361582A