Video Encoding Method, Apparatus, Electronic Device, and Storage Medium
By extracting video encoding features and using a pre-trained neural network to determine a target quality rank for hardware video editors, the method addresses the issue of unreasonable bit rates, optimizing image quality and efficiency in video encoding.
Patent Information
- Application Number
- JP2024573687
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-12-07
- Filing Date
- 2023-12-05
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2043-12-05
Smart Images

Figure 2025522455000001_ABST
Abstract
Description
Technical Field
[0001] This application was filed on December 7, 2022, claims the priority of a Chinese patent application with the title "Video Encoding Method, Apparatus, Electronic Device and Storage Medium" and the application number 202211567316.X, and the entire content thereof is incorporated herein by reference.
[0002] Embodiments of the present disclosure relate to the field of image processing technologies, and in particular, to video encoding methods, apparatuses, electronic devices, and storage media.
Background Art
[0003] A graphics processing unit (GPU) is a processor that plays a role in executing image processing tasks in terminal devices such as mobile phones and personal computers. The graphics processing unit has powerful digital computing capabilities and parallel processing capabilities. In application scenarios such as video encoding and decoding, by using the graphics processing unit, the quality and efficiency of image encoding and decoding can be effectively improved.
[0004] Some graphics processing units include a hardware-based video encoder, also called a hardware video editor, such as an Nvenc unit. The hardware video editor can encode data in YUV / RGB format into a video conforming to the H.264 / HEVC standard, realizing efficient video encoding.
[0005] In actual applications, when calling a hardware video editor to encode a video, it is necessary to set parameters characterizing the quality rank. However, in the prior art, usually, encoding is performed using a quality rank set fixedly based on experience, resulting in the problem that the bit rate of the encoded video is unreasonable.
Summary of the Invention
[0006] Embodiments of the present disclosure provide a video encoding method, an apparatus, an electronic device, and a storage medium for overcoming the problem that the bit rate of an encoded video caused by encoding using a fixed quality rank is unreasonable.
[0007] In a first aspect, embodiments of the present disclosure provide a video encoding method, the method comprising: extracting video encoding features of initial video data by a hardware video editor, where the video encoding features characterize the complexity of the video content of the initial video data; processing the video encoding features of the initial video data by a pre-trained prediction neural network model to obtain a target quality rank corresponding to the initial video data, where the target quality rank characterizes the image quality level when encoding the initial video in a constant quality variable bit rate mode; and encoding the initial video data with the target quality rank using the hardware video editor to generate a target video.
[0008] In a second aspect, embodiments of the present disclosure provide a video encoding apparatus, the apparatus comprising: a parameter acquisition module for extracting video encoding features of initial video data by a hardware video editor, where the video encoding features characterize the complexity of the video content of the initial video data; a parameter optimization module for processing the video encoding features of the initial video data by a pre-trained prediction neural network model to obtain a target quality rank corresponding to the initial video data, where the target quality rank characterizes the image quality level when encoding the initial video in a constant quality variable bit rate mode; and an encoding module for encoding the initial video data with the target quality rank using the hardware video editor to generate a target video.
[0009] In a third aspect, an embodiment of the present disclosure provides an electronic device, which includes a processor and a memory communicatively connected to the processor, wherein the memory stores computer-executable instructions, and the processor executes the computer-executable instructions stored in the memory to implement the video encoding method described in the first aspect and various possible designs of the first aspect.
[0010] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, on which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the video encoding method described in the first aspect and various possible designs of the first aspect is implemented.
[0011] In a fifth aspect, an embodiment of the present disclosure provides a computer program product including a computer program that, when executed by a processor, implements the video encoding method described in the first aspect and various possible designs of the first aspect.
[0012] The video encoding method, apparatus, electronic device, and storage medium provided by this embodiment extract the video encoding features of the initial video data by a hardware video editor, and the video encoding features characterize the complexity of the video content of the initial video data. The video encoding features of the initial video data are processed by a pre-trained prediction neural network model to obtain the target quality rank corresponding to the initial video data. The target quality rank characterizes the image quality level when encoding the initial video in a constant quality variable bitrate mode. The initial video data is encoded with the target quality rank using the hardware video editor to generate a target video. By utilizing the characteristic that the hardware video editor can conveniently output encoding parameters, the video encoding features corresponding to the initial video data are obtained, and then the target image quality level corresponding to the video encoding features is obtained by a pre-trained prediction neural network model. Furthermore, by encoding the video with the target quality rank using the hardware video editor, the bitrate of the generated target video is adapted to the video content, avoiding the problem of the bitrate being too high or too low, improving the image quality of the video, and avoiding waste of the bitrate.
Brief Description of the Drawings
[0013] To more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following briefly introduces the attached drawings that need to be used in the description of the embodiments or the prior art. However, the attached drawings in the following description are some embodiments of the present disclosure, and it is obvious for those of ordinary skill in the art that other drawings can be obtained according to these attached drawings without creative efforts.
[0014]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Embodiments for Carrying Out the Invention
[0015] In order to make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, hereinafter, the technical solutions in the embodiments of the present disclosure will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present disclosure. It is obvious that the described embodiments are only a part of the embodiments of the present disclosure, not all of them. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without creative efforts shall be included in the protection scope of the present disclosure.
[0016] The application scenarios of the embodiments of the present disclosure will be described below.
[0017] The video encoding method provided by the embodiments of the present disclosure can be applied to various application scenarios that require video encoding, such as video editing, previewing, playing, etc. More specifically, for example, it is applied to video editing software, video editing cloud platforms, and live streaming software. Exemplarily, the method provided by the embodiments of the present disclosure may be applied to terminal devices such as smartphones, tablet computers, and personal computers, or may be applied to cloud servers. FIG. 1 is a schematic diagram of an application scenario provided by the embodiments of the present disclosure. Taking the application scenario of running video editing software on the terminal device side as an example, as shown in FIG. 1, specifically, the terminal device runs the video editing software, performs video editing on the original video, for example, adds video effects, adds audio tracks, adds subtitles, and then generates video data. The video data may include a plurality of video frames and editing information corresponding to each of the frames. Then, the terminal device calls a hardware video editor in the graphics processing unit to encode the video data and generate a target video (completed video) with editing effects for playback, completing the workflow of video editing. Here, the hardware video editor is, for example, an Nvenc unit.
[0018] In the prior art, when calling a hardware video editor to process video data, it is necessary to set corresponding parameters to control a specific encoding method. Among them, variable bit rate (VBR) encoding, also called dynamic bit rate encoding or non-fixed bit rate encoding, is a commonly used encoding method. Since it can dynamically adjust the bit rate according to the content of the video, the bit rate of the encoded video can change according to the complexity of the image. Therefore, its encoding efficiency is relatively high, the capacity of the video in a still image is compressed, there are fewer mosaics in a dynamic moving image, and both the capacity and quality of the video are balanced. Among these, constant quality (CQ) variable bit rate (CQ-VBR) is an encoding method based on variable bit rate. By using a quality rank to control the image quality (bit rate) in the process of variable bit rate encoding, it is possible to realize more fine-grained control of the capacity and quality of the video generated after encoding, and further improve the flexibility and practicality of video encoding control.
[0019] However, in the actual application process, when calling a hardware video editor to encode using the constant quality variable bit rate (CQ-VBR) mode, the setting of the quality rank (cq value) is usually determined based on the user's experience. As a result, the quality rank is often set inappropriately, leading to problems such as the video bit rate being too high (the video capacity becomes too large, wasting storage and network resources) or too low (the video image quality is low, affecting the video viewing experience).
[0020] Embodiments of the present disclosure provide a video encoding method that solves the above problems by automatically generating a reasonable target quality rank (cq value) and performing video encoding based on this target quality rank.
[0021] Referring to FIG. 2, FIG. 2 is a schematic flowchart 1 of a video encoding method provided by an embodiment of the present disclosure. The method of this embodiment is applicable to an electronic device equipped with a hardware video editor, such as a terminal device or a server. In this embodiment, the terminal device is introduced as the execution subject. Exemplarily, this video encoding method includes the following steps.
[0022] Step S101: Extract the video encoding features of the initial video data by a hardware video editor, where the video encoding features characterize the complexity of the video content of the initial video data.
[0023] Exemplarily, after step S101, the terminal device first obtains the initial video data, which may be data generated based on a video editing application, a live streaming application, etc. For example, the video data in YUV / RGB format, and data such as video effects, added audio tracks, subtitles, etc. are included. The initial video data is the data to be encoded, and after encoding the initial video data, a corresponding playable video can be generated.
[0024] After that, the initial video data is processed by calling a hardware video editor. Specifically, for example, the initial video data is used as input parameters, the application interface of the hardware video editor is called, and the corresponding processing function is executed to process this initial video data, and the video encoding characteristics corresponding to this initial video data are obtained. Here, the video encoding characteristics characterize the complexity of the video content of the initial video data by the video parameters of the initial video data and the encoding parameters corresponding to the video parameters. Specifically, in one possible implementation form, the video encoding characteristics are processed based on the video parameters and the encoding parameters. Here, the video parameters are information characterizing this initial video data, for example, including the height, width (i.e., resolution) of the video, frame rate (fps), etc. The encoding parameters characterize the parameters used when the hardware video editor encodes the initial video data. For example, they characterize the frame encoding size (size), the type of each frame corresponding to the video data (including I frames, P frames, B frames), the distortion degree of the image, the sharpness of the image, etc.
[0025] Furthermore, the video parameters corresponding to the video encoding characteristics are directly obtained based on the video information of the initial video data, and the description thereof will not be repeated. The video parameters corresponding to the video encoding characteristics are obtained by taking the video parameters as input and calling the interface provided by the hardware video editor to generate the encoding parameters. Therefore, the encoding parameters have a corresponding relationship with the initial video data. In one possible implementation form, as shown in FIG. 3, the specific implementation steps of step S101 include the following steps.
[0026] Step S1011: Obtain the quality rank options of the initial video data.
[0027] Exemplarily, based on the application scenarios described so far and the prior art, when encoding using a hardware video editor in a constant quality variable bitrate mode, it is necessary to set a quality rank. The quality rank options may be predetermined default values. More specifically, the quality rank options may be values within the range of [18, 35], for example, 25. The smaller the quality rank, the higher the image quality level of the video generated after encoding, the clearer the video becomes, and relatively, the larger the video capacity becomes.
[0028] Step S1012: Input the video parameters and quality rank options of the initial video data into the hardware video editor to obtain the encoding parameters of the initial video data.
[0029] Furthermore, by calling the interface provided by the hardware video editor, input the video parameters and quality rank options of the initial video data into the hardware video editor as input quantities, and obtain the encoding parameters output by the hardware video editor, such as the frame encoding size (size), the type of each frame corresponding to the video data (including I frames, P frames, B frames), the distortion degree of the image, and the sharpness of the image. Here, exemplarily, the distortion degree of the image is represented by the sum of absolute transformed differences (SATD) of each frame, and the sharpness of the image may be represented by the quantization parameter (QP) of each frame. The specific calculation processes of the sum of absolute differences and the quantization parameter are existing technologies that can be realized by functions provided by the driver of the hardware video editor, and the description here will not be repeated.
[0030] Step S1013: Generate corresponding video encoding features based on the encoding parameters.
[0031] Exemplarily, the encoding parameters obtained by inputting video parameters and quality rank options into a hardware video editor correspond to pre-encoding of initial video data by the hardware video editor, that is, the hardware video editor predicts corresponding encoding parameters based on the initial video data, but does not actually perform encoding. Subsequently, based on the encoding parameters, video encoding features are generated by combining video parameters and quality rank options. This video encoding feature can represent the complexity of the video content of the initial video data. Subsequently, based on this video encoding feature, a quality rank that matches it can be determined for encoding, thereby achieving the purpose that the picture quality matches the complexity of the video content and avoiding the problem that the bit rate is wasted or too low.
[0032] Furthermore, in one possible implementation form, the initial video data includes a plurality of video frames. As shown in FIG. 4, the specific implementation steps of step S1013 are Step S1013A: Obtaining encoding parameters corresponding to each of the video frames, and Step S1013B: Based on the encoding parameters corresponding to each of the video frames, obtaining an encoding feature average value and an encoding feature variance value, where the encoding feature average value is the average value of the encoding parameters corresponding to each of the video frames, and the encoding feature variance value is the variance value of the encoding parameters corresponding to each of the video frames, and Step S1013C: Generating video encoding features based on the video parameters, quality rank options, and the encoding feature average value and encoding feature variance value corresponding to each of the video frames.
[0033] Exemplarily, the initial video data includes a plurality of video frames. By respectively obtaining the video parameters and encoding parameters corresponding to the video frames in the initial video data, frame encoding features corresponding to the video frames are obtained. Further, based on the average level and variation among the plurality of frame encoding features, the complexity of the video content of the initial video data is determined, and the video encoding features of the initial video data are obtained. FIG. 5 is a schematic diagram of a process for generating video encoding features provided by an embodiment of the present disclosure. Hereinafter, the above process will be described in conjunction with FIG. 5. As shown in FIG. 5, the initial video data includes N video frames, where N is an integer greater than 1. Here, taking the Mth frame as an example (M is an integer smaller than N and greater than 1), first, the video parameters of the Mth frame, for example, the height, width, and frame rate of the Mth frame are obtained. The video parameters and encoding level options are input into the interface of the hardware video editor, and using the hardware video editor, the encoding parameters of the Mth frame, for example, the type (Type as shown in the figure), frame encoding size (Size as shown in the figure), sum of absolute differences (SATD as shown in the figure), quantization parameter (QP as shown in the figure) of the Mth frame are obtained. Then, the average value and variance of the encoding parameters corresponding to the Mth frame are calculated to obtain the encoding feature average value and the encoding feature variance value. Here, specifically, the average value and variance of the encoding parameters corresponding to the Mth frame are the average value and variance of the encoding parameters of each video frame within the set composed of at least one adjacent video frame before the Mth frame (shown as the Lth frame in the figure, where L is smaller than M and an integer greater than or equal to 1) and the Mth frame. For the specific implementation of the encoding parameters, the corresponding average value and variance are calculated, and the specific calculation process for obtaining the encoding feature average value and the encoding feature variance value is not repeated.After that, based on combining video parameters, quality rank options, average encoded feature values, and variance values of encoded features, video encoding features corresponding to the M-th frame are obtained. The video encoding features indicate the complexity of the video content of the video segment corresponding to the M-th frame from the L-th frame in the initial video data. There is a possibility that L = 1, and the video encoding features indicate the complexity of the video content of the video segment before the M-th frame in the initial video data.
[0034] Furthermore, FIG. 6 is a schematic diagram of the data structure of video encoding features provided by an embodiment of the present disclosure. Referring to FIGS. 5 and 6, exemplarily, the video encoding features include a total of 21 data fields, which are shown as Field #1 to #21 in the figure.
[0035] Field #1 indicates the height of the images from frame 1 to frame M in the initial video data.
[0036] Field #2 indicates the width of the images from frame 1 to frame M in the initial video data.
[0037] Field #3 indicates the frame rate from frame 1 to frame M in the initial video data.
[0038] Field #4 indicates the quality rank options from frame 1 to frame M in the initial video data.
[0039] Field #5 indicates the number of I-frames from frame 1 to frame M in the initial video data.
[0040] Field #6 indicates the number of P-frames from frame 1 to frame M in the initial video data.
[0041] Field #7 indicates the number of B-frames from frame 1 to frame M in the initial video data.
[0042] Field #8 indicates the average size of I-frames from Frame 1 to Frame M in the initial video data.
[0043] Field #9 indicates the average size of P-frames from Frame 1 to Frame M in the initial video data.
[0044] Field #10 indicates the average size of B-frames from Frame 1 to Frame M in the initial video data.
[0045] Field #11 indicates the sum of average absolute errors of I-frames from Frame 1 to Frame M in the initial video data.
[0046] Field #12 indicates the sum of average absolute errors of P-frames from Frame 1 to Frame M in the initial video data.
[0047] Field #13 indicates the sum of average absolute errors of B-frames from Frame 1 to Frame M in the initial video data.
[0048] Field #14 indicates the average quantization parameter from Frame 1 to Frame M in the initial video data.
[0049] Field #15 indicates the size variance of I-frames from Frame 1 to Frame M in the initial video data.
[0050] Field #16 indicates the size variance of P-frames from Frame 1 to Frame M in the initial video data.
[0051] Field #17 indicates the size variance of B-frames from Frame 1 to Frame M in the initial video data.
[0052] Field #18 indicates the variance of the sum of average absolute errors of I-frames from Frame 1 to Frame M in the initial video data.
[0053] Field #19 indicates the sum of mean absolute error variances of P-frames from frame 1 to frame M in the initial video data.
[0054] Field #20 indicates the sum of mean absolute error variances of B-frames from frame 1 to frame M in the initial video data.
[0055] Field #21 indicates the quantization parameter variance from frame 1 to frame M in the initial video data.
[0056] Step S102: Process the video encoding features of the initial video data using a pre-trained prediction neural network model to obtain a target quality rank corresponding to the initial video data, where the target quality rank characterizes the image quality level when initially encoding in a constant quality variable bitrate mode for the video.
[0057] Exemplarily, after obtaining the video encoding features of the initial video data, it is necessary to determine a target quality rank that fits these video encoding features of the initial video data. Specifically, the quality rank, i.e., the cq value, characterizes the image quality level when initially encoding the initial video in a constant quality variable bitrate (CQ-VBR) mode and is one of the parameters that need to be used when calling a hardware video editor. In the steps of this embodiment, the video encoding features of the initial video data are processed by a pre-trained prediction neural network model to predict a quality rank, i.e., the target quality rank, that matches the complexity of the video content it characterizes.
[0058] Step S103: Use a hardware video editor to encode the initial video data with the target quality rank to generate a target video.
[0059] Exemplarily, further, after obtaining the target quality rank, a hardware video editor can be called to process the initial video data using the target quality rank as a parameter to generate a corresponding playback-ready video finished product, i.e., the target video.
[0060] In one possible implementation form, the target quality rank may be a level sequence including a plurality of level identifiers characterizing specific quality ranks (i.e., cq values). Here, each level identifier within the level sequence corresponds to one or more video frames within the initial video data. In a more specific possible implementation form, each level identifier corresponds to one video frame. When encoding the initial video data with the target quality rank using a hardware video editor, each video frame within the initial video data is sequentially (either in parallel or serially) obtained, and based on the level identifier corresponding to each video frame, the hardware video editor is called to encode the corresponding initial video data at a constant quality variable bit rate. As a result, each encoded video frame has a different image quality level, enabling more accurate encoding and improved encoding efficiency. Among these, the specific implementation process of calling the hardware video editor to encode at a constant quality variable bit rate is a prior art and will not be repeated here.
[0061] In this embodiment, a hardware video editor extracts the video encoding features of the initial video data. The video encoding features characterize the complexity of the video content of the initial video data. The video encoding features of the initial video data are processed by a pre-trained prediction neural network model to obtain a target quality rank corresponding to the initial video data. The target quality rank characterizes the image quality level when encoding the initial video in a constant quality variable bitrate mode. The initial video data is encoded with the target quality rank using the hardware video editor to generate a target video. Utilizing the characteristic that the hardware video editor conveniently outputs encoding parameters, the video encoding features corresponding to the initial video data are obtained. Then, a target image quality level matching the video encoding features is obtained by the pre-trained prediction neural network model. Furthermore, by using the hardware video editor to encode the video with the target quality rank, the bitrate of the generated target video is adapted to the content of the video, avoiding the problem of the bitrate being too high or too low, improving the image quality of the video, and avoiding waste of the bitrate.
[0062] Referring to FIG. 7, FIG. 7 is a schematic flowchart 2 of a video encoding method provided by an embodiment of the present disclosure. This embodiment further subdivides the implementation processes of steps S101 and S102 on the basis of the embodiment shown in FIG. 2. This video encoding method includes the following steps.
[0063] Step S201: Obtain a quality rank sequence characterizing quality rank options, where the quality rank sequence is a set of a plurality of quality ranks arranged in order.
[0064] Exemplarily, referring to the introduction to the quality rank options in the embodiment shown in FIG. 2, in this embodiment, there are multiple quality rank options, and the multiple quality rank options are characterized by a preset quality rank sequence. That is, the quality rank sequence is one of the realization forms of the quality rank options. Specifically, for example, the quality rank sequence is cq_data = [15:40], that is, it is an ordered arrangement of multiple quality ranks from quality rank 15 to quality rank 40. The multiple quality ranks from quality rank 15 to quality rank 40 arranged in this order are quality rank options. It is understood that the quality rank sequence can be implemented in various ways, such as an array, a matrix, a key-value pair, an enumerated data structure such as a structure, or a set of numerical values represented as a function, so no examples are given here.
[0065] Step S202: Obtain the current frame of the initial video data.
[0066] Step S203: Obtain the video parameters of the current frame and the current quality rank in the quality rank sequence.
[0067] Step S204: Input the video parameters of the current frame and the quality rank of the current frame into the hardware video editor to obtain the encoding parameters corresponding to the current quality rank of the current frame.
[0068] Exemplarily, from steps S202 and S203, in a way that circulates in two dimensions (video frame dimension, quality rank dimension) respectively, obtain the corresponding target quality rank that matches each video frame in the initial video parameters. Therefore, by encoding each frame based on the target quality rank corresponding to each video frame, dynamic encoding for each frame in the video is realized, improving the encoding efficiency and encoding quality.
[0069] Specifically, starting from the first frame in the initial video data and sequentially taking the current frame until the last frame of the initial video data, for each current frame, video parameters can be directly obtained, and the description here will not be repeated. Then, based on the fact that multiple quality ranks and video parameters in the quality data options are sequentially input into the hardware video editor, encoding parameters corresponding to the current quality rank of the current frame can be obtained. Among these, the encoding parameters of the first frame of the initial video data can be obtained by calling the corresponding interface of the hardware video editor in combination with the video parameters based on a predetermined default quality rank. The specific implementation form has been described in the above embodiments, so it will not be repeated here.
[0070] Step S205: Generate a first feature corresponding to the current quality rank of the current frame from the video parameters of the current frame, the current quality rank of the current frame, and the encoding parameters corresponding to the current quality rank.
[0071] Furthermore, after obtaining the video parameters of the current frame, the current quality rank of the current frame, and the encoding parameters corresponding to the current quality rank, combine the above video parameters, current quality rank, and encoding parameters to generate a frame encoding feature corresponding to the current quality rank, that is, the first feature. That is, the first feature includes video parameters, the current quality rank, and encoding parameters. Specifically, for example, the video parameters include the height, width, resolution, and frame rate of the image, and the encoding parameters include the frame encoding size, the type of each frame corresponding to the video data (including I-frame, P-frame, and B-frame), the distortion degree of the image, and the sharpness of the image. Integrate the video parameters of the current frame, the current quality rank of the current frame, and the encoding parameters corresponding to the current quality rank of the current frame to generate the first feature. F = {h, w, fps, cq, type, size, satd, gp}. Here, F represents the first feature of the current frame, h represents the height of the image, w represents the width of the image, fps represents the frame rate of the image, and the above h, w, and fps are video parameters. cq represents the current quality rank, type represents the type of the current frame, size represents the size of the current frame, satd represents the sum of absolute differences of the current frame, and gp represents the quantization parameter of the current frame. The above type, size, satd, and gp are encoding parameters.
[0072] Step S206: Obtain the second feature corresponding to the target quality rank of the previous frame corresponding to the current frame, where the previous frame is the first predetermined number of video frames adjacent to the current frame before the current frame.
[0073] Step S207: Generate an encoded feature option corresponding to the current quality rank of the current frame based on the first feature and the second feature.
[0074] Further, after obtaining the first feature corresponding to the current frame, obtain the second feature generated by the previous frame corresponding to the current frame based on the target quality rank, where the second feature is data similar to the first feature and characterizing the frame encoding feature of the previous frame. Specifically, the previous frame corresponding to the current frame is the first predetermined number of video frames adjacent to the current frame before the current frame, for example, 30 video frames before the current frame. For the specific implementation of the previous frame, reference can be made to the embodiment corresponding to FIG. 5, that is, the set of direct video frames from the L-th frame to the M-th frame is the previous frame. More specifically, when the current frame is the second frame of the initial video data, the previous frame of the current frame is the first frame of the initial video data, and the second feature corresponding to the first frame is a frame encoding feature generated based on the default quality rank, and the specific process thereof will not be repeated here. Then, based on the frame encoding features corresponding to the plurality of video frames (the first frame and the second frame of the initial video data) composed of the second feature of the first frame and the first feature of the second frame, calculate the average value and the variance value of the encoding parameters in the frame encoding feature corresponding to the current frame, and obtain the video encoding feature corresponding to the second frame, that is, the encoding feature option corresponding to the current quality rank of the current frame. For the specific implementation process, reference can be made to the detailed description of the process of obtaining the video encoding feature in the embodiment shown in FIG. 5, so it will not be repeated here. And when the current frame is a video frame after the third frame of the initial video data, since each video frame in the initial video data is sequentially processed as the current frame, when the process reaches the third frame, the target quality rank (i.e., the optimized quality rank) corresponding to the previous frame (for example, the second frame) has already been obtained. At this time, the frame encoding feature corresponding to the previous frame of the current frame, that is, the second feature, is generated based on the target quality rank corresponding to this previous frame.
[0075] In this embodiment, by obtaining a second feature corresponding to the target quality rank of the previous frame corresponding to the current frame, based on the second feature and the first feature, an encoding feature option corresponding to the current quality rank characterizing the complexity of the video content is generated. Since the second feature of the previous frame is generated based on the optimized target quality rank, the encoding feature option generated based on the second feature can more accurately represent the complexity of the video content, the accuracy of the encoding feature option is improved, and finally the accuracy of the target quality rank obtained based on the encoding feature option is improved, and the efficiency of video encoding is improved.
[0076] Step S208: Input the encoding feature option into the prediction neural network model to obtain a corresponding first evaluation value and a corresponding second evaluation value. Here, the first evaluation value characterizes the video quality evaluation value based on video multi-modal evaluation fusion, and the second evaluation value characterizes the video bit rate.
[0077] Step S209: If the current quality rank is at the end of the quality rank sequence, continue to execute Step S210; otherwise, return to execute Step S203.
[0078] Furthermore, for each current frame in the loop process, as the current quality rank changes cyclically, the occurrence of the correspondingly generated first feature also changes. Furthermore, the encoding feature option corresponding to the generated current quality rank also changes. In each cycle round corresponding to the current quality rank, input the encoding feature option obtained in Step S207 into the prediction neural network model to obtain the first evaluation value and the second evaluation value output by the prediction neural network model. Here, the first evaluation value characterizes the video quality evaluation value based on video multi-modal evaluation fusion, and the second evaluation value characterizes the video bit rate.
[0079] Here, Video Multimethod Assessment Fusion (VMAF) is a video quality assessment metric used to measure the perception of streaming video quality in a large-scale environment, which can solve the problem that conventional metrics cannot reflect video situations with multiple scenes and multiple features. The specific implementation form of video multimethod assessment fusion is prior art. The video bitrate can be obtained from the quantization parameter in the encoding parameters, which will not be repeated here. By a pre-trained prediction neural network model, the input video encoding features (encoding feature options) can be mapped to the corresponding video multimethod assessment fusion metric and video bitrate. FIG. 8 is a schematic diagram of a process for generating an evaluation value corresponding to the current quality rank of the current frame provided by an embodiment of the present disclosure. As shown in FIG. 8, first, after performing video frame traversal on the initial video data, quality rank traversal is performed on each video frame, and the current quality rank is obtained when traversing the current frame. Then, based on the steps of the above embodiment, the encoding feature option corresponding to the current quality rank is obtained, and the encoding feature option is input into the prediction neural network model. The prediction neural network model outputs a first evaluation value and a second evaluation value. Then, the first evaluation value and the second evaluation value are saved as a set of encoding feature option-evaluation value mapping data together with the corresponding encoding feature option (and / or the current quality rank). Then, if the current quality rank is at the end of the quality rank sequence, it means that all the quality ranks of the quality rank sequence corresponding to the current frame have been traversed, and step S210 is executed to select a target quality rank from multiple quality rank options. If the current quality rank is not at the end of the quality rank sequence, return to step S203, cycle to the next set of quality ranks (update the current quality rank), and repeat the above process until all the quality ranks of the quality rank sequence have been traversed.
[0080] Step S210: Generate the target quality rank of the current frame based on the first evaluation value corresponding to each of the encoding feature options of the current frame and the corresponding second evaluation value.
[0081] Since the encoding feature options are generated based on the quality rank, each encoding feature option corresponds to one quality rank. After obtaining the first evaluation value and the second evaluation value corresponding to each quality rank in the quality rank sequence of the current frame, evaluate each quality rank in the quality rank sequence based on the first evaluation value and the second evaluation value corresponding to each quality rank in the quality rank sequence of the current frame, and obtain the optimal quality rank, that is, the target quality rank.
[0082] Exemplarily, the specific implementation steps of step S210 include the following.
[0083] Step S2101: Obtain the first target encoding feature based on the first evaluation value corresponding to each encoding feature option. The first target encoding feature is an encoding feature option whose first evaluation value is greater than the first threshold.
[0084] Step S2102: Determine the second target encoding feature based on the second evaluation value of the first target encoding feature. The second target feature is the video encoding feature with the smallest second evaluation value among the first target encoding features.
[0085] Step S2103: Obtain the target quality rank based on the quality rank corresponding to the second target feature.
[0086] Exemplarily, the first evaluation value and the second evaluation value respectively characterize the video quality and the bit rate after the video is encoded. The higher the first evaluation value, the higher the video quality of the video, and the higher the second evaluation value, the higher the bit rate, that is, the larger the video capacity. In order to improve the efficiency of video encoding, it is necessary to select the quality rank that can most significantly reduce the bit rate on the premise of satisfying the requirement of the video quality of a predetermined video.
[0087] To solve the above problems, in this embodiment, first, based on the first evaluation value corresponding to each encoding feature option, an encoding feature option whose first evaluation value (that is, video quality) is greater than the first threshold is determined as the first target encoding feature. Then, from the first target encoding feature, a video encoding feature with the smallest second evaluation value (that is, the smallest bit rate) is selected. Then, the quality rank corresponding to the video encoding feature is obtained as the target quality rank. Specifically, the above process can be realized by the encoding feature option - evaluation value mapping data saved in the previous step, and the specific process is not repeated.
[0088] In this embodiment, by combining video multi - mode evaluation fusion and the bit rate index to obtain a corresponding target quality rank, on the premise that the target video encoded based on the target quality rank meets the requirement of the video quality of a predetermined video, the bit rate can be reduced, the compression of the video capacity can be realized, and the efficiency of video encoding can be improved.
[0089] Step S211: Use a hardware video editor to encode the current frame with the target quality rank and generate a target frame for constructing the target video.
[0090] Step S212: If the current frame is not the last frame of the initial video data, set the next frame of the current frame as the new current frame and return to step S202.
[0091] Exemplarily, after obtaining the target quality rank, the current frame is encoded based on the target quality rank to obtain a target frame having better video quality and a lower bit rate. At the same time, if the current frame is not the last frame of the original video data, the process returns to step S202, traverses all video frames, and continues the above process for the next video frame until corresponding target frames are generated to constitute a target video. Since dynamic encoding is performed using different target quality ranks for each video frame, it is possible to improve video quality, reduce video capacity, and improve video encoding efficiency at the same time.
[0092] Optionally, according to specific requirements, before step S208, it further includes a step of training a prediction neural network model. Exemplarily, as shown in FIG. 9, the step of training a prediction neural network model includes the following.
[0093] Step S2001: Obtain the original video data and a quality rank sequence, and the quality rank sequence includes at least two different quality ranks.
[0094] Step S2002: Based on the quality rank sequence, sequentially process the original video data using a hardware video editor to obtain video encoding features corresponding to each of the quality ranks.
[0095] Step S2003: Calculate a first evaluation value and a second evaluation value corresponding to each of the video encoding features.
[0096] Step S2004: Generate a training sample based on each video encoding feature, the corresponding first evaluation value, and the corresponding second evaluation value, and train a predetermined neural network model based on the training sample to obtain a prediction neural network model.
[0097] Exemplarily, the original video data is video data samples, the quality rank sequence is predetermined parameters, and by calling a hardware video editor with the original video data and the quality rank sequence as input parameters, the corresponding encoding parameters output by the hardware video editor can be obtained, and video encoding features can be obtained. For the specific generation process of the video encoding features, reference can be made to the description in the above embodiments and will not be repeated here. Then, based on the calculation method of the video multi-mode evaluation fusion index and the encoding parameters output by calling the hardware video editor, the first evaluation value and the second evaluation value corresponding to each of the video encoding features are obtained. Using the first evaluation value and the second evaluation value as sample labels, the original video data and the level sequence as the original samples, training samples are generated. Based on the training samples, a predetermined neural network model is trained until convergence, and a prediction neural network model is obtained. The description of the specific process is omitted.
[0098] In this embodiment, by utilizing the feature that the hardware video editor can output encoding parameters simply and quickly (without actually encoding), and combining the video multi-mode evaluation fusion index to generate training samples, efficient and high-quality training of the neural network model can be realized, and the model can be quickly converged. This improves the prediction accuracy and efficiency of the model.
[0099] Corresponding to the video encoding method of the above embodiment, FIG. 10 is a configuration block diagram of a video encoding apparatus provided by an embodiment of the present disclosure. For ease of explanation, only the parts related to the embodiments of the present disclosure are shown. Referring to FIG. 10, the video encoding apparatus 3 includes a parameter acquisition module 31, a parameter optimization module 32, and an encoding module 33.
[0100] The parameter acquisition module 31 is used by a hardware video editor to extract the video encoding features of the initial video data, and the video encoding features characterize the complexity of the video content of the initial video data.
[0101] The parameter optimization module 32 is used to process the video encoding features of the initial video data by a pre-trained prediction neural network model to obtain the target quality rank corresponding to the initial video data, and the target quality rank characterizes the image quality level when encoding the initial video in a constant quality variable bitrate mode.
[0102] The encoding module 33 is used to encode the initial video data with the target quality rank using a hardware video editor to generate a target video.
[0103] In an embodiment of the present disclosure, the parameter acquisition module 31 is specifically used for obtaining the quality rank options of the initial video data, inputting the video parameters and quality rank options of the initial video data into the hardware video editor to obtain the encoding parameters of the initial video data, and generating corresponding video encoding features based on the encoding parameters.
[0104] In an embodiment of the present disclosure, the initial video data includes a plurality of video frames. When the parameter acquisition module 31 generates corresponding video encoding features based on the encoding parameters, specifically, it acquires the encoding parameters corresponding to each of the video frames, and based on the encoding parameters corresponding to each of the video frames, it acquires the average value of the encoding features and the variance value of the encoding features. The average value of the encoding features is the average value of the encoding parameters corresponding to each of the video frames, and the variance value of the encoding features is the variance value of the encoding parameters corresponding to each of the video frames. It is used to generate video encoding features based on the video parameters, quality rank options, and the average value of the encoding features and the variance value of the encoding features corresponding to each of the video frames.
[0105] In an embodiment of the present disclosure, when the parameter acquisition module 31 inputs the video parameters and quality rank options of the initial video data into the hardware video editor to acquire the encoding parameters of the initial video data, specifically, it is used to repeatedly execute the following steps until a predetermined condition is reached. The steps include: acquiring the current frame of the initial video data; acquiring the video parameters of the current frame and the quality rank of the current frame; inputting the video parameters of the current frame and the quality rank of the current frame into the hardware video editor to acquire the encoding parameters corresponding to the current frame; and setting the next frame of the current frame as the new current frame.
[0106] In an embodiment of the present disclosure, after obtaining the encoding parameters corresponding to the current frame, the parameter acquisition module 31 is further used to generate the frame encoding features corresponding to the current frame from the video parameters of the current frame, the quality rank of the current frame, and the encoding parameters, and to obtain the frame encoding features of the previous frame corresponding to the current frame. The previous frame is the first predetermined number of video frames adjacent to the current frame before the current frame, and the frame encoding features of the previous frame are generated based on the target quality rank corresponding to the previous frame.
[0107] When the parameter acquisition module 31 generates the corresponding video encoding features based on the encoding parameters, specifically, it is used to generate the video encoding features corresponding to the current frame based on the frame encoding features corresponding to the current frame and the frame encoding features of the previous frame.
[0108] In an embodiment of the present disclosure, the encoding parameters include at least one of frame type, frame size, image distortion degree, and image sharpness.
[0109] In an embodiment of the present disclosure, the video encoding features of the initial video data include at least two encoding feature options. Each encoding feature option corresponds to a different quality rank. Specifically, the parameter optimization module 32 sequentially inputs each encoding feature option into the prediction neural network model to obtain the corresponding first evaluation value and the corresponding second evaluation value. The first evaluation value characterizes the video quality evaluation value based on video multi-modal evaluation fusion, and the second evaluation value characterizes the video bit rate. It is used to generate the target quality rank based on the first evaluation value and the corresponding second evaluation value corresponding to each of the encoding feature options.
[0110] In an embodiment of the present disclosure, when generating a target quality rank based on a first evaluation value corresponding to each of the encoded feature options and a corresponding second evaluation value, specifically, the parameter optimization module 32 obtains a first target encoded feature based on the first evaluation value corresponding to each of the encoded feature options, where the first target encoded feature is an encoded feature option for which the first evaluation value is greater than a first threshold, determines a second target encoded feature based on the second evaluation value of the first target encoded feature, where the second target feature is a video encoding feature with the smallest second evaluation value among the first target encoded features, and obtains a target quality rank based on the quality rank corresponding to the second target encoded feature.
[0111] In an embodiment of the present disclosure, before processing the video encoding features of the initial video data by a pre-trained prediction neural network model to obtain the target quality rank corresponding to the initial video data, the parameter optimization module 32 obtains the original video data and a quality rank sequence, where the quality rank sequence includes at least two different quality ranks, sequentially processes the original video data using a hardware video editor based on the quality rank sequence, and obtains the video encoding features corresponding to each of the quality ranks, calculates a first evaluation value and a second evaluation value corresponding to each video encoding feature, generates a training sample based on each video encoding feature, the corresponding first evaluation value, and the corresponding second evaluation value, and further uses the training sample to train a predetermined neural network model to obtain a prediction neural network model.
[0112] Here, the parameter acquisition module 31, the parameter optimization module 32, and the encoding module 33 are connected in sequence. The video encoding device 3 provided by this embodiment can execute the technical solution of the above-described method embodiment, and its implementation principle and technical effect are the same, and the description of this embodiment will not be repeated here.
[0113] FIG. 11 is a schematic configuration diagram of an electronic device provided according to an embodiment of the present disclosure. As shown in FIG. 11, this electronic device 4 includes a processor 41 and a memory 42 communicably connected to the processor 41. The memory 42 stores computer-executable instructions. The processor 41 executes the computer-executable instructions stored in the memory 42 to implement the video encoding method in the embodiments shown in FIGS. 2 to 9.
[0114] Here, optionally, the processor 41 and the memory 42 are connected via a bus 43.
[0115] Since the related description can be understood by referring to the related description and effects corresponding to the steps in the embodiments corresponding to FIGS. 2 to 9, the description here is omitted.
[0116] Referring to FIG. 12 showing a schematic configuration diagram of an electronic device 900 suitable for implementing the embodiments of the present disclosure, this electronic device 900 may be a terminal device or a server. Here, the terminal device may include, but is not limited to, mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, personal digital assistants (abbreviated as PDA), tablet computers (abbreviated as PAD), portable media players (abbreviated as PMP), in-vehicle terminals (for example, vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Note that the electronic device shown in FIG. 12 is merely an example and does not limit the functions and usage ranges of the embodiments of the present disclosure in any way.
[0117] As shown in FIG. 12, the electronic device 900 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 901 that can execute various appropriate operations and processes based on a program stored in a read only memory (abbreviated as ROM) 902 or a program loaded from a storage device 908 into a random access memory (abbreviated as RAM) 903. Also, various programs and data necessary for the operation of the electronic device 900 are stored in the RAM 903. The processing device 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0118] Typically, the I / O interface 905 may be connected to an input device 906 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc., an output device 907 including, for example, a liquid crystal display (abbreviated as LCD), a speaker, a vibrator, etc., a storage device 908 including, for example, a magnetic tape, a hard disk, etc., and a communication device 909. The communication device 909 may enable the electronic device 900 to communicate with other devices wirelessly or wiredly to exchange data. Although FIG. 12 shows an electronic device 900 equipped with various devices, it should be understood that it is not necessary to implement or provide all of the shown devices. More or fewer devices may be alternatively implemented or provided.
[0119] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart may be implemented as a computer software program. For example, an embodiment of the present disclosure has a computer program product including a computer program stored in a computer-readable medium, and this computer program includes program code for executing the method shown in the flowchart. In such an embodiment, this computer program may be downloaded and installed from a network via the communication device 909, or installed from the storage device 908, or downloaded from the ROM 902. When this computer program is executed by the processing device 901, the above-described functions defined in the method of the embodiment of the present disclosure are executed.
[0120] In the present disclosure, the computer-readable medium described above may be a computer-readable signal medium, a computer-readable storage medium, or any combination thereof. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of the computer-readable storage medium include, but are not limited to, electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In the present disclosure, the computer-readable storage medium may be any tangible medium that includes or stores a program that can be used by or in combination with an instruction execution system, apparatus, or device. Also, in the present disclosure, the computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier carrying computer-readable program code. Such a propagated data signal can take various forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination of the foregoing. The computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that transmits, propagates, or conveys a program for use by or in combination with an instruction execution system, apparatus, or device. The program code included in the computer-readable medium can be transmitted using any suitable medium, including, but not limited to, wires, optical fiber cables, RF (radio frequency), or any suitable combination of the foregoing.
[0121] The computer-readable medium may be included in the electronic device or may be separate and not assembled into the electronic device.
[0122] The computer-readable medium stores one or more programs, and when the one or more programs are executed by the electronic device, the electronic device is caused to execute the method shown in the above embodiments.
[0123] The computer program code for executing the operations of the present disclosure can be written in one or more programming languages or combinations thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user computer, partially on the user computer, as a stand-alone software package, partially on the user computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user computer through any type of network including a Local Area Network (LAN) or a Wide Area Network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0124] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram may represent a module, a program segment, or a portion of code, which contains one or more executable instructions for implementing the specified logical function. Note that in some alternative implementations, the functions assigned to the boxes may occur in a different order than shown in the accompanying drawings. For example, two consecutive boxes may actually be executed substantially in parallel, or depending on the related functions, may be executed in reverse order. Also, it should be noted that each box in the block diagrams and / or flowcharts, as well as combinations of boxes in the block diagrams and / or flowcharts, may be implemented in a dedicated hardware-based system that performs a given function or operation, or may be implemented in a combination of dedicated hardware and computer instructions.
[0125] The units described in the description of the embodiments of the present disclosure may be implemented by software or by hardware. Here, the name of the unit does not constitute a limitation to the unit itself in a given situation. For example, the first acquisition unit may also be described as "a unit for acquiring at least two Internet protocol addresses".
[0126] The functions described above in this specification may be executed, at least in part, by one or more hardware logic components. For example, without limitation, usable and exemplary hardware logic components include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
[0127] In the context of the present disclosure, a machine-readable medium may be a tangible medium that includes or can store a program for use by or in conjunction with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium includes, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. More specific examples of the machine-readable storage medium include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0128] In a first aspect, according to one or more embodiments of the present disclosure, a video encoding method is provided, the method comprising: extracting video encoding features of initial video data by a hardware video editor, the video encoding features characterizing the complexity of the video content of the initial video data; processing the video encoding features of the initial video data by a pre-trained prediction neural network model to obtain a target quality rank corresponding to the initial video data, the target quality rank characterizing the image quality level when encoding the initial video in a constant quality variable bitrate mode; and encoding the initial video data at the target quality rank using the hardware video editor to generate a target video.
[0129] According to one or more embodiments of the present disclosure, extracting video encoding features of initial video data by a hardware video editor includes obtaining quality rank options of the initial video data, inputting video parameters of the initial video data and the quality rank options into the hardware video editor to obtain encoding parameters of the initial video data, and generating corresponding video encoding features based on the encoding parameters.
[0130] According to one or more embodiments of the present disclosure, the initial video data includes a plurality of video frames. Generating corresponding video encoding features based on the encoding parameters includes obtaining encoding parameters corresponding to each of the video frames, obtaining an encoding feature average value and an encoding feature variance value based on the encoding parameters corresponding to each of the video frames, where the encoding feature average value is an average value of the encoding parameters corresponding to each of the video frames, and the encoding feature variance value is a variance value of the encoding parameters corresponding to each of the video frames, and generating the video encoding features based on the video parameters, the quality rank options, and the encoding feature average value and the encoding feature variance value corresponding to each of the video frames.
[0131] According to one or more embodiments of the present disclosure, inputting the video parameters of the initial video data and the quality rank options into the hardware video editor to obtain the encoding parameters of the initial video data includes repeatedly executing the following steps until a predetermined condition is reached. The steps include obtaining the current frame of the initial video data, obtaining the video parameters of the current frame and the quality rank of the current frame, inputting the video parameters of the current frame and the quality rank of the current frame into the hardware video editor to obtain the encoding parameters corresponding to the current frame, and setting the next frame of the current frame as the new current frame.
[0132] According to one or more embodiments of the present disclosure, after obtaining the encoding parameters corresponding to the current frame, from the video parameters of the current frame, the quality rank of the current frame, and the encoding parameters, generating frame encoding features corresponding to the current frame; and obtaining frame encoding features of a previous frame corresponding to the current frame, where the previous frame is a first predetermined number of video frames adjacent to the current frame before the current frame, and the frame encoding features of the previous frame are generated based on a target quality rank corresponding to the previous frame. Further including, generating corresponding video encoding features based on the encoding parameters includes generating video encoding features corresponding to the current frame based on the frame encoding features corresponding to the current frame and the frame encoding features of the previous frame.
[0133] According to one or more embodiments of the present disclosure, the encoding parameters include at least one of frame type, frame size, image distortion degree, and image sharpness.
[0134] According to one or more embodiments of the present disclosure, the video encoding features of the initial video data include at least two encoding feature options, and each of the encoding feature options corresponds to a different quality rank. Processing the video encoding features of the initial video data by a pre-trained prediction neural network model, obtaining the target quality rank corresponding to the initial video data includes sequentially inputting each of the encoding feature options into the prediction neural network model to obtain a corresponding first evaluation value and a corresponding second evaluation value, where the first evaluation value characterizes a video quality evaluation value based on video multi-modal evaluation fusion, and the second evaluation value characterizes a video bit rate; and generating the target quality rank based on the corresponding first evaluation value and the corresponding second evaluation value for each of the encoding feature options.
[0135] According to one or more embodiments of the present disclosure, generating the target quality rank based on the first evaluation value corresponding to each of the encoded feature options and the corresponding second evaluation value includes obtaining a first target encoded feature based on the first evaluation value corresponding to each of the encoded feature options, where the first target encoded feature is an encoded feature option for which the first evaluation value is greater than a first threshold value, and determining a second target encoded feature based on a second evaluation value of the first target encoded feature, where the second target feature is a video encoding feature having the smallest second evaluation value among the first target encoded features, and obtaining the target quality rank based on a quality rank corresponding to the second target feature.
[0136] According to one or more embodiments of the present disclosure, before processing video encoding features of the initial video data by a pre-trained prediction neural network model and obtaining a target quality rank corresponding to the initial video data, obtaining original video data and a quality rank sequence, where the quality rank sequence includes at least two different quality ranks, sequentially processing the original video data using the hardware video editor based on the quality rank sequence to obtain video encoding features corresponding to each of the quality ranks, calculating a first evaluation value and the second evaluation value corresponding to each of the video encoding features, generating a training sample based on each of the video encoding features, the corresponding first evaluation value, and the corresponding second evaluation value, and training a predetermined neural network model based on the training sample to obtain the prediction neural network model.
[0137] In a second aspect, according to one or more embodiments of the present disclosure, a video encoding apparatus is provided, and the apparatus includes A parameter acquisition module for extracting video encoding features of initial video data by a hardware video editor, wherein the video encoding features include a parameter acquisition module for characterizing the complexity of the video content of the initial video data, A parameter optimization module for processing the video encoding features of the initial video data by a pre-trained prediction neural network model to obtain a target quality rank corresponding to the initial video data, wherein the target quality rank includes a parameter optimization module for characterizing the image quality level when encoding the initial video in a constant quality variable bitrate mode, An encoding module for encoding the initial video data with the target quality rank using the hardware video editor to generate a target video.
[0138] According to one or more embodiments of the present disclosure, specifically, the parameter acquisition module is used for obtaining quality rank options of the initial video data, inputting the video parameters and the quality rank options of the initial video data into the hardware video editor to obtain encoding parameters of the initial video data, and generating corresponding video encoding features based on the encoding parameters.
[0139] According to one or more embodiments of the present disclosure, the initial video data includes a plurality of video frames. When the parameter acquisition module generates corresponding video encoding features based on the encoding parameters, specifically, it acquires the encoding parameters corresponding to each of the video frames, and obtains an encoding feature average value and an encoding feature variance value based on the encoding parameters corresponding to each of the video frames. The encoding feature average value is the average value of the encoding parameters corresponding to each of the video frames, and the encoding feature variance value is the variance value of the encoding parameters corresponding to each of the video frames. It is used to generate the video encoding features based on the video parameters, the quality rank options, and the encoding feature average value and encoding feature variance value corresponding to each of the video frames.
[0140] According to one or more embodiments of the present disclosure, when the parameter acquisition module inputs the video parameters of the initial video data and the quality rank options into the hardware video editor to obtain the encoding parameters of the initial video data, specifically, it is used to repeatedly execute the following steps until a predetermined condition is reached. The steps include: obtaining the current frame of the initial video data; obtaining the video parameters of the current frame and the quality rank of the current frame; inputting the video parameters of the current frame and the quality rank of the current frame into the hardware video editor to obtain the encoding parameters corresponding to the current frame; and setting the next frame of the current frame as the new current frame.
[0141] According to one or more embodiments of the present disclosure, after obtaining the encoding parameters corresponding to the current frame, the parameter acquisition module generates frame encoding features corresponding to the current frame from the video parameters of the current frame, the quality rank of the current frame, and the encoding parameters, and obtains frame encoding features of a previous frame corresponding to the current frame, where the previous frame is the first predetermined number of video frames adjacent to the current frame before the current frame, and the frame encoding features of the previous frame are generated based on a target quality rank corresponding to the previous frame, and are further used for, when the parameter acquisition module generates corresponding video encoding features based on the encoding parameters, specifically generating video encoding features corresponding to the current frame based on the frame encoding features corresponding to the current frame and the frame encoding features of the previous frame.
[0142] According to one or more embodiments of the present disclosure, the encoding parameters include at least one of frame type, frame size, image distortion degree, and image sharpness.
[0143] According to one or more embodiments of the present disclosure, the video encoding features of the initial video data include at least two encoding feature options, each of the encoding feature options corresponding to a different quality rank, and the parameter optimization module specifically sequentially inputs each of the encoding feature options into the prediction neural network model to obtain a corresponding first evaluation value and a corresponding second evaluation value, where the first evaluation value characterizes a video quality evaluation value based on video multi-modal evaluation fusion, the second evaluation value characterizes a video bit rate, and is used for generating the target quality rank based on the first evaluation value corresponding to each of the encoding feature options and the second evaluation value corresponding thereto.
[0144] According to one or more embodiments of the present disclosure, when generating the target quality rank, the parameter optimization module is specifically based on the first evaluation value corresponding to each of the encoded feature options and the corresponding second evaluation value. Specifically, a first target encoded feature is obtained based on the first evaluation value corresponding to each of the encoded feature options, where the first target encoded feature is an encoded feature option for which the first evaluation value is greater than a first threshold, and a second target encoded feature is determined based on the second evaluation value of the first target encoded feature, where the second target feature is the video encoding feature with the smallest second evaluation value among the first target encoded features, and the target quality rank is obtained based on the quality rank corresponding to the second target feature. It is used for.
[0145] According to one or more embodiments of the present disclosure, before processing the video encoding features of the initial video data by a pre-trained prediction neural network model to obtain the target quality rank corresponding to the initial video data, the parameter optimization module obtains the original video data and a quality rank sequence, where the quality rank sequence includes at least two different quality ranks, and based on the quality rank sequence, the original video data is sequentially processed using the hardware video editor to obtain video encoding features corresponding to each of the quality ranks, and the first evaluation value and the second evaluation value corresponding to each of the video encoding features are calculated, and training samples are generated based on each of the video encoding features, the corresponding first evaluation value, and the corresponding second evaluation value, and a predetermined neural network model is trained based on the training samples to obtain the prediction neural network model. It is further used for.
[0146] In a third aspect, according to one or more embodiments of the present disclosure, an electronic device is provided, which includes a processor and a memory communicably connected to the processor. The memory stores computer-executable instructions, and the processor executes the computer-executable instructions stored in the memory to implement the video encoding method described in the first aspect and various possible designs of the first aspect.
[0147] In a fourth aspect, according to one or more embodiments of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, and when a processor executes the computer-executable instructions, the video encoding method described in the first aspect and various possible designs of the first aspect is implemented.
[0148] In a fifth aspect, embodiments of the present disclosure provide a computer program product including a computer program that, when executed by a processor, implements the video encoding method described in the first aspect and various possible designs of the first aspect.
[0149] The above description is merely an explanation of the preferred embodiments of the present disclosure and an explanation of the technical principles employed. The scope of the disclosure of the present disclosure is not limited to the technical solutions formed by specific combinations of the above technical features, but should be understood by those skilled in the art that other technical solutions formed by any combination of the above disclosed technical features or features equivalent thereto without departing from the above disclosed concept are also covered. For example, it includes technical solutions formed by replacing the above features with technical features (but not limited thereto) having similar functions disclosed in the present disclosure.
[0150] Furthermore, although the operations are depicted in a particular order, these operations should not be construed as requiring that they be performed in the particular order shown or sequentially. Multitasking and parallel processing may be advantageous in certain environments. Similarly, although some specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Some features described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, each feature described in the context of a single embodiment may also be implemented in multiple embodiments, individually or in any suitable sub-combination.
[0151] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely exemplary forms of implementing the claims.
Claims
1. Extracting video encoding features of initial video data by a hardware video editor, wherein the video encoding features characterize the complexity of the video content of the initial video data, and Processing the video encoding features of the initial video data by a pre-trained prediction neural network model to obtain a target quality rank corresponding to the initial video data, wherein the target quality rank characterizes the image quality level when encoding the initial video in a constant quality variable bitrate mode, and Encoding the initial video data with the target quality rank using the hardware video editor to generate a target video, and A video encoding method comprising the above steps.
2. The step of extracting video encoding features of initial video data by a hardware video editor includes: Obtaining quality rank options of the initial video data, and Inputting the video parameters of the initial video data and the quality rank options into the hardware video editor to obtain encoding parameters of the initial video data, and Generating corresponding video encoding features based on the encoding parameters, and The method according to claim 1, comprising the above steps.
3. The initial video data includes a plurality of video frames, and The step of generating corresponding video encoding features based on the encoding parameters includes: Obtaining encoding parameters corresponding to each of the video frames, and Obtaining an encoding feature average value and an encoding feature variance value based on the encoding parameters corresponding to each of the video frames, wherein the encoding feature average value is the average value of the encoding parameters corresponding to each of the video frames, and the encoding feature variance value is the variance value of the encoding parameters corresponding to each of the video frames, and Generating the video encoding features based on the video parameters, the quality rank options, and the encoding feature average value and encoding feature variance value corresponding to each of the video frames, and The method according to claim 2, comprising the above steps.
4. Inputting the video parameters of the initial video data and the quality rank options into the hardware video editor to obtain encoding parameters of the initial video data includes repeatedly executing the following steps until a predetermined condition is met: The said steps are obtaining the current frame of the said initial video data; obtaining the video parameters of the said current frame and the quality rank of the said current frame; inputting the video parameters of the said current frame and the quality rank of the said current frame into the said hardware video editor to obtain the encoding parameters corresponding to the said current frame; setting the next frame of the said current frame as the new current frame; The method according to claim 2.
5. After obtaining the encoding parameters corresponding to the said current frame, generating frame encoding features corresponding to the said current frame from the video parameters of the said current frame, the quality rank of the said current frame and the said encoding parameters; obtaining the frame encoding features of the previous frame corresponding to the said current frame, where the previous frame is the first predetermined number of video frames adjacent to the front of the said current frame, and the frame encoding features of the previous frame are generated based on the target quality rank corresponding to the said previous frame, further including generating corresponding video encoding features based on the said encoding parameters means generating video encoding features corresponding to the said current frame based on the frame encoding features corresponding to the said current frame and the frame encoding features of the previous frame; The method according to claim 4, including.
6. The said encoding parameters include at least one of frame type, frame size, image distortion degree, and image sharpness. The method according to claim 2.
7. The video encoding features of the said initial video data include at least two encoding feature options, and each of the said encoding feature options corresponds to a different quality rank. Processing the video encoding features of the said initial video data by a pre-trained prediction neural network model, and obtaining the target quality rank corresponding to the said initial video data means sequentially inputting each of the said encoding feature options into the said prediction neural network model to obtain a corresponding first evaluation value and a corresponding second evaluation value, where the first evaluation value characterizes a video quality evaluation value based on video multi-method evaluation fusion, and the second evaluation value characterizes a video bit rate. generating the target quality rank based on the first evaluation value corresponding to each of the encoded feature options and the corresponding second evaluation value; The method according to claim 1, comprising: **Claim 8** Generating the target quality rank based on the first evaluation value corresponding to each of the encoded feature options and the corresponding second evaluation value comprises: obtaining a first target encoded feature based on the first evaluation value corresponding to each of the encoded feature options, wherein the first target encoded feature is an encoded feature option for which the first evaluation value is greater than a first threshold; determining a second target encoded feature based on the second evaluation value of the first target encoded feature, wherein the second target feature is the video encoded feature among the first target encoded features for which the second evaluation value is the smallest; obtaining the target quality rank based on the quality rank corresponding to the second target feature; The method according to claim 7, comprising: **Claim 9** Before processing the video encoded features of the initial video data by a pre-trained predictive neural network model to obtain the target quality rank corresponding to the initial video data, obtaining original video data and a quality rank sequence, wherein the quality rank sequence includes at least two different quality ranks; sequentially processing the original video data using the hardware video editor based on the quality rank sequence to obtain video encoded features corresponding to each of the quality ranks; calculating a first evaluation value and the second evaluation value corresponding to each of the video encoded features; generating a training sample based on each of the video encoded features, the corresponding first evaluation value, and the corresponding second evaluation value, and training a predetermined neural network model based on the training sample to obtain the predictive neural network model; The method according to claim 7, comprising: **Claim 10** A parameter acquisition module for extracting video encoded features of initial video data by a hardware video editor, wherein the video encoded features are parameter acquisition modules that characterize the complexity of the video content of the initial video data; A parameter optimization module for processing video encoding features of the initial video data by a pre-trained prediction neural network model to obtain a target quality rank corresponding to the initial video data, wherein the target quality rank is a parameter optimization module for characterizing a picture quality level when encoding the initial video in a constant quality variable bit rate mode, An encoding module for encoding the initial video data with the target quality rank using the hardware video editor to generate a target video, A video encoding apparatus comprising the same.
11. An electronic device comprising a processor and a memory communicably connected to the processor, The memory stores computer-executable instructions, The processor executes the computer-executable instructions stored in the memory to implement the video encoding method according to any one of claims 1 to 9.
12. A computer-readable storage medium storing computer-executable instructions, and when the processor executes the computer-executable instructions, implementing the video encoding method according to any one of claims 1 to 9.
13. A computer program product including a computer program that, when executed by a processor, implements the video encoding method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Training rate control neural networks through reinforcement learning
WO2022248736A1
Methods, systems, and media for determining perceptual quality indicators of video content items
WO2022261203A1