A method for determining quantization parameters and related devices
By obtaining the feature vector and bit allocation information of the video frame, and using the training model to predict the correction coefficient, the problem of inaccurate quantization parameters is solved, and more efficient video encoding performance and adaptability are achieved.
Patent Information
- Application Number
- CN202210883402.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-26
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-07-26
AI Technical Summary
In the prior art, the accuracy of the quantization parameters during video encoding is insufficient, resulting in a degradation of encoding performance, especially when the scene is transformed, which leads to an increase in the number of bits.
By obtaining the content attributes of the frame to be encoded, determining the target feature vector and bit allocation information, using the trained coefficient prediction model to predict the target correction coefficient, and calculating the target quantization parameters in combination with the target encoding bits, realizing the encoding of the frame to be encoded.
It improves the accuracy and encoding performance of quantization parameters, can adapt to different images, and improves encoding efficiency and video quality.
Smart Images

Figure CN115243042B_ABST
Abstract
Description
Background Art
[0002] Rate control is an important part of video coding. It can optimize the allocation of the remaining bits according to the state of the channel and dynamically adjust the quantization parameter according to the current situation of the encoder, so as to provide the optimal video quality under the condition of limited bandwidth.
[0003] In related technologies, the following method is usually adopted to determine the quantization parameter corresponding to the target frame: First, according to the frame type of the target frame and the block level it is in, the target correction coefficient corresponding to the target frame is determined from each candidate correction coefficient. Then, according to the target allocated bits and the target correction coefficient allocated to the target frame, the quantization parameter corresponding to the target frame is calculated. After that, the calculated quantization parameter can also be used to encode the target frame, and according to the encoding result, the candidate correction coefficient is updated.
[0004] However, on the one hand, since the target correction coefficient is fixed before the start of encoding and cannot be adaptive to different images, the calculated quantization parameter is not accurate enough. On the other hand, during the encoding process, although the candidate correction coefficient is updated according to the encoding result, there is a certain lag in the coefficient update and it cannot be updated in advance when encountering scene changes, resulting in inaccurate calculated quantization parameters, and further leading to a decline in the overall encoding performance. Under the same video quality, the number of consumed bits increases. Summary of the Invention
[0005] The embodiments of the present application provide a method for determining quantization parameters and related devices to improve the accuracy of quantization parameters and enhance the encoding performance.
[0006] In a first aspect, the embodiments of the present application provide a method for determining quantization parameters, including:
[0007] Obtain a frame to be encoded, and based on the content attributes of the frame to be encoded, obtain the target feature vector corresponding to the frame to be encoded, and obtain the bit allocation information corresponding to the frame to be encoded;
[0008] Based on the bit allocation information and the specified code rate, determine the corresponding target encoding bits;
[0009] Input the target feature vector into a trained coefficient prediction model to obtain the target correction coefficient corresponding to the frame to be encoded, where the coefficient prediction model is trained based on a training data set, and each training data contains a sample frame and a corresponding reference correction coefficient, and the reference correction coefficient is determined after multiple rounds of encoding for the one sample frame;
[0010] Based on the target correction coefficient and the target coding bits, obtain a target quantization parameter, and based on the target quantization parameter, encode the frame to be encoded to obtain a target encoded frame.
[0011] In a second aspect, an embodiment of the present application provides a quantization parameter determination device, including:
[0012] An acquisition unit, configured to acquire a frame to be encoded, and based on the content attribute of the frame to be encoded, obtain a target feature vector corresponding to the frame to be encoded, and obtain bit allocation information corresponding to the frame to be encoded;
[0013] A bit allocation unit, configured to determine corresponding target coding bits based on the bit allocation information and a specified code rate;
[0014] A coefficient determination unit, configured to input the target feature vector into a trained coefficient prediction model to obtain a target correction coefficient corresponding to the frame to be encoded, where the coefficient prediction model is trained based on a training data set, and each training data includes a sample frame and a corresponding reference correction coefficient, and the reference correction coefficient is determined after multiple rounds of encoding for the one sample frame;
[0015] An encoding unit, configured to obtain a target quantization parameter based on the target correction coefficient and the target coding bits, and encode the frame to be encoded based on the quantization parameter to obtain a target encoded frame.
[0016] As a possible implementation manner, when determining a target offset from the multiple candidate offsets based on the multiple encoding results, the training unit is specifically configured to:
[0017] Based on the multiple encoding results, determine the encoding error corresponding to each of the multiple candidate offsets;
[0018] From the multiple candidate offsets, determine a candidate offset whose corresponding encoding error meets a preset error condition, and use the determined candidate offset as the target offset.
[0019] As a possible implementation manner, when obtaining a target feature vector corresponding to the frame to be encoded based on the content attribute of the frame to be encoded, the training unit is specifically configured to:
[0020] Divide the frame to be encoded into N non-overlapping blocks according to a preset block size, where N is a positive integer;
[0021] Perform intra-frame prediction on the N blocks respectively, and based on the intra-frame prediction results, obtain the target intra-frame prediction cost corresponding to each of the N blocks;
[0022] Perform inter-frame prediction on the N blocks respectively, and based on the inter-frame prediction results, obtain the target inter-frame prediction costs corresponding to the N blocks respectively;
[0023] Based on the inter-frame prediction costs and intra-frame prediction costs corresponding to the N blocks respectively, obtain multiple feature attributes;
[0024] Based on the multiple feature attributes, obtain the target feature vector corresponding to the frame to be encoded.
[0025] As a possible implementation manner, the training unit is specifically configured to:
[0026] Divide the training data set into a training set, a validation set, and a test set;
[0027] Based on the training set, train the coefficient prediction model to obtain a first error metric;
[0028] Based on the validation set, train the coefficient prediction model to obtain a second error metric;
[0029] If the second error metric is greater than the first error metric, then based on the second error metric, adjust the model parameters of the coefficient prediction model, and based on the validation set, train the adjusted coefficient prediction model until the second error metric is not greater than the first error metric;
[0030] Input the test set into the coefficient prediction model to obtain a third error metric, and when the third error metric is less than a preset threshold, output the trained coefficient prediction model.
[0031] In a third aspect, an embodiment of the present application provides an electronic device, including a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor is caused to execute the steps of the above-mentioned quantization parameter determination method.
[0032] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which includes a computer program, and when the computer program runs on an electronic device, the computer program is used to cause the electronic device to execute the steps of the above-mentioned quantization parameter determination method.
[0033] In a fifth aspect, an embodiment of the present application provides a computer program product, the program product includes a computer program, the computer program is stored in a computer-readable storage medium, and a processor of an electronic device reads and executes the computer program from the computer-readable storage medium, so that the electronic device executes the steps of the above-mentioned quantization parameter determination method.
[0034] In the embodiments of the present application, based on the content attributes of the frame to be encoded, the target feature vector corresponding to the frame to be encoded is obtained, and the bit allocation information corresponding to the frame to be encoded is obtained. Based on the bit allocation information and the specified code rate, the corresponding target encoding bits are determined; the target feature vector is input into the trained coefficient prediction model to obtain the target correction coefficient corresponding to the frame to be encoded; based on the target correction coefficient and the target encoding bits, the target quantization parameter is obtained, and based on the quantization parameter, the frame to be encoded is encoded to obtain the target encoded frame.
[0035] In this way, through the machine learning algorithm, the problem of lag of the correction coefficient is solved, thereby improving the accuracy of the quantization parameter and enhancing the encoding performance. At the same time, it can be adapted to different images, and through more accurate correction coefficients, the accuracy of the quantization parameter is further improved. In addition, by performing multiple rounds of encoding on the samples to obtain training data, the accuracy of the model labels can be improved, thereby enhancing the model prediction accuracy.
[0036] Other features and advantages of the present application will be described in the following specification, and part of them will be obvious from the specification, or will be understood by implementing the present application. The objectives and other advantages of the present application can be achieved and obtained through the structures specifically pointed out in the written specification, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation to the present application. In the drawings:
[0038] Figure 1 It is a schematic diagram of an application scenario provided in the embodiments of the present application;
[0039] Figure 2 It is a schematic flowchart of a method for determining quantization parameters provided in the embodiments of the present application;
[0040] Figure 3 It is a schematic flowchart of a method for determining the target feature vector provided in the embodiments of the present application;
[0041] Figure 4 It is a schematic flowchart of a method for obtaining a training data set provided in the embodiments of the present application;
[0042] Figure 5 It is a schematic flowchart of a method for determining quantization parameters provided in the embodiments of the present application;
[0043] Figure 6 It is a schematic flowchart of a method for determining the target quantization parameter offset provided in the embodiments of the present application;
[0044] Figure 7 Schematic diagram of a candidate offset provided in an embodiment of the present application;
[0045] Figure 8 Schematic flowchart of a model training method provided in an embodiment of the present application;
[0046] Figure 9A Schematic diagram of a sequence to be processed provided in an embodiment of the present application;
[0047] Figure 9B Schematic diagram of a level provided in an embodiment of the present application;
[0048] Figure 10 Another schematic flowchart of obtaining a training data set provided in an embodiment of the present application;
[0049] Figure 11 Another schematic flowchart of determining a target feature vector provided in an embodiment of the present application;
[0050] Figure 12 Schematic diagram of the structure of a quantization parameter determination device provided in an embodiment of the present application;
[0051] Figure 13 Schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Detailed implementation manners
[0052] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments recorded in this application document without creative efforts shall fall within the scope of the technical solutions protected by the present application.
[0053] First, the terms related to the present application are explained:
[0054] Bitrate control: In the video encoding process, a technique of controlling the quantization parameter so that the output bitrate of the output video is equal to the given bitrate.
[0055] Quantization parameter: An important parameter in the video encoding process, which affects the encoding bit count and encoding quality.
[0056] Bitrate: In the video transmission process, the number of bits transmitted per second. For example, the bitrate can be the total number of bits of the video divided by the video duration.
[0057] Bit-Quantization Parameter Model: According to statistical laws, it is the relationship between the number of encoded bits and the quantization parameter. In this application, the bit-quantization parameter model can be expressed by formula (1). In formula (1), E is a constant, which takes the value of 2000000 when the frame type of the current frame is KEY_FRAME, and 1500000 otherwise. R tar is the target number of bits, q is the quantization parameter, and c is the correction factor:
[0058]
[0059] Frame: An image.
[0060] Group of pictures: A set composed of several adjacent frames in the display order.
[0061] Key frame (KEY_FRAME): A frame encoded entirely in the intra mode.
[0062] AOM-AV1: The reference software of the AV1 standard.
[0063] Coding performance: The efficiency of video compression. Under the same quality, the less the number of bits occupied, the higher the performance; under the same number of bits, the higher the quality, the higher the performance. The coding performance can be specifically measured by the Bitrate-Distortion Ratio (BDBR). Its physical meaning is the saving ratio of the bitrate under the same quality. A negative value indicates performance improvement.
[0064] Bitstream: A string of binary characters used to represent the compressed video.
[0065] Intra prediction: A prediction mode that uses the pixels above and to the left of the current block to predict the current block.
[0066] Inter prediction: A prediction mode that searches for the block (reference block) most similar to the current block in the reference frame to predict the current block.
[0067] Motion vector: A vector (x, y), where x and y respectively represent the number of pixels by which the current block is offset from the reference block in the horizontal and vertical directions.
[0068] Cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or local area network to achieve data computing, storage, processing, and sharing.
[0069] Cloud technology is a general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model. It can form a resource pool, which can be used on demand and is flexible and convenient. Cloud computing technology will become an important support. The backend services of the technical network system require a large amount of computing and storage resources, such as video websites, picture websites and more portal websites. With the rapid development and application of the Internet industry, in the future, each item may have its own identification mark, which needs to be transmitted to the backend system for logical processing. Data of different levels will be processed separately. All kinds of industry data need strong system backing support, which can only be achieved through cloud computing.
[0070] Cloud computing is a computing model that distributes computing tasks across a resource pool consisting of a large number of computers, enabling various application systems to obtain computing power, storage space, and information services as needed. The network that provides resources is called a "cloud". The resources in the "cloud" are infinitely scalable to users, and can be obtained at any time, used on demand, expanded at any time, and paid for by use.
[0071] As a provider of basic cloud computing capabilities, a cloud computing resource pool (referred to as a cloud platform, generally referred to as an Infrastructure as a Service (IaaS) platform) will be established, and various types of virtual resources will be deployed in the resource pool for external customers to choose to use. The cloud computing resource pool mainly includes: computing devices (virtualized machines, including operating systems), storage devices, and network devices.
[0072] According to the logical function division, the Platform as a Service (PaaS) layer can be deployed on the Infrastructure as a Service (IaaS) layer, and the SaaS (Software as a Service) layer can be deployed on the PaaS layer. SaaS can also be deployed directly on IaaS. PaaS is a platform for software operation, such as databases, web containers, etc. SaaS is a variety of business software, such as web portals, SMS mass senders, etc. Generally speaking, SaaS and PaaS are upper layers relative to IaaS.
[0073] Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making.
[0074] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0075] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning. The solution provided in the embodiments of this application relates to the machine learning technology of artificial intelligence. Through the machine learning technology, a coefficient prediction model is constructed, and then the coefficient prediction model is used to predict the correction coefficient.
[0076] Rate control is an important part of video coding. It can optimize the allocation of the remaining bits according to the state of the channel and dynamically adjust the quantization parameters according to the current situation of the encoder, so as to provide the best video quality under the condition of limited bandwidth.
[0077] Rate control in AV1 usually includes stages such as pre-analysis, bit allocation, quantization parameter calculation, and encoding. Among them, the quantization parameter calculation stage is usually implemented in the following way:
[0078] First, according to the frame type of the target frame and the block level it is in, determine the target correction coefficient corresponding to the target frame from each candidate correction coefficient. Then, according to the target allocation bits and the target correction coefficient assigned to the target frame, use the above formula (1) to calculate the quantization parameter corresponding to the target frame. After that, use the calculated quantization parameter to encode the target frame. After encoding, update the candidate correction coefficients according to the encoding result.
[0079] However, on the one hand, since the correction coefficient is fixed before the start of encoding and cannot be adapted to different images, the calculated quantization parameter is not accurate enough. On the other hand, during the encoding process, although the candidate correction coefficients are updated according to the encoding result, there is a certain lag in coefficient update and it cannot be updated in advance when encountering scene changes, resulting in inaccurate calculated quantization parameters, and further leading to a decline in the overall encoding performance. Under the same video quality, the number of consumed bits increases.
[0080] In the embodiments of the present application, based on the content attributes of the frame to be encoded, obtain the target feature vector corresponding to the frame to be encoded, and obtain the bit allocation information corresponding to the frame to be encoded. Based on the bit allocation information and the specified bit rate, determine the corresponding target encoding bits; input the target feature vector into the trained coefficient prediction model to obtain the target correction coefficient corresponding to the frame to be encoded; based on the target correction coefficient and the target encoding bits, obtain the target quantization parameter, and based on the quantization parameter, encode the frame to be encoded to obtain the target encoded frame. In this way, through the machine learning algorithm, the problem of lag in the correction coefficient is solved, thereby improving the accuracy of the quantization parameter and enhancing the encoding performance. At the same time, it can be adapted to different images, and through more accurate correction coefficients, the accuracy of the quantization parameter is further improved.
[0081] Refer to Figure 1 As shown, it is a schematic diagram of an application scenario provided in the embodiments of the present application. This application scenario includes at least a terminal device 110 and a server 120. The number of terminal devices 110 can be one or more, and the number of servers 120 can also be one or more. The present application does not make specific limitations on the number of terminal devices 110 and servers 120.
[0082] In the embodiments of the present application, the terminal device 110 can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, an Internet of Things device, a smart home appliance, a vehicle-mounted terminal, etc., but is not limited thereto.
[0083] The server 120 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The terminal device 110 and the server 120 can be directly or indirectly connected through wired or wireless communication methods, and this application does not limit this here.
[0084] Exemplarily, a client corresponding to a video-related application is installed in the terminal device 110, and the video-related applications include but are not limited to conference applications, video applications, live broadcast applications, etc.
[0085] It should be noted that the quantization parameter determination method mentioned in this application can be applied to the terminal device 110, can also be applied to the server 120, or can be jointly executed by the terminal device 110 and the server 120. Taking the application to the server 120 as an example only, the server 120 obtains a frame to be encoded. The frame to be encoded can be any frame in the video to be encoded. The video to be encoded can be a video such as a live broadcast, a conference, or a TV drama, but is not limited thereto. Based on the content attributes of the frame to be encoded, a target feature vector corresponding to the frame to be encoded is obtained, and bit allocation information corresponding to the frame to be encoded is obtained. Based on the bit allocation information and the specified code rate, the corresponding target encoding bits are determined. Then, the target feature vector is input into the trained coefficient prediction model to obtain the target correction coefficient corresponding to the frame to be encoded, and based on the target correction coefficient and the target encoding bits, the target quantization parameter is obtained. After that, based on the quantization parameter, the frame to be encoded is encoded to obtain the target encoded frame, and then the target frame is sent to the terminal device 110.
[0086] Refer to Figure 2 As shown, it is a schematic flowchart of a quantization parameter determination method provided in an embodiment of this application. This method is applied to an electronic device, and the electronic device can be a terminal device or a server. The specific process of this method is as follows:
[0087] S201. Obtain a frame to be encoded, and based on the content attributes of the frame to be encoded, obtain a target feature vector corresponding to the frame to be encoded, and obtain bit allocation information corresponding to the frame to be encoded.
[0088] In the embodiment of this application, the frame to be encoded can be any frame in the video to be encoded.
[0089] Specifically, refer to Figure 3 As shown, based on the content attributes of the frame to be encoded, obtaining the target feature vector corresponding to the frame to be encoded can adopt but is not limited to the following steps:
[0090] S301. Divide the frame to be encoded into N non - overlapping blocks according to a preset block size, where N is a positive integer.
[0091] The frame to be encoded can be divided into one or more largest coding units. The largest coding unit is usually 128×128. In actual application, each largest coding unit can be further divided into one or more coding units through depth division. Among them, the size of the coding unit is the preset block size.
[0092] The size of the coding unit exists but is not limited to the following cases: 4×4, 4×8, 8×4, 8×8, 8×16, 16×8, 16×16, 16×32, 32×16, 32×32, 32×64, 64×32, 64×128, 128×64, 128×128, 4×16, 16×4, 8×32, 32×8, 16×64, 64×16. For example, if the preset block size is 16×16 and the size of the frame to be encoded is 128×128, the frame to be encoded is divided into 64 blocks, and each block is 16×16.
[0093] S302. Perform intra - frame prediction on the N blocks respectively, and based on the intra - frame prediction results, obtain the target intra - frame prediction costs corresponding to the N blocks respectively.
[0094] Taking block x as an example, where block x is any one of the N blocks. Specifically, for each intra - frame prediction mode, perform prediction on block x respectively to obtain the intra - frame prediction results, and calculate the intra - frame prediction costs corresponding to each intra - frame prediction mode according to the intra - frame prediction results. Then, among the calculated intra - frame prediction costs, the intra - frame prediction mode corresponding to the intra - frame prediction cost that meets the preset cost condition is used as the target intra - frame prediction mode of block x, and the intra - frame prediction cost corresponding to the target intra - frame prediction mode is used as the target intra - frame prediction cost of block x. Among them, the preset cost condition can be that the cost value is the smallest, but is not limited to this. The intra - frame prediction cost is used to characterize the difference between the original value of block x and the intra - frame prediction result, and the intra - frame prediction result can also be called the predicted pixel value.
[0095] In the embodiments of the present application, the coding strategy can include intra - frame prediction mode and inter - frame prediction mode. Taking AV1 as an example, the intra - frame prediction mode includes a directional prediction mode and a non - directional prediction mode.
[0096] The direction prediction mode may specifically include: angle prediction mode 1 (for example, V_PRED mode, used for prediction in the vertical direction), angle prediction mode 2 (for example, H_PRED, used for prediction in the horizontal direction), angle prediction mode 3 (for example, D45_PRED, used for prediction in the 45-degree angle direction), angle prediction mode 4 (for example, D135_PRED, used for prediction in the 135-degree angle direction), angle prediction mode 5 (for example, D113_PRED, used for prediction in the 113-degree angle direction), angle prediction mode 6 (for example, D157_PRED mode, used for prediction in the 157-degree angle direction), angle prediction mode 7 (for example, D203_PRED mode, used for prediction in the 203-degree angle direction), and angle prediction mode 8 (for example, D67_PRED mode, used for prediction in the 67-degree angle direction). Among them, each angle includes 6 angle offsets, which are plus or minus 3 degrees, plus or minus 6 degrees, and plus or minus 9 degrees respectively.
[0097] There may be multiple non-direction prediction modes. For example, intra-frame prediction mode 1 (for example, DC_PRED mode, applicable to large flat areas, and predicted based on the average value of the left and / or upper reference pixels), intra-frame prediction mode 2 (for example, SMOOTH_PRED mode, used for prediction using quadratic interpolation in the horizontal and vertical directions), intra-frame prediction mode 3 (for example, SMOOTH_V_PRED mode, used for prediction using quadratic interpolation in the vertical direction), intra-frame prediction mode 4 (for example, SMOOTH_H_PRED mode, used for prediction using quadratic interpolation in the horizontal direction), intra-frame prediction mode 5 (PAETH_PRED, used for prediction in the direction of the minimum gradient). In addition, the intra-frame prediction mode may also include a palette prediction mode and an intra-frame block copy prediction mode.
[0098] It should be noted that during the intra-frame prediction process, the video encoder provided by the embodiments of the present application further upgrades the granularity of direction prediction, incorporates gradients and correlations into non-directional prediction, and fully utilizes the luminance consistency and chrominance signals.
[0099] Exemplarily, for any intra-frame prediction mode, the intra-frame prediction cost can be calculated using the following formula (2):
[0100] intra_error = ∑ i (intra_ori i - intra_pred i ) Formula (2)
[0101] Wherein, intra_error represents the intra-frame prediction cost, and intra_ori iRepresents the original value of the i-th pixel in block x, intra_pred i Represents the predicted value of the i-th pixel in block x obtained by using the intra prediction mode. In this article, the original value represents the original pixel value, and the predicted value represents the predicted pixel value.
[0102] Taking the angular prediction mode 1 in the intra prediction mode as an example, for pixel P in block x, according to the prediction angle in the vertical direction, determine the position of the reference pixel from the already encoded pixel row above block x, and then use the value of the reference pixel as the predicted value of pixel P. Furthermore, based on the element value of pixel P and the predicted value of pixel P, obtain the intra prediction cost of pixel P. In this way, for the angular prediction mode 1, the intra prediction cost corresponding to each pixel in block x can be obtained respectively, and then based on the intra prediction cost corresponding to each pixel respectively, obtain the intra prediction cost corresponding to block x.
[0103] Assume that there are 10 intra prediction modes. For the 10 intra prediction modes, perform predictions on block x respectively to obtain the corresponding intra prediction results. According to the intra prediction results, calculate the intra prediction costs corresponding to the 10 intra prediction modes as: 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1. Then, take the intra prediction cost 0.1 as the target intra prediction cost corresponding to block x.
[0104] S303. Perform inter predictions on N blocks respectively, and based on the inter prediction results, obtain the target inter prediction costs corresponding to the N blocks respectively.
[0105] Inter prediction mainly utilizes the correlation in the video time domain, uses the pixels in other adjacent already encoded images to predict the pixels in the current image, so as to effectively remove the redundancy in the video time domain and can effectively save the bits of the encoded residual data.
[0106] The inter prediction mode can include the single reference frame mode and the combined reference frame mode. Among them, the single reference frame mode can include the motion estimation mode and the non-motion estimation mode.
[0107] The motion estimation mode can be the NEWMV mode. The NEWMV mode needs to transmit the residual (Motion Vector Difference, MVD) between the unit to be encoded and the prediction unit during the encoding process.
[0108] The non-motion estimation modes may include non-motion estimation mode 1 (e.g., NEARESTMV mode), non-motion estimation mode 2 (e.g., NEARMV mode), and non-motion estimation mode 3 (e.g., GLOBALMV mode). Among them, the motion vectors (MVs) of the prediction blocks in the NEARESTMV mode and the NEARMV mode are derived based on the surrounding block information, and there is no need to transmit residuals; the motion vectors of the prediction blocks in the GLOBALMV mode need to be derived based on global motion.
[0109] The combined reference frame modes may include combined reference frame mode 1 (e.g., NEAREST_NEARESTMV mode), combined reference frame mode 2 (e.g., NEAR_NEARMV mode), combined reference frame mode 3 (e.g., NEAREST_NEWMV mode), combined reference frame mode 4 (e.g., NEW_NEARESTMV mode), combined reference frame mode 5 (e.g., NEAR_NEWMV mode), combined reference frame mode 6 (e.g., NEW_NEARMV mode), combined reference frame mode 7 (e.g., GLOBAL_GLOBALMV mode), combined reference frame mode 8 (e.g., NEW_NEWMV mode).
[0110] In the embodiments of the present application, the inter-frame prediction cost is used to characterize the difference between the original value of block x and the inter-frame prediction result, and the target inter-frame prediction cost includes one or more of the first target inter-frame prediction cost or the second target inter-frame prediction cost.
[0111] Specifically, the first target inter-frame prediction cost can be obtained in the following manner:
[0112] Using the LAST_FRAME of the frame to be encoded as the target reference frame, perform inter-frame prediction on block x. LAST_FRAME represents the reference frame in the video to be encoded that is closest to the frame to be encoded and has a frame number less than that of the frame to be encoded. Specifically, for each inter-frame prediction mode, perform prediction on block x respectively to obtain the corresponding inter-frame prediction results, and calculate the first inter-frame prediction cost corresponding to each inter-frame prediction mode according to the inter-frame prediction results. Then, among the calculated first inter-frame prediction costs, the inter-frame prediction mode corresponding to the smallest first inter-frame prediction cost is used as the first target inter-frame prediction mode of block x, and the first inter-frame prediction cost corresponding to the first target inter-frame prediction mode is used as the first target inter-frame prediction cost corresponding to block x.
[0113] It should be noted that the first inter-frame prediction cost is used to characterize the difference between the original value of block x and the inter-frame prediction result when performing inter-frame prediction on block x using the LAST_FRAME of the frame to be encoded as the reference frame.
[0114] After determining the first target inter-frame prediction mode corresponding to block x, the motion vector mv corresponding to block x is also recorded last , mv last represents the motion vector corresponding to block x in the first target inter-frame prediction mode, that is, the number of pixels by which block x is offset in the horizontal and vertical directions compared to the reference frame in the first target inter-frame prediction mode.
[0115] Exemplarily, for each inter-frame prediction mode, the first inter-frame prediction cost can be calculated using the following formula (3):
[0116] coded_error = ∑ i (coded_ori i - coded_pred i ) Formula (3)
[0117] where coded_error represents the first inter-frame prediction cost, coded_ori i represents the original value of the i-th pixel in block x, and coded_pred i represents the predicted value of the i-th pixel in block x obtained by the inter-frame prediction mode with LAST_FRAME as the reference frame. Specifically, the second target inter-frame prediction cost can be obtained in the following manner:
[0118] Using the GOLDEN_FRAME of the frame to be encoded as the target reference frame, perform inter-frame prediction on block x. GOLDEN_FRAME represents an I-frame or a Generalized P and B picture (GPB) with a frame number smaller than the I-frame corresponding to the frame to be encoded. Specifically, for each inter-frame prediction mode, perform prediction on block x separately, calculate the second inter-frame prediction cost corresponding to each inter-frame prediction mode according to the inter-frame prediction result, and then, among the calculated second inter-frame prediction costs, take the inter-frame prediction mode corresponding to the second inter-frame prediction cost with the smallest value as the second target inter-frame prediction mode of block x, and take the second inter-frame prediction cost corresponding to the second target inter-frame prediction mode as the second target inter-frame prediction cost of block x.
[0119] It should be noted that the second inter-frame prediction cost is used to characterize the difference between the original value of block x and the inter-frame prediction result when performing inter-frame prediction on block x with the GOLDEN_FRAME of the frame to be encoded as the reference frame.
[0120] After determining the second target inter-frame prediction mode corresponding to block x, the motion vector mv corresponding to block x is also recorded golden , mv goldenDenote the motion vector corresponding to block x in the second target inter-frame prediction mode, that is, the number of pixel offsets in the horizontal and vertical directions of block x in the first target inter-frame prediction mode compared to the reference frame.
[0121] Exemplarily, for each inter-frame prediction mode, the second inter-frame prediction cost can be calculated using the following formula (4):
[0122] sr_coded_error = ∑ i (sr_coded_ori i -sr_coded_pred i ) Formula (4)
[0123] where sr_coded_error represents the second inter-frame prediction cost, sr_coded_ori i is the original value of the i-th pixel point in block x, and sr_coded_pred i represents the predicted value of the i-th pixel point in block x obtained by the inter-frame prediction mode when using GOLDEN_FRAME as the reference frame.
[0124] It should be noted that in the implementation of this application, there are 7 reference frames in each of the 4 single-reference-frame prediction modes: LAST_FRAME, LAST2_FRAME, LAST3_FRAME, GOLDEN_FRAME, BWDREF_FRAME, ALTREF2_FRAME, and ALTREF_FRAME respectively. Among them, LAST2_FRAME represents the reference frame that is the second closest to the frame to be encoded and has a frame number smaller than the frame to be encoded (forward reference), LAST3_FRAME represents the reference frame that is the third closest to the frame to be encoded and has a frame number smaller than the frame to be encoded (forward reference), BWDREF_FRAME represents the reference frame that is the closest to the frame to be encoded and has a frame number larger than the frame to be encoded (backward reference), ALTREF2_FRAME represents the reference frame that is the second closest to the frame to be encoded and has a frame number larger than the frame to be encoded (backward reference), and ALTREF_FRAME represents the reference frame that is the third closest to the frame to be encoded and has a frame number larger than the frame to be encoded (backward reference).
[0125] Under each of the 8 combined reference frame prediction modes, there are 16 reference frame combinations, namely {LAST_FRAME, ALTREF_FRAME}, {LAST2_FRAME, ALTREF_FRAME}, {LAST3_FRAME, ALTREF_FRAME}, {GOLDEN_FRAME, ALTREF_FRAME}, {LAST_FRAME, BWDREF_FRAME}, {LAST2_FRAME, BWDREF_FRAME}, {LAST3_FRAME, BWDREF_FRAME}, {GOLDEN_FRAME, BWDREF_FRAME}, {LAST_FRAME, ALTREF2_FRAME}, {LAST2_FRAME, ALTREF2_FRAME}, {LAST3_FRAME, ALTREF2_FRAME}, {GOLDEN_FRAME, ALTREF2_FRAME}, {LAST_FRAME, LAST2_FRAME}, {LAST_FRAME, LAST3_FRAME}, {LAST_FRAME, GOLDEN_FRAME}, {BWDREF_FRAME, ALTREF_FRAME}.
[0126] In the above text, only LAST_FRAME and GOLDEN_FRAME are taken as examples of the target reference frames for illustration. In the actual application process, any one of the above reference frames or reference frame combinations can also be used for the first target inter-frame prediction cost or the second target inter-frame prediction cost.
[0127] S304. Obtain multiple feature attributes based on the target inter-frame prediction cost and the target intra-frame prediction cost corresponding to each of the N blocks.
[0128] Taking the vector dimension of the target feature vector corresponding to the frame to be encoded as 12 as an example, correspondingly, 12 feature attributes are obtained based on the target inter-frame prediction cost and the target intra-frame prediction cost corresponding to each of the N blocks.
[0129] Using the feature vector X to represent the target feature vector, as shown in Table 1, the physical meanings of the 12 feature attributes X[0] to X
[11] are as follows:
[0130] Table 1 Feature Attributes
[0131]
[0132]
[0133] It should be noted that in the embodiments of the present application, when initializing the feature vector X, the value of each feature vector is set to a first numerical value, for example, 0.
[0134] Next, the calculation processes of 12 feature attributes X[0] to X
[11] will be described separately.
[0135] Among them, the value of X[0] is N.
[0136] X[1] can be calculated by formula (5):
[0137]
[0138] where intra_error i represents the target intra-frame prediction cost of the i-th block among N blocks.
[0139] X[2] can be calculated by formula (6):
[0140]
[0141] where intra_error i represents the target intra-frame prediction cost of the i-th block among N blocks, and coded_error i represents the first target inter-frame prediction cost of the i-th block among N blocks, and min() is the minimum value function.
[0142] X[3] can be calculated by formula (7):
[0143]
[0144] where intra_error i represents the target intra-frame prediction cost of the i-th block among N blocks, and sr_coded_error i represents the second target inter-frame prediction cost of the i-th block among N blocks, and min() is the minimum value function.
[0145] X[4] can be calculated by formula (8):
[0146] X[4] = num1 / X[1] Formula (8)
[0147] where num1 represents the number of blocks among N blocks that satisfy condition 1 or condition 2, where condition 1 is intra_error > coded_error, and condition 2 is intra_error > sr_coded_error.
[0148] X[5] can be calculated by formula (9):
[0149] X[5] = num2 / X[1] Equation (9)
[0150] where num3 represents the number of blocks among N blocks that satisfy Condition 3 or Condition 4, where Condition 3 is intra_error > coded_error and ||mv last || > 0, and Condition 4 is intra_error > sr_coded_error and ||mv golden || > 0.
[0151] X[6] can be calculated using Equation (10):
[0152] X[6] = num3 / X[1] Equation (10)
[0153] where num3 represents the number of blocks among N blocks that satisfy Condition 5, where Condition 5 is intra_error > sr_coded_error and coded_error > sr_coded_error.
[0154] X[7] can be calculated using Equation (11):
[0155] X[7] = num4 / X[1] Equation (11)
[0156] where num4 represents the number of blocks among N blocks that satisfy Condition 6, where Condition 6 is intra_error = 0.
[0157] In the embodiments of the present application, a set s of motion vectors that is initially empty is established, and the motion vectors of inter-frame prediction are stored when the cost of inter-frame prediction is less than the cost of intra-frame prediction. Specifically, the mv corresponding to the blocks that satisfy Condition 7 last are added to the set s, and for the blocks that do not satisfy Condition 7, the mv corresponding to the blocks that satisfy Condition 8 golden are added to the set s.
[0158] X[8] can be calculated using Equation (12):
[0159]
[0160] where mv i,x is the value of x in the i-th motion vector (x, y) in s, and M is the number of motion vectors in s.
[0161] X[9] can be calculated using Equation (13):
[0162]
[0163] where mvi,y is the value of y in the i-th motion vector (x, y) in s, and M is the number of motion vectors in s.
[0164] X
[10] can be calculated using formula (14):
[0165]
[0166] where mv i,x is the value of x in the i-th motion vector (x, y) in s, and M is the number of motion vectors in s.
[0167] X
[11] can be calculated using formula (15):
[0168]
[0169] where mv i,y is the value of y in the i-th motion vector (x, y) in s, and M is the number of motion vectors in s.
[0170] S305. Obtain the target feature vector corresponding to the frame to be encoded based on multiple feature attributes.
[0171] Specifically, use the feature vector X to represent the target feature vector. Referring to Table 1, the target feature vector can be expressed as {X[0], X[1], X[2], X[3], X[4], X[5], X[6], X[7], X[8], X[9], X
[10] , X
[11] }.
[0172] It should be noted that in the embodiments of the present application, S302 can be executed first and then S303, or S303 can be executed first and then S302, and there is no limitation on this.
[0173] S202. Determine the corresponding target encoding bits based on the bit allocation information and the specified coding rate.
[0174] It should be noted that in the embodiments of the present application, the bit allocation information refers to the relevant variables used for bit allocation. In AV1, bit allocation includes bit allocation for groups of key frames, bit allocation for key frames, bit allocation for groups of pictures, and bit allocation for each frame in a group of pictures. In the embodiments of the present application, it mainly relates to bit allocation for each frame in a group of pictures. Specifically, the bit allocation information includes the frame rate, and based on the bit allocation information and the specified coding rate, determine the target encoding bits allocated for the frame to be encoded. Among them, the frame rate refers to the number of frames transmitted per second.
[0175] S203. Input the target feature vector into the trained coefficient prediction model to obtain the target correction coefficient corresponding to the frame to be encoded. The coefficient prediction model is trained based on a training data set. Each training data contains a sample frame and the corresponding reference correction coefficient, and the reference correction coefficient is determined after multiple rounds of encoding for a sample frame.
[0176] S204. Obtain the target quantization parameter based on the target correction coefficient and the target coding bits, and encode the frame to be encoded based on the target quantization parameter to obtain the target encoded frame.
[0177] Specifically, the target quantization parameter can be calculated using formula (1), that is, the target quantization parameter = target correction coefficient × (E / target coding bits).
[0178] In the process of encoding the frame to be encoded based on the target quantization parameter to obtain the target encoded frame, first, the frame to be encoded is quantized based on the target quantization parameter to obtain the quantized segment data, and then, the quantized segment data is encoded to obtain the target encoded frame.
[0179] The principle of quantization is to divide the transformed matrix by a constant, which can be called the quantization parameter. The quantization parameter is used to indicate the quantization fineness in the quantization process. When the QP value is larger, coefficients in a larger value range will be quantized to the same output, so generally, it will bring larger distortion and lower bit rate. On the contrary, when the QP value is smaller, coefficients in a smaller value range will be quantized to the same output, so generally, it will bring smaller distortion and at the same time, a higher bit rate.
[0180] Coding is performed on the obtained quantized segment data after the frame to be encoded is quantized using the target quantization parameter. When coding, entropy coding or statistical coding methods can be used for coding processing.
[0181] In the implementation of this application, through machine learning algorithms, the problem of lagging correction coefficients is solved, thereby improving the accuracy of quantization parameters and enhancing the coding performance. At the same time, it can be adapted to different images. Through more accurate correction coefficients, the accuracy of quantization parameters is further improved. In addition, by performing multiple rounds of encoding on samples to obtain training data, the accuracy of model labels can be improved, thereby enhancing the model prediction accuracy.
[0182] Next, the model training process involved in this application will be described. The model training process includes a sample collection stage and a model training stage.
[0183] In the sample acquisition stage, in order to obtain model labels with better performance, in the implementation of this application, for sample frames, through multiple encodings, the target quantization parameter offset corresponding to the sample frame can be determined, and then the corresponding reference correction coefficient can be determined according to the target quantization parameter offset. In this way, through multiple encodings, the target quantization parameter offset that enables better encoding performance can be determined, so as to obtain the model label corresponding to the sample frame under better encoding performance, thereby improving the accuracy of training data and enhancing the model training effect.
[0184] Specifically, refer to Figure 4 As shown, it is a process for obtaining a training data set provided in an embodiment of this application, and the specific process is as follows:
[0185] S401. Obtain a sample sequence containing each sample frame, and based on the position of each sample frame in a Group of Pictures (GOP), determine the respective layer corresponding to each sample frame according to the preset corresponding relationship between each position and each layer.
[0186] In the embodiment of this application, the sample sequence can be any sequence from the standard test set. For example, 6 sequences are selected from the standard test set, and the 6 sequences are: FoodMarket4, CatRobot, BasketballDrive, PartyScene, BQSquare, and KristenAndSara. Any one of the 6 sequences can be used as the sample sequence. In this way, each sample sequence selected from each sequence in the standard test set can be used as the sample sequence and processed through S401 - S404 to obtain the training data set.
[0187] Among them, the position of each sample frame in the GOP can also be understood as the display order of each sample frame in the GOP. The position of each sample frame in the GOP can be represented by a serial number. Usually, the GOP consists of 17 frames, so the serial numbers are 0 - 16.
[0188] In the embodiment of this application, only 6 layers are taken as an example for illustration. The 6 layers are layer 0, layer 1, layer 2, layer 3, layer 4, and layer 5. Layer 0 can be called the 0th layer, layer 1 can be called the 1st layer, layer 2 can be called the 2nd layer, layer 3 can be called the 3rd layer, layer 4 can be called the 4th layer, and layer 5 can be called the 5th layer.
[0189] For example, refer to Figure 5As shown in the figure, the frames included in the GOP are successively IBBBBBBBBBBBBBBB. The sequence number of the I frame is the 0th frame. The I frame is located at the 0th layer, the 16th frame is located at the 1st layer, the 8th frame is located at the 2nd layer, the 4th and 12th frames are located at the 3rd layer, the 2nd, 6th, 10th, and 14th frames are located at the 4th layer, and the 1st, 3rd, 5th, 7th, 9th, 11th, 13th, and 15th frames are located at the 5th layer.
[0190] S402. Based on the initial quantization parameters preset for the sample sequence and the initial quantization parameter offsets corresponding to each respective layer preset, perform multiple rounds of encoding on each sample frame to obtain the target quantization parameter offsets corresponding to each respective layer.
[0191] In the embodiments of the present application, for each layer, the sample frames can be encoded respectively according to multiple candidate offsets to determine the quantization parameter offset with the best encoding performance. Specifically, refer to Figure 6 As shown in the figure, when executing S402, for each layer, perform the following steps:
[0192] S601. Based on the initial quantization parameter offset corresponding to layer L, as well as the preset offset step size and number of offsets, determine multiple candidate offsets corresponding to layer L.
[0193] Among them, layer L can be any one of layer 0, layer 1, layer 2, layer 3, layer 4, and layer 5. The initial quantization parameter offsets corresponding to each layer can be the same or different, and there is no limitation in this regard. In this article, only the case where the initial quantization parameter offsets corresponding to each layer are the same is taken as an example for illustration.
[0194] Suppose the preset offset step size is 1, the preset number of offsets is 10 times, and the initial quantization parameter offset corresponding to layer L is -5. Determine 11 candidate offsets corresponding to layer L. The 11 candidate offsets are respectively: -5, -4, -3, -2, -1, 0, 1, 2, 3, 4, 5. In the following text, △q L is used to represent the candidate offset corresponding to layer L. The candidate offsets corresponding to each layer can be expressed as △q0, △q1, △q2, △q3, △q4, and △q5.
[0195] S602. Based on the initial quantization parameters preset for the sample sequence and multiple candidate offsets, respectively encode at least one sample frame corresponding to layer L to obtain multiple encoding results.
[0196] Specifically, when executing S602, the following steps can be adopted:
[0197] S6021. Based on the initial quantization parameters preset for the sample sequence, determine the initial quantization parameters corresponding to at least one sample frame respectively.
[0198] The initial quantization parameter preset for the sample sequence is called quantization parameter Q, and the initial quantization parameter corresponding to the sample frame is called quantization parameter q. Exemplarily, the value of quantization parameter Q is 128.
[0199] Based on quantization parameter Q, using the AOM-AV1 method, calculate the quantization parameter q corresponding to each of at least one sample frame.
[0200] S6022: Based on multiple candidate offsets and based on the initial quantization parameter corresponding to each of at least one sample frame, perform encoding on at least one sample frame respectively to obtain multiple encoding results.
[0201] Specifically, for sample frame y, based on multiple candidate offsets and based on quantization parameter q corresponding to sample frame y, it can be determined that the quantization parameter used for encoding sample frame y is q + △q L and then, based on quantization parameter q + △q L perform encoding on sample frame y. Sample frame y is any one of at least one sample frame corresponding to layer L.
[0202] For example, as shown in Figure 7 , the 11 candidate offsets include: -5, -4, -3, -2, -1, 0, 1, 2, 3, 4, 5. At the first encoding, the value of △q L is -5, and sample frame y is encoded according to quantization parameter q - 5 to obtain encoding result 1. At the second encoding, the value of △q L is -4, and sample frame y is encoded according to quantization parameter q - 4 to obtain encoding result 2. At the third encoding, the value of △q L is -3, and sample frame y is encoded according to quantization parameter q - 3 to obtain encoding result 3. At the fourth encoding, the value of △q L is -2, and sample frame y is encoded according to quantization parameter q - 2 to obtain encoding result 4. Similarly, sample frame y is encoded 11 times to obtain the encoding results corresponding to each of the 11 encodings.
[0203] Through the above implementation method, according to the values of candidate offsets from -5 to 5, the sequence can be encoded 11 times. When determining the target offset according to the encoding results, the accuracy of the target offset can be improved, so that the target offset is the offset with the best performance. In addition, in the embodiments of the present application, to improve the data processing efficiency, the offsets of the sample frames included in other layers can be fixed, and the offset of the current layer can be changed.
[0204] S603. Based on multiple coding results, determine a target offset from multiple candidate offsets, and use the target offset as the target quantization parameter offset corresponding to level L.
[0205] Specifically, when executing S603, it is possible to determine the coding errors corresponding to multiple candidate offsets based on multiple coding results, and determine, from the multiple candidate offsets, the candidate offsets whose corresponding coding errors meet the preset error conditions, and use the determined candidate offsets as the target offsets.
[0206] Among them, the preset error condition can be any one of the following conditions:
[0207] Condition A: Among all the coding errors, the coding error with the smallest value.
[0208] For example, if the coding errors corresponding to 11 candidate offsets are 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.8, 0.8, 0.9 respectively, determine, from the 11 candidate offsets, the candidate offset corresponding to the smallest coding error. Suppose the candidate offset corresponding to the coding error 0.1 is 5, then use the candidate offset 5 as the target offset.
[0209] Condition B: Among all the coding errors, the coding error with the value closest to the average value.
[0210] For example, if the coding errors corresponding to 11 candidate offsets are 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.8, 0.8, 0.9 respectively, and the average value of the coding errors is 0.7, the coding error closest to the average value is 0.7. Determine, from the 11 candidate offsets, the candidate offset corresponding to the coding error whose value is closest to the average value. Suppose the candidate offset corresponding to the coding error 0.7 is 4, then use the candidate offset 4 as the target offset.
[0211] Through the above implementation method, the target offset has the best coding performance. Furthermore, when determining the correction coefficient of the sample frame, the determined sample correction coefficient is also the correction coefficient under the best coding performance, thereby ensuring to a certain extent that the correction coefficient output by the model can make the frame to be coded have better coding performance and improve the model prediction accuracy.
[0212] It should be noted that in the embodiments of the present application, during the process of processing level L, it is also possible to perform multiple encodings on each sample frame corresponding to other levels except level L. When encoding each sample frame corresponding to other levels, the quantization parameter of each sample frame corresponding to other levels is q + △q 其他 .
[0213] S403. Determine the sample quantization parameter for the corresponding sample frame based on the preset initial quantization parameter for the sample sequence and the target quantization parameter offset corresponding to each level.
[0214] Specifically, based on the quantization parameter Q, calculate the quantization parameter q corresponding to each sample frame using the AOM - AV1 method. Then, use the target quantization parameter offset corresponding to each level as the target quantization parameter offset for the sample frame at the corresponding level. Based on the quantization parameter q corresponding to each sample frame and the target quantization parameter offset corresponding to each sample frame, determine the sample quantization parameter corresponding to each sample frame.
[0215] Still taking the sample frame y as an example, the level corresponding to the sample frame y is level L. Use the target quantization parameter offset Δq L corresponding to level L as the target quantization parameter offset for the sample frame y. Then, based on the quantization parameter q corresponding to the sample frame y and the target quantization parameter offset Δq L corresponding to the sample frame y, determine that the sample quantization parameter corresponding to the sample frame y is q + ΔqL.
[0216] S404. Obtain the reference correction coefficient corresponding to each sample frame based on the sample quantization parameter corresponding to each sample frame, and obtain the training data set based on each sample frame and the corresponding reference correction coefficient.
[0217] Still taking the sample frame y as an example, specifically, when executing S404, encode the sample frame y based on the sample quantization parameter corresponding to the sample frame y, record the sample quantization parameter and the coding bit number R corresponding to the sample frame y. Then, calculate the reference correction coefficient c corresponding to the sample frame y according to formula (16). After that, obtain a set of training data based on the sample frame y and the corresponding reference correction coefficient c.
[0218]
[0219] Where E is a constant. Exemplarily, when the frame type of the sample frame y is KEY_FRAME, take the value 2000000, otherwise take the value 1500000.
[0220] It should be noted that in the embodiments of the present application, the feature acquisition process of the sample frame y is the same as that in S301 - S305 above, which will not be elaborated here. Based on the feature vector corresponding to the sample frame y and the reference correction coefficient c corresponding to the sample frame y, form a training data (X, c).
[0221] Taking the support vector regression algorithm as an example, the model training stage of the present application will be described. It should be noted that in the embodiments of the present application, for the prediction of the correction coefficient, other machine learning or deep learning methods can also be used for substitution, such as random forest, etc.
[0222] The purpose of model training is to find a relationship f between the input and output through the given training samples D = {(X1, c1), (X2, c2)…(X n , c n n)], such that the value f(X) predicted according to the features is as close as possible to c. The principle of using support vector regression is to find f in the high-dimensional feature space such that
[0223]
[0224] By introducing Lagrange multipliers and transforming it into a dual problem, and then introducing a kernel function, the solution f of support vector regression is expressed as:
[0225]
[0226] where w and b are parameters to be determined, k() represents the kernel function, kernel represents the type of kernel function, C represents the penalty factor, and ∈ represents the termination condition.
[0227] Refer to Figure 8 As shown, in the embodiments of the present application, the model training process is as follows:
[0228] S801. Divide the training data set into a training set, a validation set, and a test set.
[0229] It should be noted that in the embodiments of the present application, the training data set can be divided into a training set, a validation set, and a test set according to a set ratio, or can be randomly divided into a training set, a validation set, and a test set.
[0230] For example, in the random division method, 60% of the training data in the training data set is divided into the training set, 20% of the training data is divided into the validation set, and the remaining 20% of the training data is divided into the test set.
[0231] S802. Based on the training set, train the coefficient prediction model to obtain a first error index.
[0232] In the embodiments of the present application, when initializing the coefficient prediction model, the parameters of the coefficient prediction model are respectively set as follows: the kernel function uses the radial basis kernel function (RBF), the penalty factor C = 3, the termination condition ∈ = 2e -3 , and the kernel coefficient γ of RBF = 0.1.
[0233] The input of the coefficient prediction model is a feature vector formed by a subset of the set composed of feature quantities and the results of mathematical operations on feature quantities, and the output is the predicted model parameters. The first error metric can be, but is not limited to, the mean absolute error (MAE).
[0234] S803. Train the coefficient prediction model based on the validation set to obtain a second error metric.
[0235] In the embodiments of the present application, the types of the second error metric, the first error metric, and the third error metric below can be the same. For example, the second error metric, the first error metric, and the third error metric are all MAE.
[0236] S804. If the second error metric is greater than the first error metric, adjust the model parameters of the coefficient prediction model based on the second error metric, and train the adjusted coefficient prediction model based on the validation set until the second error metric is not greater than the first error metric.
[0237] When adjusting the model parameters of the coefficient prediction model, the adjusted parameters are C, ∈, and γ.
[0238] S805. Input the test set into the coefficient prediction model to obtain a third error metric, and output the trained coefficient prediction model when the third error metric is less than a preset threshold.
[0239] Finally, on the training set and the validation set, the MAE is 0.0511, and the correlation between the predicted value and the label value is 0.9192. On the test set, the MAE is 0.1184, and the correlation between the predicted value and the label value is 0.7359.
[0240] It should be noted that in the embodiments of the present application, only the AVI coding standard is taken as an example for illustration. In actual application, the above quantization parameter determination method can also be applied to other coding standards. When applying the above quantization parameter determination method to other coding standards, specifically, the bitrate control process of other coding standards can be replaced with the bitrate control process proposed in the present application.
[0241] In AOM-AV1, referring to Table 2, compared with the prior art, the present application has improved coding performance in different videos, with an average improvement of 2.01% in coding performance. Among them, training sequences represents the video sequences used for training the machine learning model, and test sequences represents the video sequences used for testing.
[0242] Table 2 Test Results
[0243]
[0244] Next, the present application will be described in conjunction with specific embodiments.
[0245] Refer to Figure 9A As shown, the sequence to be processed contains multiple frames. Refer to Figure 9B As shown, taking only the first 17 frames as an example, level 0 contains frame 1, level 1 contains frame 16, level 2 contains frame 8, level 3 contains frames 4 and 12, level 4 contains frames 2, 6, 10, and 14, and level 5 contains frames 1, 3, 5, 7, 9, 11, 13, and 15. Initialize the offsets △q0 to △q5 of each level to -5, and the optimal offset △qb0 to △qb5 of each level is -5. It should be noted that the optimal offset can be understood as the quantization parameter offset in the above text. Set the current level i = 0, and set the quantization parameter Q for encoding the current sequence to 128.
[0246] Refer to Figure 10 As shown, during the acquisition process of the training dataset, for each sequence to be processed, the following operations are performed:
[0247] S1001. Obtain the sequence to be processed;
[0248] S1002. Determine whether i is equal to 5. If so, execute S1013; otherwise, execute S1003.
[0249] S1003. Determine whether △qi is equal to 5. If so, execute S1012; otherwise, execute S1004.
[0250] S1004. Determine whether all the frames contained in the sequence to be processed have been encoded. If so, execute S1009; otherwise, execute S1005.
[0251] S1005. Read a frame as the current frame.
[0252] S1006. Calculate the quantization parameter q corresponding to the current frame according to the given Q.
[0253] S1007. If the current frame belongs to level i, set the quantization parameter of the current frame to △q + △qi; otherwise, set the quantization parameter of the current frame to △q + △qbj, where j is the level of the current frame.
[0254] S1008. Encode the current frame according to the quantization parameter q corresponding to the current frame calculated in S1007, and return to execute S1004.
[0255] S1009. If the performance when the current level i uses △qi is better than the performance when the current level i uses △qbi, execute S1010; otherwise, execute S1011.
[0256] S1010. Set the value of △qbi to △qi.
[0257] S1011. △qi = △qi + 1.
[0258] S1012. i = i + 1.
[0259] S1013. Encode the sequence to be processed according to △qb0 to △qb5 obtained during the encoding process.
[0260] S1014. Record the quantization parameter q and the number of encoded bits R for each frame.
[0261] S1015. Calculate the check parameter c for each frame based on the quantization parameter q and the number of encoded bits R for each frame.
[0262] S1016. Based on the feature vector X of each frame and the check parameter c of each frame, form the training data (X, c).
[0263] Refer to Figure 11 As shown, the feature vector X of each frame in the video to be processed can be obtained by the following steps:
[0264] S1101. Read the video to be processed.
[0265] S1102. Determine whether all the frames included in the video to be processed have been pre - analyzed. If so, end; otherwise, execute S1003.
[0266] S1103. Read a frame that has not been pre - analyzed from the current video, and initialize the frame - level feature vector X, where the values of each feature attribute in the feature vector X are initially 0.
[0267] S1104. Divide the current frame into non - overlapping blocks, each with a width and height of 16, assign the number of blocks to the feature attribute X[0], and establish an initially empty set s of motion vectors. The set s is used to store the motion vectors of inter - frame prediction when the cost of inter - frame prediction is less than the cost of intra - frame prediction.
[0268] S1105. Perform pre - analysis on each block included in the current frame respectively to obtain the feature vector X corresponding to the current frame. For details, refer to S302 to S305, which will not be elaborated here.
[0269] Based on the same inventive concept, an embodiment of the present application provides a quantization parameter determination device. As Figure 12 shown, it is a schematic structural diagram of the quantization parameter determination device 1200, which may include:
[0270] An acquisition unit 1201, configured to acquire a frame to be encoded, obtain a target feature vector corresponding to the frame to be encoded based on the content attribute of the frame to be encoded, and obtain bit allocation information corresponding to the frame to be encoded;
[0271] A bit allocation unit 1202, configured to determine corresponding target encoding bits based on the bit allocation information and a specified code rate;
[0272] A coefficient determination unit 1203, configured to input the target feature vector into a trained coefficient prediction model to obtain a target correction coefficient corresponding to the frame to be encoded, where the coefficient prediction model is trained based on a training data set, and each training data includes a sample frame and a corresponding reference correction coefficient, and the reference correction coefficient is determined after multiple rounds of encoding for the one sample frame;
[0273] An encoding unit 1204, configured to obtain a target quantization parameter based on the target correction coefficient and the target encoding bits, and encode the frame to be encoded based on the target quantization parameter to obtain a target encoded frame.
[0274] As a possible implementation manner, it further includes a training unit 1205, and the training unit 1205 is configured to:
[0275] Acquire a sample sequence including each sample frame, and determine the level corresponding to each sample frame based on the position of each sample frame in the GOP according to the corresponding relationship between each preset position and each level;
[0276] Perform multiple rounds of encoding on each sample frame based on the initial quantization parameter preset for the sample sequence and the initial quantization parameter offset corresponding to each preset level, to obtain the target quantization parameter offset corresponding to each level;
[0277] Determine the sample quantization parameter of the corresponding sample frame based on the initial quantization parameter preset for the sample sequence and the target quantization parameter offset corresponding to each level;
[0278] Obtain the reference correction coefficient corresponding to each sample frame based on the sample quantization parameter corresponding to each sample frame, and obtain the training data set based on each sample frame and the corresponding reference correction coefficient.
[0279] As a possible implementation manner, when performing multiple rounds of encoding on each sample frame based on the initial quantization parameter preset for the sample sequence and the initial quantization parameter offset corresponding to each preset level to obtain the target quantization parameter offset corresponding to each level, the training unit 1205 is specifically configured to:
[0280] For each of the above-mentioned levels, the following operations are respectively performed:
[0281] Based on the initial quantization parameter offset corresponding to a level, as well as a preset offset step size and number of offsets, determine multiple candidate offsets corresponding to the level;
[0282] Based on the initial quantization parameter preset for the sample sequence and the multiple candidate offsets, respectively encode at least one sample frame corresponding to the level to obtain multiple encoding results;
[0283] Based on the multiple encoding results, determine a target offset from the multiple candidate offsets, and use the target offset as the target quantization parameter offset corresponding to the level.
[0284] As a possible implementation, when encoding at least one sample frame corresponding to the level respectively based on the initial quantization parameter preset for the sample sequence and the multiple candidate offsets to obtain multiple encoding results, the training unit 1205 is specifically configured to:
[0285] Based on the initial quantization parameter preset for the sample sequence, determine the initial quantization parameter corresponding to each of the at least one sample frame;
[0286] Based on the multiple candidate offsets and based on the initial quantization parameter corresponding to each of the at least one sample frame, respectively encode the at least one sample frame to obtain multiple encoding results.
[0287] As a possible implementation, when determining a target offset from the multiple candidate offsets based on the multiple encoding results, the training unit 1205 is specifically configured to:
[0288] Based on the multiple encoding results, determine the encoding error corresponding to each of the multiple candidate offsets;
[0289] Determine a candidate offset whose corresponding encoding error meets a preset error condition from the multiple candidate offsets, and use the determined candidate offset as the target offset.
[0290] As a possible implementation, when determining a target offset from the multiple candidate offsets based on the multiple encoding results, the training unit 1205 is specifically configured to:
[0291] Based on the multiple encoding results, determine the encoding error corresponding to each of the multiple candidate offsets;
[0292] From the multiple candidate offsets, determine the candidate offset corresponding to the coding error meeting the preset error condition, and use the determined candidate offset as the target offset.
[0293] As a possible implementation, when obtaining the target feature vector corresponding to the frame to be coded based on the content attribute of the frame to be coded, the training unit 1205 is specifically configured to:
[0294] Divide the frame to be coded into N non-overlapping blocks according to a preset block size, where N is a positive integer;
[0295] Perform intra-frame prediction on the N blocks respectively, and obtain the target intra-frame prediction cost corresponding to each of the N blocks based on the intra-frame prediction result;
[0296] Perform inter-frame prediction on the N blocks respectively, and obtain the target inter-frame prediction cost corresponding to each of the N blocks based on the inter-frame prediction result;
[0297] Obtain multiple feature attributes based on the inter-frame prediction cost and intra-frame prediction cost corresponding to each of the N blocks;
[0298] Obtain the target feature vector corresponding to the frame to be coded based on the multiple feature attributes.
[0299] As a possible implementation, the training unit 1205 is specifically configured to:
[0300] Divide the training data set into a training set, a validation set, and a test set;
[0301] Train the coefficient prediction model based on the training set to obtain a first error metric;
[0302] Train the coefficient prediction model based on the validation set to obtain a second error metric;
[0303] If the second error metric is greater than the first error metric, adjust the model parameters of the coefficient prediction model based on the second error metric, and train the adjusted coefficient prediction model based on the validation set until the second error metric is not greater than the first error metric;
[0304] Input the test set into the coefficient prediction model to obtain a third error metric, and output the trained coefficient prediction model when the third error metric is less than a preset threshold.
[0305] For the convenience of description, the above parts are divided into various modules (or units) according to functions and described separately. Of course, when implementing the present application, the functions of the various modules (or units) can be implemented in the same or multiple software or hardware.
[0306] Regarding the devices in the above embodiments, the specific manners in which each unit executes requests have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0307] In the embodiments of the present application, based on the content attributes of the frame to be encoded, the target feature vector corresponding to the frame to be encoded is obtained, and the bit allocation information corresponding to the frame to be encoded is obtained. Based on the bit allocation information and the specified code rate, the corresponding target encoding bits are determined; the target feature vector is input into the trained coefficient prediction model to obtain the target correction coefficient corresponding to the frame to be encoded; based on the target correction coefficient and the target encoding bits, the target quantization parameter is obtained, and based on the quantization parameter, the frame to be encoded is encoded to obtain the target encoded frame. In this way, through the machine learning algorithm, the problem of lagging correction coefficient is solved, thereby improving the accuracy of the quantization parameter and enhancing the encoding performance. At the same time, it can be adapted to different images, and through more accurate correction coefficients, the accuracy of the quantization parameter is further improved.
[0308] Those skilled in the art can understand that various aspects of the present application can be implemented as a system, a method, or a program product. Therefore, various aspects of the present application can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuit", "module", or "system" here.
[0309] Based on the same inventive concept, the embodiments of the present application also provide an electronic device. In one embodiment, the electronic device can be a server or a terminal device. Refer to Figure 13 As shown, it is a schematic structural diagram of a possible electronic device provided in the embodiments of the present application. Figure 13 In it, the electronic device 1300 includes: a processor 1310 and a memory 1320.
[0310] Among them, the memory 1320 stores a computer program executable by the processor 1310, and the processor 1310 can execute the steps of the above quantization parameter determination method by executing the instructions stored in the memory 1320.
[0311] The memory 1320 can be a volatile memory, such as a random-access memory (RAM); the memory 1320 can also be a non-volatile memory, such as a Read-Only Memory (ROM), a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD); or the memory 1320 is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 1320 can also be a combination of the above memories.
[0312] The processor 1310 can include one or more central processing units (CPUs) or be a digital processing unit, etc. When the processor 1310 executes the computer program stored in the memory 1320, the above method for determining quantization parameters is implemented.
[0313] In some embodiments, the processor 1310 and the memory 1320 can be implemented on the same chip. In some embodiments, they can also be separately implemented on independent chips.
[0314] In the embodiments of the present application, the specific connection medium between the above-mentioned processor 1310 and the memory 1320 is not limited. In the embodiments of the present application, taking the connection between the processor 1310 and the memory 1320 through a bus as an example, the bus is Figure 13 described by a thick line in []. The connection manners between other components are only for illustrative purposes and are not to be taken as a limitation. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of description, Figure 13 only a thick line is used to describe it in [], but it does not describe that there is only one bus or one type of bus.
[0315] Based on the same inventive concept, the embodiments of the present application provide a computer-readable storage medium, which includes a computer program. When the computer program runs on an electronic device, the computer program is used to cause the electronic device to execute the steps of the above method for determining quantization parameters. In some possible implementation manners, each aspect of the method for determining quantization parameters provided in the present application can also be implemented in the form of a program product, which includes a computer program. When the program product runs on an electronic device, the computer program is used to cause the electronic device to execute the steps in the above method for determining quantization parameters. For example, the electronic device can execute as Figure 2 the steps shown in [].
[0316] The program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, RAM, ROM, erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0317] The program product of the embodiments of the present application can adopt a CD-ROM and include a computer program, and can run on an electronic device. However, the program product of the present application is not limited thereto. In this document, the readable storage medium can be any tangible medium that contains or stores a computer program, and this computer program can be used by or in combination with a command execution system, apparatus, or device.
[0318] The readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries a readable computer program. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable signal medium can also be any readable medium other than the readable storage medium, and this readable medium can send, propagate, or transmit a computer program for use by or in combination with a command execution system, apparatus, or device.
[0319] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications to these embodiments once they know the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.
[0320] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application also intends to include these modifications and variations.
Claims
1. A method for determining quantization parameters, characterized in that, The method includes: Obtaining a frame to be encoded, and based on the content attributes of the frame to be encoded, obtaining a target feature vector corresponding to the frame to be encoded, and obtaining bit allocation information corresponding to the frame to be encoded; Based on the bit allocation information and a specified code rate, determining corresponding target encoding bits; Inputting the target feature vector into a trained coefficient prediction model to obtain a target correction coefficient corresponding to the frame to be encoded, where the coefficient prediction model is trained based on a training data set, and each training data includes a sample frame and a corresponding reference correction coefficient, and the reference correction coefficient is determined after multiple rounds of encoding for the one sample frame; Based on the target correction coefficient and the target encoding bits, obtaining a target quantization parameter, and based on the target quantization parameter, encoding the frame to be encoded to obtain a target encoded frame; Wherein, the obtaining a target feature vector corresponding to the frame to be encoded based on the content attributes of the frame to be encoded includes: Dividing the frame to be encoded into N non-overlapping blocks according to a preset block size, where N is a positive integer; Performing intra-frame prediction on the N blocks respectively, and based on the intra-frame prediction results, obtaining target intra-frame prediction costs corresponding to the N blocks respectively; Performing inter-frame prediction on the N blocks respectively, and based on the inter-frame prediction results, obtaining target inter-frame prediction costs corresponding to the N blocks respectively; Based on the inter-frame prediction costs and intra-frame prediction costs corresponding to the N blocks respectively, obtaining multiple feature attributes; Obtaining a target feature vector corresponding to the frame to be encoded based on the multiple feature attributes.
2. The method according to claim 1, characterized in that, The training data set is obtained through the following method: Obtaining a sample sequence including each sample frame, and based on the positions of the sample frames in a group of pictures (GOP), determining the levels corresponding to the sample frames respectively according to a preset correspondence between each position and each level; Based on an initial quantization parameter preset for the sample sequence, and based on an initial quantization parameter offset corresponding to each of the preset levels respectively, performing multiple rounds of encoding on the sample frames to obtain target quantization parameter offsets corresponding to each of the levels respectively; Based on the initial quantization parameter preset for the sample sequence, and based on the target quantization parameter offsets corresponding to each of the levels respectively, determining sample quantization parameters corresponding to the corresponding sample frames; Based on the sample quantization parameters corresponding to the sample frames respectively, obtaining reference correction coefficients corresponding to the sample frames respectively, and based on the sample frames and the corresponding reference correction coefficients, obtaining the training data set.
3. The method according to claim 2, wherein The performing multiple rounds of encoding on the sample frames based on the initial quantization parameter preset for the sample sequence, and based on the initial quantization parameter offsets corresponding to each of the preset levels respectively, to obtain target quantization parameter offsets corresponding to each of the levels respectively includes: For each level in the levels, performing the following operations respectively: Based on an initial quantization parameter offset corresponding to a level, and a preset offset step size and number of offsets, determining multiple candidate offsets corresponding to the level; Based on the preset initial quantization parameters for the sample sequence and the multiple candidate offsets, encode at least one sample frame corresponding to the one level respectively to obtain multiple encoding results; Based on the multiple encoding results, determine a target offset from the multiple candidate offsets, and use the target offset as the target quantization parameter offset corresponding to the one level.
4. The method according to claim 3, characterized in that, The encoding at least one sample frame corresponding to the one level respectively to obtain multiple encoding results based on the preset initial quantization parameters for the sample sequence and the multiple candidate offsets includes: Based on the preset initial quantization parameters for the sample sequence, determine the initial quantization parameters corresponding to each of the at least one sample frame; Based on the multiple candidate offsets and based on the initial quantization parameters corresponding to each of the at least one sample frame, encode the at least one sample frame respectively to obtain multiple encoding results.
5. The method according to claim 3, wherein The determining a target offset from the multiple candidate offsets based on the multiple encoding results includes: Based on the multiple encoding results, determine the encoding errors corresponding to the multiple candidate offsets respectively; Determine a candidate offset whose corresponding encoding error meets the preset error condition from the multiple candidate offsets, and use the determined candidate offset as the target offset.
6. The method according to any one of claims 1-5, characterized in that, The coefficient prediction model is trained in the following manner: Divide the training data set into a training set, a validation set and a test set; Based on the training set, train the coefficient prediction model to obtain a first error metric; Based on the validation set, train the coefficient prediction model to obtain a second error metric; If the second error metric is greater than the first error metric, then based on the second error metric, adjust the model parameters of the coefficient prediction model, and based on the validation set, train the adjusted coefficient prediction model until the second error metric is not greater than the first error metric; Input the test set into the coefficient prediction model to obtain a third error metric, and output the trained coefficient prediction model when the third error metric is less than a preset threshold.
7. A quantization parameter determination device, characterized in that including: An acquisition unit, configured to acquire a frame to be encoded, and based on the content attribute of the frame to be encoded, obtain a target feature vector corresponding to the frame to be encoded and obtain bit allocation information corresponding to the frame to be encoded; A bit allocation unit, configured to determine corresponding target encoding bits based on the bit allocation information and a specified code rate; A coefficient determination unit, configured to input the target feature vector into the trained coefficient prediction model to obtain a target correction coefficient corresponding to the frame to be encoded, wherein the coefficient prediction model is trained based on a training data set, and each training data includes a sample frame and a corresponding reference correction coefficient, and the reference correction coefficient is determined after multiple rounds of encoding for the one sample frame; An encoding unit, configured to obtain a target quantization parameter based on the target correction coefficient and the target encoding bits, and encode the frame to be encoded based on the target quantization parameter to obtain a target encoded frame; When the obtaining unit obtains the target feature vector corresponding to the frame to be encoded based on the content attribute of the frame to be encoded, it specifically is used for: dividing the frame to be encoded into N non-overlapping blocks according to a preset block size, where N is a positive integer; respectively performing intra-frame prediction on the N blocks, and obtaining the target intra-frame prediction cost corresponding to each of the N blocks based on the intra-frame prediction result; respectively performing inter-frame prediction on the N blocks, and obtaining the target inter-frame prediction cost corresponding to each of the N blocks based on the inter-frame prediction result; obtaining multiple feature attributes based on the inter-frame prediction cost and the intra-frame prediction cost corresponding to each of the N blocks; and obtaining the target feature vector corresponding to the frame to be encoded based on the multiple feature attributes.
8. The device according to claim 7, characterized in that, It further includes a training unit, and the training unit is used for: obtaining a sample sequence including each sample frame, and determining the level corresponding to each sample frame based on the position of each sample frame in a group of pictures (GOP) according to the corresponding relationship between each preset position and each level; performing multi-round encoding on each sample frame based on the initial quantization parameter preset for the sample sequence and the initial quantization parameter offset corresponding to each preset level, to obtain the target quantization parameter offset corresponding to each level; determining the sample quantization parameter of the corresponding sample frame based on the initial quantization parameter preset for the sample sequence and the target quantization parameter offset corresponding to each level; obtaining the reference correction coefficient corresponding to each sample frame based on the sample quantization parameter corresponding to each sample frame, and obtaining the training data set based on each sample frame and the corresponding reference correction coefficient.
9. The device according to claim 8, characterized in that When performing multi-round encoding on each sample frame based on the initial quantization parameter preset for the sample sequence and the initial quantization parameter offset corresponding to each preset level to obtain the target quantization parameter offset corresponding to each level, the training unit specifically is used for: performing the following operations respectively for each level in each of the levels: determining multiple candidate offsets corresponding to one level based on the initial quantization parameter offset corresponding to one level, a preset offset step size, and the number of offsets; encoding at least one sample frame corresponding to one level respectively based on the initial quantization parameter preset for the sample sequence and the multiple candidate offsets, to obtain multiple encoding results; determining a target offset from the multiple candidate offsets based on the multiple encoding results, and using the target offset as the target quantization parameter offset corresponding to one level.
10. The device according to claim 9, characterized in that, When encoding at least one sample frame corresponding to one level respectively based on the initial quantization parameter preset for the sample sequence and the multiple candidate offsets to obtain multiple encoding results, the training unit specifically is used for: determining the initial quantization parameter corresponding to each of the at least one sample frame based on the initial quantization parameter preset for the sample sequence; Encoding is respectively performed on the at least one sample frame based on the multiple candidate offsets and based on the initial quantization parameter corresponding to each of the at least one sample frame, to obtain a plurality of encoding results.
11. The device according to claim 9, characterized in that When determining a target offset from the multiple candidate offsets based on the plurality of encoding results, the training unit is specifically configured to: Determine the encoding error corresponding to each of the multiple candidate offsets based on the plurality of encoding results; Determine, from the multiple candidate offsets, a candidate offset whose corresponding encoding error meets a preset error condition, and use the determined candidate offset as the target offset.
12. An electronic device, characterized in that, It includes a processor and a memory. Among them, the memory stores a computer program, and when the computer program is executed by the processor, the processor is caused to execute the steps of the method according to any one of claims 1 to 6.
13. A computer-readable storage medium, characterized in that, It includes a computer program, and when the computer program runs on an electronic device, the computer program is used to cause the electronic device to execute the steps of the method according to any one of claims 1 to 6.
14. A computer program product, characterized in that, It includes a computer program, the computer program is stored in a computer-readable storage medium, and a processor of an electronic device reads and executes the computer program from the computer-readable storage medium, so that the electronic device executes the steps of the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Video coding method and device based on video quality and electronic equipment
CN110418134A
Moving image encoding device, moving image decoding device, moving image encoding method and moving image decoding method
JP2012023613A