Quantization parameter prediction model training, image coding method and device

By performing block-level quantization parameter prediction and encoding/decoding processing on the sample image set, a quantization parameter prediction model is trained, which solves the problems of high training cost and low efficiency in the existing technology, and achieves efficient bitrate allocation and improved coding performance.

CN116546225BActive Publication Date: 2025-12-02BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310440277.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-23
Publication Date
2025-12-02
Estimated Expiration
2043-04-23

AI Technical Summary

Technical Problem

Existing technologies for deep learning-based quantization parameter prediction models suffer from high training costs and low efficiency. Inappropriate quantization parameters lead to low encoding efficiency, poor performance, and an inability to effectively allocate bitrate.

Method used

By acquiring a sample image set, block-level quantization parameter prediction is performed, and encoding and decoding processing is carried out in combination with a pre-defined guide codec. The quantization parameter prediction model is trained using the encoding performance data to achieve unsupervised quantization parameter prediction.

Benefits of technology

It reduces model training time, improves training efficiency and system performance, enhances the effectiveness of quantization parameters and coding efficiency, and achieves effective bitrate allocation and improved coding performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116546225B_ABST
    Figure CN116546225B_ABST
Patent Text Reader

Abstract

This disclosure relates to a quantization parameter prediction model training method and apparatus for image encoding. The training method includes acquiring a first set of sample images; performing block-level quantization parameter prediction on the first sample images based on a first preset model to be trained, obtaining sample block-level quantization parameters; using the sample block-level quantization parameters as quantization parameters, performing encoding and decoding processing on the first sample images based on a preset differentiable codec, obtaining a first sample reconstructed image and a first sample block-level bitrate; performing encoding performance analysis based on each first sample image, the first sample reconstructed image, and the first sample block-level bitrate, obtaining first encoding performance data; and using the first encoding performance data as the quantization prediction loss of the first preset model to be trained, training the first preset model to be trained, to obtain a quantization parameter prediction model. Utilizing embodiments of this disclosure can reduce model training time, improve model training efficiency, achieve block-level bitrate allocation, and improve encoding performance and efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of Internet technology, and in particular to a method and apparatus for training a quantitative parameter prediction model and encoding images. Background Technology

[0002] With the rapid development of the internet, the amount of multimedia data such as images and videos is growing exponentially, putting enormous pressure on storage and transmission. Therefore, there is an urgent need for more efficient compression encoding of multimedia data such as images and videos.

[0003] Bitrate is a crucial factor affecting coding efficiency, and bitrate allocation is achieved through adaptive adjustment of quantization parameters. Traditional coding techniques, relying on manual design and optimization, have limited capabilities and cannot adaptively adjust quantization parameters based on different multimedia content. While some deep learning-based quantization parameter prediction models have emerged, these require supervised learning. This involves exhaustive search to obtain optimized quantization parameters for the image to be encoded, which then guide the deep learning model to learn from these labels. Furthermore, obtaining optimized quantization parameters is time-consuming and computationally resource-intensive. Consequently, deep learning-based quantization parameter prediction models often assign a single quantization parameter label to each image, leading to high training costs, low training efficiency, and the possibility of inappropriate quantization parameters, resulting in ineffective bitrate allocation and ultimately, low coding efficiency and poor coding performance. Summary of the Invention

[0004] This disclosure provides a method and apparatus for training a quantization parameter prediction model and encoding images, to at least solve the technical problems in related technologies such as high training cost, low training efficiency, unreasonable quantization parameters, inability to effectively allocate bitrate, low encoding efficiency, and poor encoding performance. The technical solution of this disclosure is as follows:

[0005] According to a first aspect of the present disclosure, a method for training a quantization parameter prediction model is provided, comprising:

[0006] Obtain the first sample image set;

[0007] Based on the first preset training model, block-level quantization parameters are predicted for each first sample image in the first sample image set to obtain the sample block-level quantization parameters corresponding to each first sample image.

[0008] Using the sample block-level quantization parameters as the quantization parameters corresponding to each first sample image, each first sample image is encoded and decoded based on a preset guide codec to obtain the first sample reconstructed image of each first sample image and the first sample block-level bitrate corresponding to each first sample image;

[0009] Based on each first sample image, the first sample reconstructed image, and the first sample block-level bitrate, coding performance analysis is performed to obtain the first coding performance data corresponding to each first sample image;

[0010] Using the first encoding performance data as the quantization prediction loss of the first preset training model, the first preset training model is trained to obtain a quantization parameter prediction model.

[0011] According to a second aspect of the present disclosure, an image encoding method is provided, comprising:

[0012] Obtain at least one first image to be encoded;

[0013] Each first image to be encoded is input into the first prediction model to predict spatial block-level quantization parameters, thereby obtaining the sixth block-level quantization parameters corresponding to each first image to be encoded.

[0014] Based on the sixth block-level quantization parameters, each first image to be encoded is quantized and encoded to obtain image encoding information corresponding to each first image to be encoded.

[0015] Wherein, the first prediction model is a quantized parameter prediction model obtained by training the first training model with spatial domain coding performance data as the quantized prediction loss of the first training model. The spatial domain coding performance data is obtained by coding performance analysis based on each first sample image in the first sample image set, the first sample reconstructed image of each first sample image, and the first sample block-level code rate corresponding to each first sample image. The first sample reconstructed image and the first sample block-level code rate are obtained by encoding and decoding each first sample image based on a preset guided codec.

[0016] According to a third aspect of the embodiments of this disclosure, another image encoding method is provided, comprising:

[0017] Obtain at least one second image to be encoded corresponding to the first video to be encoded;

[0018] A first target associated image is determined for each second image to be encoded; the first target associated image includes at least one of the following: at least one neighboring frame image of each second image to be encoded, a prediction frame image corresponding to the at least one neighboring frame image, and a residual frame image corresponding to the at least one neighboring frame image;

[0019] Each second image to be encoded and the first target associated image corresponding to each second image to be encoded are input into the second prediction model to predict the temporal block-level quantization parameters, thereby obtaining the seventh block-level quantization parameters corresponding to each second image to be encoded.

[0020] Based on the seventh block-level quantization parameters, each second image to be encoded is quantized and encoded to obtain the image encoding information corresponding to each second image to be encoded.

[0021] The second prediction model is a quantized parameter prediction model obtained by training the second training model using temporal coding performance data as the quantized prediction loss of the second training model. The temporal coding performance data is obtained by analyzing the coding performance of each first sample image in the first sample image set, the first sample reconstructed image of each first sample image, the first sample associated image corresponding to each first sample image, the second sample reconstructed image of the first sample associated image, the first sample block-level code rate corresponding to each first sample image, and the second sample block-level code rate corresponding to the first sample associated image. The first sample reconstructed image, the first sample block-level code rate, the second sample reconstructed image, and the second sample block-level code rate are obtained by encoding and decoding each first sample image and the first sample associated image using a preset guided codec.

[0022] According to a fourth aspect of the embodiments of this disclosure, another image encoding method is provided, comprising:

[0023] Obtain at least one third image corresponding to the second video to be encoded;

[0024] Determine a second target associated image corresponding to each third image to be encoded; the second target associated image includes at least one of the following: at least one neighboring frame image of each third image to be encoded, a prediction frame image corresponding to the at least one neighboring frame image, and a residual frame image corresponding to the at least one neighboring frame image;

[0025] Each third image to be encoded and the second target associated image corresponding to each third image to be encoded are input into the third prediction model to perform spatiotemporal joint block-level quantization parameter prediction, so as to obtain the eighth block-level quantization parameter and the ninth block-level quantization parameter corresponding to each third image to be encoded.

[0026] Based on the eighth block-level quantization parameter or the ninth block-level quantization parameter, each third image to be encoded is quantized and encoded to obtain the image encoding information corresponding to each third image to be encoded.

[0027] The third prediction model is a quantized parameter prediction model obtained by training the third training model using joint coding performance data as the quantized prediction loss. The joint coding performance data is obtained by analyzing the coding performance of each first sample image in the first sample image set, the first sample reconstructed image of each first sample image, the second sample associated image corresponding to each first sample image, the third sample reconstructed image of the second sample associated image, the first sample block-level code rate corresponding to each first sample image, and the third sample block-level code rate corresponding to the second sample associated image. The first sample reconstructed image, the first sample block-level code rate, the point-multiple third sample reconstructed image, and the third sample block-level code rate are obtained by encoding and decoding each first sample image and the second sample associated image using a preset guided codec.

[0028] According to a fifth aspect of the present disclosure, a quantization parameter prediction model training apparatus is provided, comprising:

[0029] The first sample image set acquisition module is configured to acquire the first sample image set.

[0030] The first quantization parameter prediction module is configured to perform block-level quantization parameter prediction on each first sample image in the first sample image set based on a first preset training model, so as to obtain the sample block-level quantization parameter corresponding to each first sample image.

[0031] The first encoding / decoding processing module is configured to perform encoding / decoding processing on each first sample image based on a preset doubly oriented codec, using the sample block-level quantization parameters as the quantization parameters corresponding to each first sample image, to obtain the first sample reconstructed image of each first sample image and the first sample block-level bitrate corresponding to each first sample image;

[0032] The first coding performance analysis module is configured to perform coding performance analysis based on each first sample image, the first sample reconstructed image and the first sample block-level bitrate, and obtain the first coding performance data corresponding to each first sample image.

[0033] The first model training module is configured to execute a quantization prediction loss using the first encoding performance data as the first preset model to be trained, and train the first preset model to be trained to obtain a quantization parameter prediction model.

[0034] In an optional embodiment, the sample block-level quantization parameter includes a first block-level quantization parameter, which is a spatial block-level quantization parameter; the first preset training model includes a first training model; the first coding performance data includes spatial coding performance data; and the quantization parameter prediction model includes a first prediction model.

[0035] The first quantization parameter prediction module is specifically configured to perform spatial block-level quantization parameter prediction by inputting each first sample image into the first model to be trained, so as to obtain the first block-level quantization parameter corresponding to each first sample image;

[0036] The first encoding / decoding processing module is specifically configured to perform encoding / decoding processing on each first sample image based on the preset doubly quantized codec, using the first block-level quantization parameter as the quantization parameter corresponding to each first sample image, to obtain the first sample reconstructed image and the first sample block-level bitrate;

[0037] The first coding performance analysis module is specifically configured to perform coding performance analysis based on each first sample image, the first sample reconstructed image, and the first sample block-level bitrate, to obtain the spatial coding performance data corresponding to each first sample image;

[0038] The first model training module is specifically configured to perform quantization prediction loss on the first model to be trained using the spatial coding performance data as the first model to be trained, and to train the first model to be trained to obtain the first prediction model.

[0039] In an optional embodiment, the first quantization parameter prediction module is specifically configured to perform spatial block-level quantization parameter prediction by inputting each first sample image and the first preset block-level quantization parameter corresponding to each first sample image into the first model to be trained, thereby obtaining the first block-level quantization parameter.

[0040] In an optional embodiment, the sample block-level quantization parameter includes a second block-level quantization parameter, which is a time-domain block-level quantization parameter; the first preset training model includes a second training model; the first sample image set includes at least one sample video frame image corresponding to a sample video; the device further includes:

[0041] The first sample associated image determination module is configured to determine the first sample associated image corresponding to each first sample image; the first sample associated image includes at least one of the following: at least one adjacent frame image of each first sample image, a prediction frame image corresponding to the at least one adjacent frame image, and a residual frame image corresponding to the at least one adjacent frame image;

[0042] The first quantization parameter prediction module is specifically configured to perform temporal block-level quantization parameter prediction by inputting each first sample image and the first sample associated image corresponding to each first sample image into the second model to be trained, so as to obtain the second block-level quantization parameter corresponding to each first sample image.

[0043] In an optional embodiment, the first quantization parameter prediction module is specifically configured to perform temporal block-level quantization parameter prediction by inputting each first sample image, the second preset block-level quantization parameter corresponding to each first sample image, and the first sample associated image corresponding to each first sample image into the second model to be trained, so as to obtain the second block-level quantization parameter.

[0044] In an optional embodiment, the first coding performance data includes time-domain coding performance data; the quantization parameter prediction model includes a second prediction model; the apparatus further includes:

[0045] The second encoding / decoding processing module is configured to perform encoding / decoding processing on the first sample associated image based on the preset controllable codec, using the first sample reconstructed image corresponding to each first sample image as the reference frame image of the first sample associated image, using the third preset block-level quantization parameter corresponding to the first sample associated image as the quantization parameter corresponding to the first sample associated image, to obtain the second sample reconstructed image of the first sample associated image and the second sample block-level bitrate corresponding to the first sample associated image.

[0046] The first coding performance analysis module is specifically configured to perform coding performance analysis based on each first sample image, the first sample associated image, the first sample reconstructed image, the second sample reconstructed image, the first sample block-level code rate, and the second sample block-level code rate to obtain the temporal coding performance data.

[0047] The first model training module is specifically configured to perform quantization prediction loss on the second model to be trained using the time-domain coding performance data, and to train the second model to obtain the second prediction model.

[0048] In an optional embodiment, the first quantization parameter prediction module is specifically configured to perform block-level quantization parameter prediction by inputting each first sample image, the second preset block-level quantization parameter corresponding to each first sample image, the first sample associated image corresponding to each first sample image, and the third preset block-level quantization parameter corresponding to the first sample associated image into the second model to be trained, so as to obtain the second block-level quantization parameter and the third block-level quantization parameter corresponding to the first sample associated image.

[0049] The second encoding / decoding processing module is specifically configured to perform encoding / decoding processing on the first sample associated image based on the preset guided codec, using the first sample reconstructed image corresponding to each first sample image as the reference frame image of the first sample associated image, using the third block-level quantization parameter corresponding to the first sample associated image as the quantization parameter corresponding to the first sample associated image, to obtain the second sample reconstructed image and the second sample block-level bitrate.

[0050] In an optional embodiment, the sample block-level quantization parameters include a fourth block-level quantization parameter and a fifth block-level quantization parameter, wherein the fourth block-level quantization parameter is a spatial domain block-level quantization parameter, and the fifth block-level quantization parameter is a temporal domain block-level quantization parameter; the first preset training model includes a third training model; the first sample image set includes at least one sample video frame image corresponding to a sample video; the device further includes:

[0051] The second sample association image determination module is configured to determine the second sample association image corresponding to each first sample image; the second sample association image includes at least one of the following: at least one neighboring frame image of each first sample image, a prediction frame image corresponding to the at least one neighboring frame image, and a residual frame image corresponding to the at least one neighboring frame image;

[0052] The first quantization parameter prediction module is specifically configured to perform spatiotemporal joint block-level quantization parameter prediction by inputting each first sample image and the second sample associated image corresponding to each first sample image into the third model to be trained, thereby obtaining the fourth block-level quantization parameter and the fifth block-level quantization parameter corresponding to each first sample image.

[0053] In an optional embodiment, the first coding performance data includes joint coding performance data; the quantization parameter prediction model includes a third prediction model; the apparatus further includes:

[0054] The third encoding and decoding processing module is configured to perform encoding and decoding processing on the second sample associated image based on the preset controllable codec, using the first sample reconstructed image corresponding to each first sample image as the reference frame image of the second sample associated image corresponding to each first sample image, using the fourth preset block-level quantization parameter corresponding to the second sample associated image as the quantization parameter corresponding to the second sample associated image, to obtain the third sample reconstructed image of the second sample associated image and the corresponding third sample block-level bitrate of the second sample associated image;

[0055] The first coding performance analysis module is specifically configured to perform coding performance analysis based on each first sample image, the second sample associated image, the first sample reconstructed image, the third sample reconstructed image, the first sample block-level code rate, and the third sample block-level code rate to obtain the joint coding performance data;

[0056] The first model training module is specifically configured to perform quantization prediction loss on the third model to be trained using the joint coding performance data as the quantization prediction loss, and to train the third model to obtain the third prediction model.

[0057] In an optional embodiment, the apparatus further includes:

[0058] The data acquisition module is configured to acquire the second sample image set and the fifth preset block-level quantization parameters corresponding to each second sample image in the second sample image set;

[0059] The fourth encoding / decoding processing module is configured to perform encoding / decoding processing on each second sample image based on the second preset training model, using the fifth preset block-level quantization parameter as the quantization parameter corresponding to each second sample image, to obtain the fourth sample reconstructed image of each second sample image and the fourth sample block-level bitrate corresponding to each first sample image;

[0060] The second coding performance analysis module is configured to perform coding performance analysis based on each second sample image, the fourth sample reconstructed image and the fourth sample block-level bitrate, and obtain the second coding performance data corresponding to each second sample image;

[0061] The second model training module is configured to perform quantization encoding loss on the second preset model to be trained using the second encoding performance data as the second preset model to be trained, and to obtain the preset differentiable codec.

[0062] According to a sixth aspect of the present disclosure, an image encoding apparatus is provided, comprising:

[0063] The first image to be encoded acquisition module is configured to acquire at least one first image to be encoded.

[0064] The second quantization parameter prediction module is configured to perform spatial block-level quantization parameter prediction by inputting each first image to be encoded into the first prediction model, thereby obtaining the sixth block-level quantization parameter corresponding to each first image to be encoded.

[0065] The first quantization encoding module is configured to perform quantization encoding on each first image to be encoded based on the sixth block-level quantization parameters to obtain image encoding information corresponding to each first image to be encoded.

[0066] Wherein, the first prediction model is a quantized parameter prediction model obtained by training the first training model with spatial domain coding performance data as the quantized prediction loss of the first training model. The spatial domain coding performance data is obtained by coding performance analysis based on each first sample image in the first sample image set, the first sample reconstructed image of each first sample image, and the first sample block-level code rate corresponding to each first sample image. The first sample reconstructed image and the first sample block-level code rate are obtained by encoding and decoding each first sample image based on a preset guided codec.

[0067] According to a seventh aspect of the present disclosure, another image encoding apparatus is provided, comprising:

[0068] The second image to be encoded acquisition module is configured to acquire at least one second image to be encoded corresponding to the first video to be encoded.

[0069] The first target association image determination module is configured to determine a first target association image corresponding to each second image to be encoded; the first target association image includes at least one of the following: at least one neighboring frame image of each second image to be encoded, a prediction frame image corresponding to the at least one neighboring frame image, and a residual frame image corresponding to the at least one neighboring frame image;

[0070] The third quantization parameter prediction module is configured to perform temporal block-level quantization parameter prediction by inputting each second image to be encoded and the first target associated image corresponding to each second image to be encoded into the second prediction model to obtain the seventh block-level quantization parameter corresponding to each second image to be encoded.

[0071] The second quantization encoding module is configured to perform quantization encoding on each second image to be encoded based on the seventh block-level quantization parameters, so as to obtain image encoding information corresponding to each second image to be encoded.

[0072] The second prediction model is a quantized parameter prediction model obtained by training the second training model using temporal coding performance data as the quantized prediction loss of the second training model. The temporal coding performance data is obtained by analyzing the coding performance of each first sample image in the first sample image set, the first sample reconstructed image of each first sample image, the first sample associated image corresponding to each first sample image, the second sample reconstructed image of the first sample associated image, the first sample block-level code rate corresponding to each first sample image, and the second sample block-level code rate corresponding to the first sample associated image. The first sample reconstructed image, the first sample block-level code rate, the second sample reconstructed image, and the second sample block-level code rate are obtained by encoding and decoding each first sample image and the first sample associated image using a preset guided codec.

[0073] According to an eighth aspect of the present disclosure, another image encoding apparatus is provided, comprising:

[0074] The third image to be encoded acquisition module is configured to acquire at least one third image to be encoded corresponding to the second video to be encoded.

[0075] The second target associated image determination module is configured to determine a second target associated image corresponding to each third image to be encoded; the second target associated image includes at least one of the following: at least one neighboring frame image of each third image to be encoded, a prediction frame image corresponding to the at least one neighboring frame image, and a residual frame image corresponding to the at least one neighboring frame image;

[0076] The fourth quantization parameter prediction module is configured to perform spatiotemporal joint block-level quantization parameter prediction by inputting each third image to be encoded and the second target associated image corresponding to each third image to be encoded into the third prediction model, thereby obtaining the eighth block-level quantization parameter and the ninth block-level quantization parameter corresponding to each third image to be encoded.

[0077] The third quantization encoding module is configured to perform quantization encoding on each third image to be encoded based on the eighth block-level quantization parameter or the ninth block-level quantization parameter, so as to obtain the image encoding information corresponding to each third image to be encoded.

[0078] The third prediction model is a quantized parameter prediction model obtained by training the third training model using joint coding performance data as the quantized prediction loss. The joint coding performance data is obtained by analyzing the coding performance of each first sample image in the first sample image set, the first sample reconstructed image of each first sample image, the second sample associated image corresponding to each first sample image, the third sample reconstructed image of the second sample associated image, the first sample block-level code rate corresponding to each first sample image, and the third sample block-level code rate corresponding to the second sample associated image. The first sample reconstructed image, the first sample block-level code rate, the point-multiple third sample reconstructed image, and the third sample block-level code rate are obtained by encoding and decoding each first sample image and the second sample associated image using a preset guided codec.

[0079] According to a ninth aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to execute the instructions to implement the method as described in any one of the first, second, third, or fourth aspects above.

[0080] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided such that, when instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the method described in any one of the first, second, third, or fourth aspects of the present disclosure.

[0081] According to a fifth aspect of the present disclosure, a computer program product comprising instructions is provided that, when run on a computer, causes the computer to perform the method as described in any one of the first, second, third, or fourth aspects above.

[0082] The technical solutions provided by the embodiments of this disclosure have at least the following beneficial effects:

[0083] In the training process of the quantization parameter prediction model, firstly, based on the first preset model to be trained, block-level quantization parameters are predicted for each first sample image in the first sample image set to obtain the sample block-level quantization parameters corresponding to each first sample image. Then, using the sample block-level quantization parameters as the quantization parameters corresponding to each first sample image, each first sample image is encoded and decoded based on a preset doubly oriented codec to obtain the first sample reconstructed image used for coding performance analysis and the first sample block-level bitrate corresponding to each first sample image. Furthermore, since the first sample reconstructed image and the first sample block-level bitrate are obtained based on the preset doubly oriented codec, it can be guaranteed that the quantization parameters for each first sample image are consistent with the predicted quantization parameters for each first sample image. The first coding performance data obtained by encoding analysis of the first sample reconstructed image and the first sample block-level bitrate is differentiable and can be used as the quantization prediction loss of the first preset training model to train the block-level quantization parameter prediction model. This allows for unsupervised quantization parameter prediction model training without obtaining model labels, greatly reducing model training time and improving model training efficiency, which in turn can improve system performance during model training. Furthermore, combining the first coding performance data with the block-level quantization parameter prediction training of the first preset training model can effectively improve the effectiveness of quantization parameters, achieve block-level bitrate allocation, and improve coding performance and coding efficiency.

[0084] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0085] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.

[0086] Figure 1 This is a schematic diagram illustrating an application environment according to an exemplary embodiment;

[0087] Figure 2 This is a flowchart illustrating a method for training a quantization parameter prediction model according to an exemplary embodiment;

[0088] Figure 3 This is a schematic diagram illustrating an encoding / decoding process using a preset reversible codec, according to an exemplary embodiment.

[0089] Figure 4 This is a schematic diagram illustrating a process for training a predefined, self-contained codec according to an exemplary embodiment;

[0090] Figure 5 This is a schematic diagram of a first prediction model training process provided according to an exemplary embodiment;

[0091] Figure 6 This is a schematic diagram of a second prediction model training process provided according to an exemplary embodiment;

[0092] Figure 7 This is a flowchart illustrating an image encoding method according to an exemplary embodiment;

[0093] Figure 8 This is a flowchart illustrating another image encoding method according to an exemplary embodiment;

[0094] Figure 9 This is a flowchart illustrating another image encoding method according to an exemplary embodiment;

[0095] Figure 10 This is a block diagram of a quantization parameter prediction model training device according to an exemplary embodiment;

[0096] Figure 11 This is a block diagram of an image encoding apparatus according to an exemplary embodiment;

[0097] Figure 12 This is a block diagram of another image encoding device according to an exemplary embodiment;

[0098] Figure 13 This is a block diagram of another image encoding device according to an exemplary embodiment;

[0099] Figure 14 This is a block diagram illustrating an electronic device for image encoding according to an exemplary embodiment;

[0100] Figure 15 This is a block diagram illustrating an electronic device for training a quantization parameter prediction model according to an exemplary embodiment. Detailed Implementation

[0101] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0102] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0103] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data used for analysis, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties.

[0104] Please see Figure 1 , Figure 1 This is a schematic diagram illustrating an application environment according to an exemplary embodiment, which may include a server 100 and a terminal 200.

[0105] In an optional embodiment, server 100 can be used to train a quantization parameter prediction model; server 100 can be an independent physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides cloud computing services.

[0106] In an optional embodiment, terminal 200 can use a server-trained quantization parameter prediction model to encode and decode multimedia data such as images and videos to be encoded. Specifically, terminal 200 can be, but is not limited to, electronic devices such as smartphones, desktop computers, tablets, laptops, smart speakers, digital assistants, augmented reality (AR) / virtual reality (VR) devices, and smart wearable devices, or software running on these electronic devices, such as applications. Optionally, the operating system running on the electronic device can include, but is not limited to, Android, iOS, Linux, and Windows.

[0107] In addition, it should be noted that, Figure 1 The example shown is merely one application environment provided by this disclosure. In practical applications, other application environments may also be included, such as terminal 300. Optionally, terminal 200 can send the encoded multimedia data to server 100, and server 100 can transmit the encoded multimedia data to terminal 300. Correspondingly, terminal 300 can decode the encoded multimedia data to obtain the multimedia data.

[0108] In the embodiments described in this specification, the server 100 and the terminal 200 can be directly or indirectly connected via wired or wireless communication, and this disclosure does not impose any restrictions.

[0109] Figure 2 This is a flowchart illustrating a method for training a quantization parameter prediction model according to an exemplary embodiment. Optionally, this method can be applied to electronic devices such as servers and terminals. Figure 2 As shown, the following steps may be included:

[0110] S201: Obtain the first sample image set.

[0111] In one specific embodiment, the first sample image set can be a collection of multiple first sample images. Optionally, the first sample image can include the original first sample image or the edge image corresponding to the original first sample image (an image obtained after edge extraction of the original first sample image); the original first sample image can be a video frame image or an image acquired in an image format.

[0112] In practical applications, during model training, training data (a large number of first sample images) can be obtained in advance, and the current set of sample images (i.e., the first sample image set) used for model training can be obtained sequentially from the training data.

[0113] S203: Based on the first preset model to be trained, perform block-level quantization parameter prediction on each first sample image in the first sample image set to obtain the sample block-level quantization parameters corresponding to each first sample image.

[0114] In one specific embodiment, the first preset model to be trained can be a pre-set deep learning model to be trained for block-level quantization parameter prediction. The specific model structure can be set according to the actual application requirements.

[0115] In an optional embodiment, in the scenario of predicting block-level quantization parameters in the spatial domain, the first preset model to be trained may include a first model to be trained; the sample block-level quantization parameters include first block-level quantization parameters, which are block-level quantization parameters in the spatial domain; optionally, the above-mentioned prediction of block-level quantization parameters for each first sample image in the first sample image set based on the first preset model to be trained, to obtain the sample block-level quantization parameters corresponding to each first sample image, may include:

[0116] Each first sample image is input into the first model to be trained to predict the spatial block-level quantization parameters, thereby obtaining the first block-level quantization parameters corresponding to each first sample image.

[0117] In an optional embodiment, the output of the first model to be trained can be either the first block-level quantization parameter corresponding to each first sample image, or the adjustment parameters of the block-level quantization algorithm corresponding to each first sample image. Accordingly, after obtaining the adjustment parameters of the block-level quantization algorithm, the first block-level quantization parameter can be determined by combining the adjustment parameters of the spatial domain block-level quantization algorithm corresponding to each first sample image output by the first model to be trained. Optionally, traditional spatial domain block-level quantization algorithms may include, but are not limited to, the AQ (Adaptive Quantization) algorithm. Correspondingly, the adjustment parameters of the spatial domain block-level quantization algorithm may include: adjustable intensity values ​​in the AQ algorithm, the SATD (Sum of Absolute Transformed Difference) cost between the predicted and original values ​​of the intra-frame mode, etc.

[0118] In the above embodiments, inputting a single sample image into the first model to be trained to predict block-level quantization parameters in the spatial domain can improve the effectiveness of spatial quantization parameters, thereby achieving block-level bitrate allocation and improving the efficiency and performance of subsequent image coding.

[0119] In an optional embodiment, the above-mentioned inputting each first sample image into the first model to be trained for spatial block-level quantization parameter prediction, to obtain the first block-level quantization parameters corresponding to each first sample image, includes:

[0120] Each first sample image and the first preset block-level quantization parameter corresponding to each first sample image are input into the first model to be trained to predict the spatial block-level quantization parameters, thereby obtaining the first block-level quantization parameters.

[0121] In one specific embodiment, the first preset block-level quantization parameter can be a spatial block-level quantization parameter determined by combining traditional spatial block-level quantization algorithms, or it can be a spatial block-level quantization parameter preset according to actual application requirements; optionally, the first preset block-level quantization parameter corresponding to each first sample image can be determined by combining a block-level AQ algorithm that considers spatial correlation.

[0122] In the above embodiments, using each first sample image and the corresponding first preset block-level quantization parameter as input to the first model to be trained allows the first model to be trained to perform block-level quantization parameter correction and adjustment based on the first preset block-level quantization parameter, thereby improving the accuracy of block-level quantization parameter prediction in the spatial domain.

[0123] In an optional embodiment, when the first sample image is the original image of the first sample, the edge image corresponding to the original image of the first sample can also be used as the input of the first model to be trained; when the first sample image is the edge image corresponding to the first sample image set, the original image of the first sample can also be used as the input of the first model to be trained, so that the first model to be trained can better learn image features during the process of predicting block-level quantization parameters in the spatial domain, thereby improving the prediction accuracy of block-level quantization parameters in the spatial domain.

[0124] S205: Using the sample block-level quantization parameters for each first sample image, perform encoding and decoding processing on each first sample image based on a preset doubly oriented codec to obtain the first sample reconstructed image of each first sample image and the first sample block-level bitrate corresponding to each first sample image;

[0125] In an optional embodiment, a preset, self-contained codec can be used to encode and decode images for image reconstruction and block-level bitrate prediction. Optionally, the preset, self-contained codec can be a deep learning model used for image encoding and decoding for image reconstruction and block-level bitrate prediction. Specifically, the model structure can be set according to actual application requirements. Optionally, in this embodiment, the same preset, self-contained codec can be used during the training of the temporal quantization parameter prediction model and the spatial quantization parameter prediction model; correspondingly, when performing motion estimation and motion compensation for each frame of image in the spatial domain in the preset, self-contained codec, it can be inter-frame prediction considering the correlation between each residual block in each frame; when performing motion estimation and motion compensation for each frame of image in the temporal domain in the preset, self-contained codec, it can be inter-frame prediction considering the correlation between each frame and adjacent frames. Optionally, different preset guided codecs can be used during the training of the temporal quantization parameter prediction model and the spatial quantization parameter prediction model; correspondingly, the preset guided codec for the spatial quantization parameter prediction model can perform motion estimation and motion compensation for each frame of the spatial image, and can be a single-frame intra-spatial prediction.

[0126] In an optional embodiment, the aforementioned preset differentiable codec may include: a differentiable motion estimation module, a differentiable motion compensation module, a differentiable transform module, a differentiable quantization module, a differentiable entropy encoding module, a differentiable inverse quantization module, a differentiable inverse transform module, and a differentiable loop filtering module; specifically, the differentiable motion estimation module, differentiable motion compensation module, differentiable transform module, differentiable quantization module, differentiable entropy encoding module, differentiable inverse quantization module, differentiable inverse transform module, and differentiable loop filtering module can all be constructed by combining neural networks. Optionally, the modules in the preset differentiable codec can also be constructed based on differentiable functions that can realize the functions of each module. Figure 3 As shown, Figure 3 This is a schematic diagram illustrating an encoding / decoding process using a preset, differentiated codec, according to an exemplary embodiment. Specifically, taking the encoding / decoding process for each frame of image (each first sample image) in the temporal domain as an example, the above-mentioned sample block-level quantization parameters are used as the quantization parameters corresponding to each first sample image. Encoding / decoding each first sample image based on the preset, differentiated codec to obtain the first sample reconstructed image of each first sample image and the first sample block-level bitrate corresponding to each first sample image may include: inputting each first sample image (hereinafter referred to as the current sample image) and the preset reference frame image corresponding to the current sample image (the preset reference frame image can be the reconstructed image of the previous frame sample image in the reconstructed image buffer; the first preset reference frame image can be obtained through intra-frame encoding / decoding) into a differentiated motion estimation module for motion estimation to obtain the MV (Motion Rendering) corresponding to the current sample image. The motion vector of the current sample image and the preset reference frame image corresponding to the current sample image are input into the differentiable motion compensation module for motion compensation to obtain the prediction image corresponding to each first sample image. Next, the residual image between the current sample image and the prediction image corresponding to the current sample image is input into the differentiable transform module for block transform processing to obtain the block transform information corresponding to the current sample image. Then, the block transform information corresponding to the current sample image and the sample block-level quantization parameters corresponding to the current sample image are input into the differentiable quantization module for block quantization to obtain the block quantization transform information corresponding to the current sample image. Finally, the block quantization transform information corresponding to the current sample image and the motion vector corresponding to the current sample image are input into the differentiable entropy encoding module for encoding processing to determine the current sample image. The corresponding block-level bitrate (optionally, the output of the differentiable entropy coding module can be the block-level bitrate directly, or it can be the probability distribution in the entropy coding process; correspondingly, the block-level bitrate can be determined by combining the probability distribution); in addition, the block quantization transformation information corresponding to the current sample image can be input into the differentiable inverse quantization module for block inverse quantization processing to obtain the inverse quantization transformation information corresponding to the current sample image; then, the inverse quantization transformation information is input into the differentiable inverse transformation module for block inverse transformation processing to obtain the restoration information corresponding to the current sample image; then, the fused image between the restoration information corresponding to the current sample image and the prediction image corresponding to the current sample image is input into the differentiable loop filtering module for loop filtering to obtain the reconstructed image corresponding to the current sample image; furthermore, the reconstructed image can be cached in the reconstructed image cache for encoding and decoding of subsequent frame images.

[0127] In one specific embodiment, both the encoding and decoding processes require the use of the predicted image. To save bitrate, the decoding and decoding processes can share a differentiable motion compensation module. Accordingly, during the decoding process, motion compensation can be performed by combining the motion vectors from the encoding process to obtain the predicted image again.

[0128] In an optional embodiment, the above method may further include: a training step for a predefined, self-directed codec, specifically, as follows: Figure 4 As shown, the following steps may be included:

[0129] S401: Obtain the fifth preset block-level quantization parameter corresponding to each second sample image in the second sample image set;

[0130] S403: Using the fifth preset block-level quantization parameter as the quantization parameter corresponding to each second sample image, perform encoding and decoding processing on each second sample image based on the second preset training model to obtain the fourth sample reconstructed image of each second sample image and the fourth sample block-level bitrate corresponding to each first sample image;

[0131] S405: Based on each second sample image, the reconstructed fourth sample image, and the fourth sample block-level bitrate, perform coding performance analysis to obtain the second coding performance data corresponding to each second sample image;

[0132] S407: Using the second coding performance data as the quantization coding loss of the second preset training model, train the second preset training model to obtain a preset differentiable codec.

[0133] In one specific embodiment, the second preset training model can be a pre-set deep learning model to be trained for encoding and decoding processing. The specific model structure can be found in the above-described pre-trained, trainable codec.

[0134] In one specific embodiment, the second sample image set can be a collection of multiple second sample images. Optionally, the second sample images can include the original second sample image or the edge image corresponding to the original second sample image (an image obtained after edge extraction from the original second sample image). The original second sample image can be a video frame image or an image acquired in an image format. Optionally, the fifth preset block-level quantization parameter corresponding to each second sample image can be a block-level quantization parameter determined by combining traditional block-level quantization algorithms, or it can be a block-level quantization parameter preset according to actual application requirements.

[0135] In practical applications, during the training process of the pre-defined guide codec, the corresponding training data (a large number of second sample images) can be obtained in advance, and the sample image set (i.e. the second sample image set) used for training the guide codec can be obtained from the training data in sequence.

[0136] In a specific embodiment, the detailed steps of using the fifth preset block-level quantization parameter as the quantization parameter corresponding to each second sample image, and performing encoding and decoding processing on each second sample image based on the second preset training model to obtain the fourth sample reconstructed image of each second sample image and the fourth sample block-level bitrate corresponding to each first sample image can be found in the detailed steps of using the sample block-level quantization parameter as the quantization parameter corresponding to each first sample image, and performing encoding and decoding processing on each first sample image based on the preset differential codec to obtain the first sample reconstructed image of each first sample image and the first sample block-level bitrate corresponding to each first sample image. These details will not be repeated here.

[0137] In a specific embodiment, the second coding performance data can characterize the coding performance of encoding the corresponding second sample image based on the current second preset training model; optionally, the smaller the coding performance data, the better the coding performance; conversely, the larger the coding performance data, the worse the coding performance.

[0138] In an optional embodiment, coding performance analysis is performed based on each second sample image, the fourth sample reconstructed image, and the fourth sample block-level bitrate. The second coding performance data corresponding to each second sample image can be obtained by combining the following formula:

[0139]

[0140] in, This represents the second coding performance data corresponding to the i-th second sample image; The default loss function is used. This represents the i-th second sample image; Represents the reconstructed image (fourth sample reconstructed image) of the i-th second sample image; This represents the pixel data loss caused by encoding and decoding the i-th second sample image based on the current second preset training model; This represents the block-level bitrate (fourth sample block-level bitrate) corresponding to the i-th second sample image. The preset hyperparameters for controlling the bitrate (can be adjusted according to actual application requirements); the preset loss function can be a subjective quality metric, MSE (Mean Square Error), SSIM (Structural Similarity), and some other deep learning model loss functions.

[0141] In one specific embodiment, the second encoding performance data is used as the quantization encoding loss of the second preset training model to train the second preset training model and obtain a preset differentiable codec.

[0142] In a specific embodiment, the quantization coding loss of the second preset training model using the second coding performance data can include using the second coding performance data (the sum of the second coding performance data corresponding to all second sample images in the second sample image set) as the quantization coding loss of the second preset training model. Specifically, the quantization coding loss of the second preset training model can characterize the coding performance of the second preset training model. Optionally, training the second preset training model using the second coding performance data as the quantization coding loss of the second preset training model to obtain a preset differentiable codec can include: minimizing the quantization coding loss using gradient descent to update the model parameters corresponding to the second preset training model; based on the updated model parameters, and combined with re-obtaining the current sample image set (second sample image set) used for model training from the training data, repeating the training iteration steps of S403, S405, and minimizing the quantization coding loss using gradient descent to update the model parameters corresponding to the second preset training model, until the second preset convergence condition is met, and using the second preset training model that meets the second preset convergence condition as the preset differentiable codec.

[0143] In a specific embodiment, the above-mentioned second preset convergence condition can be that the quantization coding loss is less than or equal to the second preset loss threshold, or the number of training iterations reaches the second preset number, etc. Specifically, the second preset loss threshold and the second preset number can be set in combination with the preset guideable codec accuracy and training speed requirements in actual applications.

[0144] In the above embodiments, by combining the sample images in the second sample image set, the reconstructed images obtained by encoding and decoding the sample images based on the second preset training model, and the block-level bitrate, coding performance analysis is performed to obtain second coding performance data that can be used as the quantization coding loss of the second preset training model. This enables the training of unsupervised, differentiated codecs, greatly reducing the training cost and time of differentiated codecs and improving the training efficiency of differentiated codecs.

[0145] S207: Perform coding performance analysis based on each first sample image, the first sample reconstructed image, and the first sample block-level bitrate to obtain the first coding performance data corresponding to each first sample image;

[0146] In an optional embodiment, the first coding performance data can characterize the coding performance corresponding to the quantization coding of the corresponding first sample image based on the block-level quantization parameters predicted by the current first preset training model; optionally, the smaller the coding performance data, the better the coding performance; conversely, the larger the coding performance data, the worse the coding performance.

[0147] In an optional embodiment, in the training scenario of the spatial quantization parameter prediction model, the above-mentioned spatial coding performance data can be determined by combining the following formula:

[0148]

[0149] in, This represents the spatial coding performance data corresponding to the i-th first sample image (characterizing the coding performance of the corresponding first sample image quantized based on the block-level quantization parameters predicted by the current first model to be trained); The default loss function is used. This represents the i-th first sample image; Represents the reconstructed image of the i-th first sample image (first sample reconstructed image); This represents the pixel data loss caused by quantizing and encoding the i-th first sample image based on the block-level quantization parameters predicted by the current first model to be trained. This represents the bitrate (first sample block-level bitrate) corresponding to the i-th first sample image; Preset hyperparameters for controlling bitrate (can be adjusted according to actual application requirements).

[0150] S209: Using the first encoding performance data as the quantization prediction loss of the first preset training model, train the first preset training model to obtain the quantization parameter prediction model.

[0151] In a specific embodiment, the quantization prediction loss of the first preset training model using the first coding performance data may include: using the first coding performance data (the sum of the first coding performance data corresponding to all first sample images in the first sample image set) as the quantization prediction loss of the first preset training model. Specifically, the quantization prediction loss of the first preset training model can characterize the prediction performance of the first preset training model in predicting block-level quantization parameters. Optionally, training the first preset training model using the first coding performance data as the quantization prediction loss to obtain a quantization parameter prediction model may include: minimizing the quantization prediction loss using gradient descent to update the model parameters corresponding to the first preset training model; based on the updated model parameters, and combined with re-obtaining the current sample image set (first sample image set) used for model training from the training data, repeating the training iteration steps of S203, S205, S207 and minimizing the quantization prediction loss using gradient descent to update the model parameters corresponding to the first preset training model, until the first preset convergence condition is met, and using the first preset training model that meets the first preset convergence condition as the quantization parameter prediction model.

[0152] In a specific embodiment, the aforementioned first preset convergence condition can be that the quantized prediction loss is less than or equal to the first preset loss threshold, or that the number of training iterations reaches the first preset number, etc. Specifically, the first preset loss threshold and the first preset number can be set in combination with the model accuracy and training speed requirements in actual applications.

[0153] In a specific embodiment, in the scenario of performing block-level quantization parameter prediction in the spatial domain, the first preset training model includes a first training model; the first coding performance data includes spatial coding performance data; and the aforementioned quantization parameter prediction model includes a first prediction model.

[0154] Accordingly, the above-mentioned quantization parameters corresponding to each first sample image are quantized using sample block-level quantization parameters, and each first sample image is encoded and decoded based on a preset doubly-guided codec to obtain the first sample reconstructed image and the first sample block-level bitrate corresponding to each first sample image, including: quantizing parameters corresponding to each first sample image using first block-level quantization parameters, and encoding and decoding each first sample image based on a preset doubly-guided codec to obtain the first sample reconstructed image and the first sample block-level bitrate;

[0155] Accordingly, the above-mentioned coding performance analysis based on each first sample image, the first sample reconstructed image, and the first sample block-level bitrate to obtain the spatial coding performance data corresponding to each first sample image may include: coding performance analysis based on each first sample image, the first sample reconstructed image, and the first sample block-level bitrate to obtain the spatial coding performance data corresponding to each first sample image;

[0156] The above-mentioned coding performance analysis based on each first sample image, the first sample reconstructed image, and the first sample block-level bitrate, to obtain the first coding performance data corresponding to each first sample image, may include: performing coding performance analysis based on each first sample image, the first sample reconstructed image, and the first sample block-level bitrate to obtain the spatial coding performance data corresponding to each first sample image;

[0157] Accordingly, the above-mentioned method of using the first coding performance data as the quantization prediction loss of the first preset training model to train the first preset training model and obtaining the quantization parameter prediction model includes: using the spatial coding performance data as the quantization prediction loss of the first training model to train the first training model and obtaining the first prediction model.

[0158] In one specific embodiment, the first prediction model can be a block-level quantization parameter prediction model for the spatial domain.

[0159] In a specific embodiment, the detailed steps of using the first block-level quantization parameter as the quantization parameter corresponding to each first sample image, and performing encoding and decoding processing on each first sample image based on a preset doubly-guided codec to obtain the first sample reconstructed image and the first sample block-level bitrate, can be found in the detailed steps of using the sample block-level quantization parameter as the quantization parameter corresponding to each first sample image, and performing encoding and decoding processing on each first sample image based on a preset doubly-guided codec to obtain the first sample reconstructed image and the first sample block-level bitrate corresponding to each first sample image. These details will not be repeated here.

[0160] In a specific embodiment, the detailed steps of using spatial coding performance data as the quantization prediction loss of the first training model to train the first training model and obtain the first prediction model can be found in the detailed steps of using the first coding performance data as the quantization prediction loss of the first preset training model to train the first preset training model and obtain the quantization parameter prediction model, which will not be repeated here.

[0161] In a specific embodiment, taking the scenario of training a spatial domain quantization parameter prediction model by combining a first preset block-level quantization parameter as an example, such as... Figure 5 As shown, Figure 5 This is a schematic diagram illustrating a training process for a first prediction model according to an exemplary embodiment. Specifically, the current sample image (first sample image) and the first preset block-level quantization parameters corresponding to the current sample image can be input into the first model to be trained to predict the block-level quantization parameters in the spatial domain, thereby obtaining the first block-level quantization parameters corresponding to the current sample image. Next, using the first block-level quantization parameters as the quantization parameters of the current sample image, encoding and decoding processing is performed in conjunction with a preset differentiable encoder to obtain the reconstructed image and block-level bitrate corresponding to the current sample image. Then, encoding performance analysis can be performed by combining all first sample images in the first sample image set, the reconstructed images corresponding to all first sample images, and the block-level bitrate to obtain spatial domain encoding performance data. Finally, the model parameters of the first model to be trained can be updated based on the spatial domain encoding performance data to train the first model to be trained, thereby obtaining the first prediction model.

[0162] In the above embodiments, during the training of the spatial domain quantization prediction model, a single sample image is input into the first model to be trained to predict the block-level quantization parameters in the spatial domain. The first block-level quantization parameters predicted by the first model to be trained are used as the quantization parameters of the corresponding sample image. Combined with a preset differentiable codec, encoding and decoding processing is performed to obtain the sample reconstruction image and bitrate corresponding to the sample image. This ensures that the first coding performance data obtained from the coding analysis based on the sample image, the sample reconstruction image, and the bitrate is differentiable and can be used as the quantization prediction loss of the first model to be trained for training the spatial domain block-level quantization parameter prediction model. Without the need for model labels, the effectiveness of the spatial domain quantization parameters can be effectively improved, block-level bitrate allocation can be achieved, coding performance can be improved, and model training time can be greatly reduced, improving model training efficiency and thus improving the system performance during the model training process.

[0163] In an optional embodiment, in the scenario of predicting block-level quantization parameters in the temporal domain, the first preset model to be trained includes a second model to be trained; the sample block-level quantization parameters may include second block-level quantization parameters, which are temporal block-level quantization parameters; the first sample image set includes at least one sample video frame image corresponding to a sample video; optionally, the method may further include:

[0164] Determine the first sample associated image corresponding to each first sample image;

[0165] Accordingly, based on the first preset model to be trained, the block-level quantization parameter prediction for each first sample image in the first sample image set, to obtain the sample block-level quantization parameters corresponding to each first sample image, may include:

[0166] Each first sample image and its corresponding first sample associated image are input into the second model to be trained to predict temporal block-level quantization parameters, thereby obtaining the second block-level quantization parameters corresponding to each first sample image.

[0167] In one specific embodiment, the first sample associated image includes at least one of the following: at least one adjacent frame image of each first sample image, a predicted frame image corresponding to at least one adjacent frame image, and a residual frame image corresponding to at least one adjacent frame image. Optionally, at least one adjacent frame image of each first sample image (sample video frame image) may include adjacent first M frame sample video frame images and / or adjacent last N frame sample video frame images; specifically, M and N are both positive integers greater than or equal to 1. The predicted frame image corresponding to any adjacent frame image can be the image after motion compensation of any adjacent frame image. Optionally, the reference frame image in the motion compensation process of each adjacent frame image can be selected according to the actual application. The residual frame image corresponding to any adjacent frame image can be the residual image between the adjacent frame image and the predicted frame image corresponding to the adjacent frame image.

[0168] In an optional embodiment, the output of the second model to be trained can be either the second block-level quantization parameter corresponding to each first sample image, or the adjustment parameters of the temporal block-level quantization algorithm corresponding to each first sample image. Accordingly, after obtaining the adjustment parameters of the temporal block-level quantization algorithm, the second block-level quantization parameter can be determined by combining the adjustment parameters of the temporal block-level quantization algorithm corresponding to each first sample image output by the second model to be trained. Optionally, traditional temporal block-level quantization algorithms may include, but are not limited to, MB-tree (macroblock tree) algorithms and CU-tree (Coding Unit tree) algorithms. Optionally, taking the CU-tree algorithm as an example, the adjustment parameters of the temporal block-level quantization algorithm may include: adjustable intensity values ​​in the CU-tree algorithm, the SATD cost between the predicted and original values ​​of intra-frame modes, and the SATD cost between the predicted and original values ​​of inter-frame modes, etc.

[0169] In the above embodiments, by using each first sample image (video frame image), and at least one of the adjacent frame images of the first sample image, the prediction frame image corresponding to the at least one adjacent frame image, and the residual frame image corresponding to the at least one adjacent frame image as input to the quantization parameter prediction model to be trained, the quantization parameter prediction model to be trained can learn the temporal relationship between each frame image and the adjacent frame images and then predict the quantization parameters. This enables temporal quantization parameter prediction, better block-level bitrate allocation, and improved efficiency and performance of subsequent image coding.

[0170] In an optional embodiment, the above-mentioned inputting each first sample image and the first sample associated image corresponding to each first sample image into the second model to be trained for temporal block-level quantization parameter prediction, to obtain the second block-level quantization parameters corresponding to each first sample image, may include:

[0171] Each first sample image, the second preset block-level quantization parameter corresponding to each first sample image, and the first sample associated image corresponding to each first sample image are input into the second model to be trained to predict the temporal block-level quantization parameter, thereby obtaining the second block-level quantization parameter.

[0172] In one specific embodiment, the aforementioned second preset block-level quantization parameter can be a time-domain block-level quantization parameter determined by combining traditional time-domain block-level quantization algorithms, or it can be a time-domain block-level quantization parameter preset according to actual application requirements. Optionally, the time-domain block-level quantization algorithm may include the MB-tree algorithm, the CU-tree algorithm, etc.

[0173] In the above embodiments, each first sample image, the second preset block-level quantization parameter corresponding to each first sample image, and the first sample associated image corresponding to each first sample image are used as inputs to the second model to be trained. This allows the first model to be trained to perform temporal block-level quantization parameter correction and adjustment based on the second preset block-level quantization parameter, thereby improving the prediction accuracy of temporal block-level quantization parameters.

[0174] In an optional embodiment, if the first sample image is the original image of the first sample, the edge image corresponding to the original image of the first sample can also be used as the input of the second model to be trained; if the first sample image is the edge image corresponding to the first sample image set, the original image of the first sample can also be used as the input of the second model to be trained, so that the second model to be trained can better learn image features during the prediction of block-level quantization parameters in the temporal domain, thereby improving the prediction accuracy of block-level quantization parameters in the temporal domain.

[0175] In an optional embodiment, in the scenario of predicting block-level quantization parameters in the time domain, the first coding performance data includes time-domain coding performance data; the first preset model to be trained includes a second model to be trained; the quantization parameter prediction model includes a second prediction model; optionally, the method may further include:

[0176] The first sample reconstructed image corresponding to each first sample image is used as the reference frame image of the first sample associated image corresponding to each first sample image. The third preset block-level quantization parameter corresponding to the first sample associated image is used as the quantization parameter corresponding to the first sample associated image. The first sample associated image is encoded and decoded based on the preset doubly oriented codec to obtain the second sample reconstructed image of the first sample associated image and the second sample block-level bitrate corresponding to the first sample associated image.

[0177] Accordingly, the above-mentioned coding performance analysis based on each first sample image, the first sample reconstructed image, and the first sample block-level bitrate, to obtain the first coding performance data corresponding to each first sample image, may include: performing coding performance analysis based on each first sample image, the first sample associated image, the first sample reconstructed image, the second sample reconstructed image, the first sample block-level bitrate, and the second sample block-level bitrate to obtain temporal coding performance data;

[0178] Accordingly, the above-mentioned method of using the first coding performance data as the quantization prediction loss of the first preset training model to train the first preset training model and obtaining the quantization parameter prediction model may include: using the temporal coding performance data as the quantization prediction loss of the second training model to train the second training model and obtaining the second prediction model.

[0179] In one specific embodiment, the second prediction model can be a time-domain block-level quantization parameter prediction model.

[0180] In a specific embodiment, the first sample reconstructed image corresponding to each first sample image is used as the reference frame image of the first sample associated image corresponding to each first sample image. The third preset block-level quantization parameter corresponding to the first sample associated image is used as the quantization parameter corresponding to the first sample associated image. The first sample associated image is encoded and decoded based on a preset doubly oriented codec to obtain the second sample reconstructed image of the first sample associated image and the detailed refinement of the second sample block-level bitrate corresponding to the first sample associated image. For details, please refer to the above-mentioned detailed refinement of the first sample reconstructed image of each first sample image and the first sample block-level bitrate corresponding to each first sample image by using the sample block-level quantization parameter as the quantization parameter corresponding to each first sample image and encoding and decoding based on a preset doubly oriented codec. This will not be repeated here.

[0181] In a specific embodiment, coding performance analysis is performed based on each first sample image, first sample associated image, first sample reconstructed image, second sample reconstructed image, first sample block-level code rate, and second sample block-level code rate. The resulting temporal coding performance data can be obtained by combining the following formula:

[0182]

[0183] in, This represents the temporal coding performance data corresponding to the t-th frame sample image (a certain first sample image), which characterizes the coding performance of quantizing the corresponding frame sample image based on the block-level quantization parameters predicted by the current second training model. The default loss function is used. This represents the i-th frame sample image (taking the adjacent frame image of the t-th frame sample image as an example, where the first sample is associated with the sample image of the t-th frame sample image); Represents the reconstructed image (second sample reconstructed image) of the i-th frame sample image; This represents the pixel data loss caused by quantizing and encoding the i-th first sample image based on the block-level quantization parameters predicted by the current second training model. The preset weights represent the pixel data loss corresponding to the i-th frame sample image (which can be set according to actual application requirements). This represents the sample block-level bitrate corresponding to the i-th frame sample image; The preset weights represent the sample block-level bitrates corresponding to the i-th frame sample image (which can be set according to actual application requirements). Preset hyperparameters for controlling bitrate (can be adjusted according to actual application requirements).

[0184] In a specific embodiment, the detailed steps for training the second training model using temporal coding performance data as the quantization prediction loss can be found in the above-described detailed steps for training the first preset training model using the first coding performance data as the quantization prediction loss, and obtaining the quantization parameter prediction model. These steps will not be repeated here. However, it should be noted that during the training of the second training model, the temporal coding performance data is obtained through coding performance analysis based on each first sample image in the first sample image set, the first sample reconstructed image of each first sample image, the first sample associated image corresponding to each first sample image, the second sample reconstructed image of the first sample associated image, the first sample block-level code rate corresponding to each first sample image, and the second sample block-level code rate corresponding to the first sample associated image. Correspondingly, the training iteration steps also include the aforementioned steps of obtaining the first sample associated image, obtaining the second sample reconstructed image of the first sample associated image, and obtaining the second sample block-level code rate corresponding to the first sample associated image.

[0185] In the above embodiments, during the training of the temporal quantization prediction model, each first sample image (video frame image), along with at least one first sample associated image from at least one of its neighboring frame images, the prediction frame image corresponding to at least one neighboring frame image, and the residual frame image corresponding to at least one neighboring frame image, are used together as input to the quantization parameter prediction model to be trained. This allows the second model to be trained to perform block-level quantization parameter prediction based on multi-frame correlation (temporal correlation). Furthermore, by combining the second block-level quantization parameters predicted by the first model to be trained, and the third preset block-level quantization parameters corresponding to the first sample associated image, a preset guideable codec is used to process each first sample image... By performing encoding and decoding processing on the image associated with the first sample, the corresponding reconstructed image and bitrate are obtained. This ensures that the temporal coding performance data obtained from encoding analysis based on each first sample image, first sample associated image, first sample reconstructed image, second sample reconstructed image, first sample block-level bitrate, and second sample block-level bitrate is differentiable. This data can be used as the quantization prediction loss for the second model to be trained, enabling temporal block-level quantization parameter prediction model training. Without the need for model labels, the effectiveness of temporal quantization parameters can be effectively improved, block-level bitrate allocation can be achieved, coding performance can be improved, and model training time can be greatly reduced, thus improving model training efficiency and ultimately enhancing system performance during model training.

[0186] In an optional embodiment, each first sample image and its corresponding first sample associated image are input into a second model to be trained for temporal block-level quantization parameter prediction, resulting in second block-level quantization parameters for each first sample image, including:

[0187] Each first sample image, the second preset block-level quantization parameter corresponding to each first sample image, the first sample associated image corresponding to each first sample image, and the third preset block-level quantization parameter corresponding to the first sample associated image are input into the second training model to perform block-level quantization parameter prediction, so as to obtain the second block-level quantization parameter and the third block-level quantization parameter corresponding to the first sample associated image.

[0188] Accordingly, the reference frame image of the first sample associated image corresponding to each first sample image is the first sample reconstructed image corresponding to each first sample image. The third preset block-level quantization parameter corresponding to the first sample associated image is used as the quantization parameter corresponding to the first sample associated image. The first sample associated image is encoded and decoded based on a preset doubly-controllable codec to obtain the second sample reconstructed image of the first sample associated image and the second sample block-level bitrate corresponding to the first sample associated image, including:

[0189] The first sample reconstructed image corresponding to each first sample image is used as the reference frame image of the first sample associated image corresponding to each first sample image. The third block-level quantization parameter corresponding to the first sample associated image is used as the quantization parameter corresponding to the first sample associated image. The first sample associated image is encoded and decoded based on a preset guided codec to obtain the second sample reconstructed image and the second sample block-level bitrate.

[0190] In the above embodiments, combining the second network to be trained to predict the block-level quantization parameters of the sample association image can better improve the quantization coding performance in the encoding and decoding process of the sample association image, and thus better improve the model's prediction accuracy of the temporal quantization parameters, better perform block-level bitrate allocation, and improve coding performance.

[0191] In a specific embodiment, such as Figure 6 As shown, Figure 6 This is a schematic diagram of a second prediction model training process provided according to an exemplary embodiment. Optionally, taking the first sample associated image as an example of adjacent frame images (M frames before and N frames after), the current sample image (first sample image), the second preset block-level quantization parameters corresponding to the current sample image, and the adjacent frame images corresponding to the previous sample image can be input into the second model to be trained for temporal block-level quantization parameter prediction to obtain the second block-level quantization parameters corresponding to the current sample image. Then, the second block-level quantization parameters are used as the quantization parameters of the current sample image, and the third preset block-level quantization parameters corresponding to the adjacent frame images are used as the quantization parameters of the adjacent frame images. Encoding and decoding processing is performed in combination with a preset differentiable encoder to obtain the reconstructed image and block-level bitrate corresponding to the current sample image, as well as the reconstructed image and block-level bitrate corresponding to the adjacent frame images. Next, coding performance analysis can be performed by combining the current sample image, the reconstructed image and block-level bitrate corresponding to the current sample image, the adjacent frame images, and the reconstructed image and block-level bitrate corresponding to the adjacent frame images to obtain the temporal coding performance data corresponding to the current sample image. Then, the model parameters of the second model to be trained are updated by combining the temporal coding performance data corresponding to all the first sample images in the first sample image set to train the second model to obtain the second prediction model.

[0192] In a specific embodiment, a model capable of joint spatiotemporal block-level quantization parameter prediction can be trained; correspondingly, the aforementioned sample block-level quantization parameters include a fourth block-level quantization parameter and a fifth block-level quantization parameter, wherein the fourth block-level quantization parameter is a spatial block-level quantization parameter and the fifth block-level quantization parameter is a temporal block-level quantization parameter; the first preset model to be trained includes a third model to be trained; the first sample image set includes at least one sample video frame image corresponding to a sample video; correspondingly, the above method may further include:

[0193] Determine the second sample associated image corresponding to each first sample image; the second sample associated image includes at least one of the following: at least one neighboring frame image of each first sample image, at least one prediction frame image corresponding to at least one neighboring frame image, and at least one residual frame image corresponding to at least one neighboring frame image;

[0194] Accordingly, based on the first preset model to be trained, the block-level quantization parameters of each first sample image in the first sample image set are predicted to obtain the sample block-level quantization parameters corresponding to each first sample image, including:

[0195] Each first sample image and the corresponding second sample associated image are input into the third model to be trained for spatiotemporal joint block-level quantization parameter prediction, resulting in the fourth block-level quantization parameter and the fifth block-level quantization parameter corresponding to each first sample image.

[0196] Furthermore, the aforementioned first coding performance data includes joint coding performance data; the quantization parameter prediction model includes a third prediction model; correspondingly, the aforementioned method may also include:

[0197] The first sample reconstructed image corresponding to each first sample image is used as the reference frame image of the second sample associated image corresponding to each first sample image. The fourth preset block-level quantization parameter corresponding to the second sample associated image is used as the quantization parameter corresponding to the second sample associated image. The second sample associated image is encoded and decoded based on the preset guide codec to obtain the third sample reconstructed image of the second sample associated image and the third sample block-level bitrate corresponding to the second sample associated image.

[0198] Accordingly, the above-mentioned coding performance analysis based on each first sample image, the first sample reconstructed image, and the first sample block-level bitrate, to obtain the first coding performance data corresponding to each first sample image, includes: coding performance analysis based on each first sample image, the second sample associated image, the first sample reconstructed image, the third sample reconstructed image, the first sample block-level bitrate, and the third sample block-level bitrate, to obtain joint coding performance data;

[0199] Accordingly, the above-mentioned method of using the first coding performance data as the quantization prediction loss of the first preset training model to train the first preset training model and obtaining the quantization parameter prediction model includes: using the joint coding performance data as the quantization prediction loss of the third training model to train the third training model and obtaining the third prediction model.

[0200] In a specific embodiment, the third prediction model can be a spatiotemporally joint block-level quantization parameter prediction model. The model input during the training process of the third prediction model can refer to the model input during the training process of the second prediction model in the temporal domain, but the model output of the third prediction model needs to include spatial domain block-level quantization parameters. Specifically, the detailed steps in the training process of the third prediction model can be found in the detailed steps of the training processes of the first prediction model in the spatial domain and the second prediction model in the temporal domain, as described above, and will not be repeated here.

[0201] In the above embodiments, during the training of the spatiotemporal joint quantization prediction model, each first sample image (video frame image), along with at least one second sample associated image from at least one of the first sample image's neighboring frame images, the prediction frame image corresponding to at least one neighboring frame image, and the residual frame image corresponding to at least one neighboring frame image, are used as input to the third model to be trained. This allows the third model to be trained to perform block-level quantization parameter prediction based on considering inter-frame correlation (temporal correlation) and spatial correlation. Furthermore, a preset, self-contained codec is used to encode and decode each first sample image and the second sample associated image to obtain the corresponding... By reconstructing the image and bitrate, the spatiotemporal joint coding performance data obtained from coding analysis based on the first sample image, the associated second sample image, the reconstructed first sample image, the reconstructed fourth sample image, and the third sample block-level bitrate is differentiable. This data can be used as the quantization prediction loss for the third model to be trained, enabling the training of the spatiotemporal block-level quantization parameter prediction model. Without requiring model labels, this effectively improves the effectiveness of the spatiotemporal joint block-level quantization parameters, achieves spatiotemporal joint block-level bitrate allocation, improves coding performance, and significantly reduces model training time, thereby improving model training efficiency and ultimately enhancing system performance during model training.

[0202] In an optional embodiment, the quantization parameters used in the encoding and decoding process of each first sample image in conjunction with a preset guideable codec can also be fused block-level quantization parameters. Specifically, they can be obtained by weighted fusion of the sample block-level quantization parameters corresponding to each first sample image and the traditional block-level quantization parameters (block-level quantization parameters determined based on traditional block-level quantization algorithms) corresponding to the first sample image. Optionally, the weights corresponding to the sample block-level quantization parameters and the traditional block-level quantization parameters can be adjusted according to actual application requirements. Optionally, when the sample block-level quantization parameters include spatial domain block-level quantization parameters, the traditional block-level quantization parameters can be block-level quantization parameters determined based on traditional spatial domain block-level quantization algorithms. Optionally, when the sample block-level quantization parameters include temporal domain block-level quantization parameters, the traditional block-level quantization parameters can be block-level quantization parameters determined based on traditional temporal domain block-level quantization algorithms.

[0203] As can be seen from the technical solutions provided in the embodiments of this specification above, in the training process of the quantization parameter prediction model in this application embodiment, firstly, based on the first preset model to be trained, block-level quantization parameters are predicted for each first sample image in the first sample image set to obtain the sample block-level quantization parameters corresponding to each first sample image; then, using the sample block-level quantization parameters as the quantization parameters corresponding to each first sample image, each first sample image is encoded and decoded based on a preset doubly compliant codec to obtain the first sample reconstructed image for encoding performance analysis and the first sample block-level bitrate corresponding to each first sample image; and since the first sample reconstructed image and the first sample block-level bitrate are based on the preset doubly compliant codec... The first coding performance data obtained by the instrument can guarantee that the first coding performance data obtained by coding analysis based on the first sample image, the first sample reconstructed image, and the first sample block-level bitrate is differentiable. It can be used as the quantization prediction loss of the first preset training model to train the block-level quantization parameter prediction model. This can achieve unsupervised quantization parameter prediction model training without obtaining model labels, which greatly reduces the model training time and improves the model training efficiency, thereby also improving the system performance during the model training process. Furthermore, combining the first coding performance data with the block-level quantization parameter prediction training of the first preset training model can effectively improve the effectiveness of quantization parameters, realize block-level bitrate allocation, and improve coding performance and coding efficiency.

[0204] Based on the first prediction model trained according to the above embodiments of this application, the following describes an image encoding method of this application. Optionally, this method can be applied to electronic devices such as servers and terminals. Figure 7 As shown, the following steps may be included:

[0205] S701: Acquire at least one first image to be encoded;

[0206] S703: Input each first image to be encoded into the first prediction model to predict the spatial block-level quantization parameters, and obtain the sixth block-level quantization parameters corresponding to each first image to be encoded.

[0207] S705: Based on the sixth block-level quantization parameters, each first image to be encoded is quantized and encoded to obtain the image encoding information corresponding to each first image to be encoded.

[0208] In a specific embodiment, the first prediction model is a quantized parameter prediction model obtained by training the first training model with spatial domain coding performance data as the quantized prediction loss of the first training model. The spatial domain coding performance data is obtained by analyzing the coding performance of each first sample image in the first sample image set, the first sample reconstructed image of each first sample image, and the first sample block-level code rate corresponding to each first sample image. The first sample reconstructed image and the first sample block-level code rate are obtained by encoding and decoding each first sample image using a preset doubly oriented codec.

[0209] In one specific embodiment, at least one first image to be encoded may include at least one original image to be encoded or an edge image corresponding to at least one original image to be encoded; optionally, the original image to be encoded may be a video frame image corresponding to the video to be encoded or an image to be encoded acquired in an image format.

[0210] In a specific embodiment, each first image to be encoded is input into the first prediction model to predict spatial block-level quantization parameters, and the detailed refinement of the sixth block-level quantization parameters corresponding to each first image to be encoded can be found in the above-described detailed refinement of each first sample image being input into the first training model to predict spatial block-level quantization parameters, and the first block-level quantization parameters corresponding to each first sample image, which will not be repeated here.

[0211] In an optional embodiment, the preset block-level quantization parameters corresponding to the first image to be encoded can also be used as input to the first prediction model. This allows the first prediction model to adjust the block-level quantization parameters based on the preset block-level quantization parameters corresponding to the first image to be encoded, thereby improving the accuracy of spatial block-level quantization parameter prediction. In a specific embodiment, the preset block-level quantization parameters corresponding to the first image to be encoded can be determined based on a traditional spatial block-level quantization algorithm.

[0212] In an optional embodiment, when the first image to be encoded is the original image to be encoded, the edge image corresponding to the original image to be encoded can also be used as the input of the first prediction model; when the first image to be encoded is the edge image corresponding to the original image to be encoded, the original image to be encoded can also be used as the input of the first prediction model, so that the first prediction model can better learn image features during the prediction of block-level quantization parameters in the spatial domain, thereby improving the prediction accuracy of block-level quantization parameters in the spatial domain.

[0213] In a specific embodiment, the above-mentioned quantization encoding of each first image to be encoded based on the sixth block-level quantization parameter to obtain the image encoding information corresponding to each first image to be encoded may include: quantizing the block transformation information corresponding to each first image to be encoded according to the sixth block-level quantization parameter to obtain quantization transformation information; and performing entropy encoding on each first image to be encoded based on the quantization transformation information to obtain the image encoding information corresponding to each first image to be encoded.

[0214] As can be seen from the technical solutions provided in the embodiments of this specification above, in the process of image quantization coding, combining the block-level quantization parameter prediction model (first prediction model) in the spatial domain to determine the block-level quantization parameters can effectively improve the effectiveness of quantization parameters, realize block-level bitrate allocation, and improve the coding performance and coding efficiency of multimedia data such as images and videos.

[0215] The following describes an image encoding method based on the second prediction model trained according to the above embodiments of this application. Optionally, this method can be applied to electronic devices such as servers and terminals. Figure 8 As shown, the following steps may be included:

[0216] S801: Obtain at least one second image to be encoded corresponding to the first video to be encoded;

[0217] S803: Determine the first target association image corresponding to each second image to be encoded; the first target association image includes at least one of the following: at least one neighboring frame image of each second image to be encoded, a prediction frame image corresponding to at least one neighboring frame image, and a residual frame image corresponding to at least one neighboring frame image;

[0218] S805: Input each second image to be encoded and the first target associated image corresponding to each second image to be encoded into the second prediction model to predict the temporal block-level quantization parameters, and obtain the seventh block-level quantization parameters corresponding to each second image to be encoded.

[0219] S807: Based on the seventh block-level quantization parameters, each second image to be encoded is quantized and encoded to obtain the image encoding information corresponding to each second image to be encoded.

[0220] In one specific embodiment, the second prediction model is a quantized parameter prediction model obtained by training the second training model using temporal coding performance data as the quantized prediction loss of the second training model. The temporal coding performance data is obtained by analyzing the coding performance of each first sample image, the first sample reconstructed image of each first sample image, the first sample associated image of each first sample image, the second sample reconstructed image of the first sample associated image, the first sample block-level code rate of each first sample image, and the second sample block-level code rate of the first sample associated image in the first sample image set. The first sample reconstructed image, the first sample block-level code rate, the second sample reconstructed image, and the second sample block-level code rate are obtained by encoding and decoding each first sample image and the first sample associated image using a preset guided codec.

[0221] In one specific embodiment, at least one second image to be encoded corresponding to the first video to be encoded may include at least one first original video frame image in the first video to be encoded or an edge image corresponding to at least one first original video frame image.

[0222] In a specific embodiment, the detailed determination of the first target associated image corresponding to each second image to be encoded can be found in the above-described detailed determination of the first sample associated image corresponding to each first sample image, which will not be repeated here.

[0223] In a specific embodiment, each second image to be encoded and the first target associated image corresponding to each second image to be encoded are input into the second prediction model to predict temporal block-level quantization parameters, and the specific refinement of the seventh block-level quantization parameters corresponding to each second image to be encoded can be found in the above-described process of inputting each first sample image and the first sample associated image corresponding to each first sample image into the second training model to predict temporal block-level quantization parameters, and the specific refinement of the second block-level quantization parameters corresponding to each first sample image, which will not be repeated here.

[0224] In an optional embodiment, the preset block-level quantization parameters corresponding to the second image to be encoded can also be used as input to the second prediction model. This allows the second prediction model to adjust the block-level quantization parameters based on the preset block-level quantization parameters corresponding to the second image to be encoded, thereby improving the prediction accuracy of the block-level quantization parameters in the temporal domain. In a specific embodiment, the preset block-level quantization parameters corresponding to the second image to be encoded can be determined based on a traditional temporal block-level quantization algorithm.

[0225] In addition, it should be noted that for other related refinements of the second prediction model for time-domain block-level quantization parameter prediction, please refer to the above-mentioned refinements of the second model to be trained for time-domain block-level quantization parameter prediction, which will not be repeated here.

[0226] In a specific embodiment, the above-mentioned quantization encoding of each second image to be encoded based on the seventh block-level quantization parameter to obtain the image encoding information corresponding to each second image to be encoded may include: quantizing the block transformation information corresponding to each second image to be encoded according to the seventh block-level quantization parameter to obtain quantization transformation information; and performing entropy encoding on each second image to be encoded based on the quantization transformation information to obtain the image encoding information corresponding to each second image to be encoded.

[0227] As can be seen from the technical solutions provided in the embodiments of this specification above, in the process of image quantization coding, combining the temporal block-level quantization parameter prediction model (second prediction model) to determine the block-level quantization parameters can effectively improve the effectiveness of quantization parameters, realize block-level bitrate allocation, and improve video coding performance and coding efficiency.

[0228] Based on the third prediction model trained according to the above embodiments of this application, the following describes an image encoding method of this application. Optionally, this method can be applied to electronic devices such as servers and terminals. Figure 9 As shown, the following steps may be included:

[0229] S901: Obtain at least one third image to be encoded corresponding to the second video to be encoded;

[0230] S903: Determine the second target association image corresponding to each third image to be encoded; the second target association image includes at least one of the following: at least one neighboring frame image of each third image to be encoded, at least one prediction frame image corresponding to at least one neighboring frame image, and at least one residual frame image corresponding to at least one neighboring frame image;

[0231] S905: Input each third image to be encoded and the second target associated image corresponding to each third image to be encoded into the third prediction model to perform spatiotemporal joint block-level quantization parameter prediction, and obtain the eighth block-level quantization parameter and the ninth block-level quantization parameter corresponding to each third image to be encoded.

[0232] S907: Based on the eighth block-level quantization parameter or the ninth block-level quantization parameter, quantize and encode each third image to be encoded to obtain the image encoding information corresponding to each third image to be encoded.

[0233] In a specific embodiment, the third prediction model is a quantized parameter prediction model obtained by training the third training model using joint coding performance data as the quantized prediction loss of the third training model. The joint coding performance data is obtained by analyzing the coding performance of each first sample image in the first sample image set, the first sample reconstructed image of each first sample image, the second sample associated image corresponding to each first sample image, the third sample reconstructed image of the second sample associated image, the first sample block-level code rate corresponding to each first sample image, and the third sample block-level code rate corresponding to the second sample associated image. The first sample reconstructed image, the first sample block-level code rate, the point-multiple third sample reconstructed image, and the third sample block-level code rate are obtained by encoding and decoding each first sample image and the second sample associated image using a preset guided codec.

[0234] In one specific embodiment, at least one third image to be encoded corresponding to the second video to be encoded may include at least one second original video frame image or an edge image corresponding to at least one second original video frame image in the second video to be encoded.

[0235] In a specific embodiment, the detailed determination of the second target associated image corresponding to each third image to be encoded can be referred to the detailed determination of the first sample associated image corresponding to each first sample image described above, and will not be repeated here.

[0236] In a specific embodiment, the detailed steps for predicting the joint block-level quantization parameters in the spatiotemporal domain using the third prediction model are described above, and will not be repeated here.

[0237] In a specific embodiment, the detailed process of quantizing and encoding each third image to be encoded based on the eighth or ninth block-level quantization parameters to obtain the image encoding information corresponding to each third image to be encoded can be found in the above description, and will not be repeated here.

[0238] As can be seen from the technical solutions provided in the embodiments of this specification above, in the process of image quantization coding, combining the spatiotemporal joint block-level quantization parameter prediction model (second prediction model) to determine the block-level quantization parameters can improve the diversity and comprehensiveness of quantization parameter prediction, and can effectively improve the effectiveness of quantization parameters, realize block-level bitrate allocation, and improve coding performance and coding efficiency.

[0239] Furthermore, it should be noted that the quantization parameters in the above image encoding process can also be the block-level quantization parameters obtained by weighted fusion of the block-level quantization parameters predicted based on the corresponding quantization parameter prediction model and the traditional block-level quantization parameters determined based on the corresponding traditional block-level quantization algorithm.

[0240] Figure 10 This is a block diagram illustrating a quantization parameter prediction model training device according to an exemplary embodiment. (Refer to...) Figure 10 The device includes:

[0241] The first sample image set acquisition module 1010 is configured to acquire the first sample image set.

[0242] The first quantization parameter prediction module 1020 is configured to perform block-level quantization parameter prediction on each first sample image in the first sample image set based on the first preset training model, so as to obtain the sample block-level quantization parameter corresponding to each first sample image.

[0243] The first encoding and decoding processing module 1030 is configured to execute the quantization parameters corresponding to each first sample image with sample block-level quantization parameters, and perform encoding and decoding processing on each first sample image based on a preset doubly oriented codec to obtain the first sample reconstructed image of each first sample image and the first sample block-level bitrate corresponding to each first sample image.

[0244] The first coding performance analysis module 1040 is configured to perform coding performance analysis based on each first sample image, the first sample reconstructed image and the first sample block-level bitrate, and obtain the first coding performance data corresponding to each first sample image.

[0245] The first model training module 1050 is configured to execute a quantization prediction loss using the first encoding performance data as the first preset model to be trained, and to train the first preset model to be trained to obtain a quantization parameter prediction model.

[0246] In an optional embodiment, the sample block-level quantization parameters include a first block-level quantization parameter, which is a spatial block-level quantization parameter; the first preset training model includes a first training model; the first coding performance data includes spatial coding performance data; and the quantization parameter prediction model includes a first prediction model.

[0247] The first quantization parameter prediction module 1020 is specifically configured to perform spatial block-level quantization parameter prediction by inputting each first sample image into the first model to be trained, so as to obtain the first block-level quantization parameter corresponding to each first sample image.

[0248] The first encoding and decoding processing module 1030 is specifically configured to execute the quantization parameters corresponding to each first sample image with the first block-level quantization parameters, and perform encoding and decoding processing on each first sample image based on a preset doubly oriented codec to obtain the first sample reconstructed image and the first sample block-level bitrate.

[0249] The first coding performance analysis module 1040 is specifically configured to perform coding performance analysis based on each first sample image, the first sample reconstructed image and the first sample block-level bitrate, and obtain the spatial coding performance data corresponding to each first sample image;

[0250] The first model training module 1050 is specifically configured to execute a quantized prediction loss using spatial coding performance data as the first model to be trained, and train the first model to be trained to obtain the first prediction model.

[0251] In an optional embodiment, the first quantization parameter prediction module 1020 is specifically configured to perform spatial domain block-level quantization parameter prediction by inputting each first sample image and the first preset block-level quantization parameter corresponding to each first sample image into the first model to be trained, thereby obtaining the first block-level quantization parameter.

[0252] In an optional embodiment, the sample block-level quantization parameters include second block-level quantization parameters, which are time-domain block-level quantization parameters; the first preset training model includes a second training model; the first sample image set includes at least one sample video frame image corresponding to a sample video; the above apparatus further includes:

[0253] The first sample associated image determination module is configured to determine the first sample associated image corresponding to each first sample image; the first sample associated image includes at least one of the following: at least one adjacent frame image of each first sample image, at least one prediction frame image corresponding to at least one adjacent frame image, and at least one residual frame image corresponding to at least one adjacent frame image.

[0254] The first quantization parameter prediction module 1020 is specifically configured to perform temporal block-level quantization parameter prediction by inputting each first sample image and the first sample associated image corresponding to each first sample image into the second model to be trained, thereby obtaining the second block-level quantization parameter corresponding to each first sample image.

[0255] In an optional embodiment, the first quantization parameter prediction module 1020 is specifically configured to perform temporal block-level quantization parameter prediction by inputting each first sample image, the second preset block-level quantization parameter corresponding to each first sample image, and the first sample associated image corresponding to each first sample image into the second model to be trained, so as to obtain the second block-level quantization parameter.

[0256] In an optional embodiment, the first coding performance data includes temporal coding performance data; the quantization parameter prediction model includes a second prediction model; the above apparatus further includes:

[0257] The second encoding and decoding processing module is configured to execute a reference frame image of the first sample associated image corresponding to each first sample image, using the first sample reconstructed image corresponding to each first sample image as the reference frame image of the first sample associated image, using the third preset block-level quantization parameter corresponding to the first sample associated image as the quantization parameter corresponding to the first sample associated image, and performing encoding and decoding processing on the first sample associated image based on a preset doubly oriented codec to obtain the second sample reconstructed image of the first sample associated image and the second sample block-level bitrate corresponding to the first sample associated image.

[0258] The first coding performance analysis module 1040 is specifically configured to perform coding performance analysis based on each first sample image, first sample associated image, first sample reconstructed image, second sample reconstructed image, first sample block-level code rate and second sample block-level code rate to obtain temporal coding performance data.

[0259] The first model training module 1050 is specifically configured to execute the quantization prediction loss of the second model to be trained using time-domain coding performance data, and train the second model to obtain the second prediction model.

[0260] In an optional embodiment, the first quantization parameter prediction module 1020 is specifically configured to perform block-level quantization parameter prediction by inputting each first sample image, the second preset block-level quantization parameter corresponding to each first sample image, the first sample associated image corresponding to each first sample image, and the third preset block-level quantization parameter corresponding to the first sample associated image into the second model to be trained, so as to obtain the second block-level quantization parameter and the third block-level quantization parameter corresponding to the first sample associated image.

[0261] The second encoding / decoding processing module is specifically configured to execute a reference frame image for each first sample image, using the first sample reconstructed image corresponding to each first sample image as the reference frame image for each first sample image, using the third block-level quantization parameter corresponding to the first sample associated image as the quantization parameter corresponding to the first sample associated image, and performing encoding / decoding processing on the first sample associated image based on a preset doubly oriented codec to obtain the second sample reconstructed image and the second sample block-level bitrate.

[0262] In an optional embodiment, the sample block-level quantization parameters include a fourth block-level quantization parameter and a fifth block-level quantization parameter, wherein the fourth block-level quantization parameter is a spatial domain block-level quantization parameter and the fifth block-level quantization parameter is a temporal domain block-level quantization parameter; the first preset training model includes a third training model; the first sample image set includes at least one sample video frame image corresponding to a sample video; the above apparatus further includes:

[0263] The second sample association image determination module is configured to determine the second sample association image corresponding to each first sample image; the second sample association image includes at least one of the following: at least one adjacent frame image of each first sample image, at least one prediction frame image corresponding to at least one adjacent frame image, and at least one residual frame image corresponding to at least one adjacent frame image.

[0264] The first quantization parameter prediction module 1020 is specifically configured to perform spatiotemporal joint block-level quantization parameter prediction by inputting each first sample image and the corresponding second sample associated image into the third model to be trained, thereby obtaining the fourth block-level quantization parameter and the fifth block-level quantization parameter corresponding to each first sample image.

[0265] In an optional embodiment, the first coding performance data includes joint coding performance data; the quantization parameter prediction model includes a third prediction model; the above apparatus further includes:

[0266] The third encoding and decoding processing module is configured to execute a reference frame image of the second sample associated image corresponding to each first sample image, using the first sample reconstructed image corresponding to each first sample image as the reference frame image of the second sample associated image, using the fourth preset block-level quantization parameter corresponding to the second sample associated image as the quantization parameter corresponding to the second sample associated image, and performing encoding and decoding processing on the second sample associated image based on a preset doubly-controllable codec to obtain the third sample reconstructed image of the second sample associated image and the third sample block-level bitrate corresponding to the second sample associated image.

[0267] The first coding performance analysis module 1040 is specifically configured to perform coding performance analysis based on each first sample image, second sample associated image, first sample reconstructed image, third sample reconstructed image, first sample block-level bitrate, and third sample block-level bitrate to obtain joint coding performance data;

[0268] The first model training module 1050 is specifically configured to execute a quantized prediction loss using joint coding performance data as the third model to be trained, and train the third model to obtain the third prediction model.

[0269] In an optional embodiment, the above-described apparatus further includes:

[0270] The data acquisition module is configured to acquire the fifth preset block-level quantization parameters corresponding to each second sample image in the second sample image set;

[0271] The fourth encoding and decoding processing module is configured to execute the quantization parameters corresponding to each second sample image with the fifth preset block-level quantization parameters, and to perform encoding and decoding processing on each second sample image based on the second preset training model to obtain the fourth sample reconstructed image of each second sample image and the fourth sample block-level bitrate corresponding to each first sample image.

[0272] The second coding performance analysis module is configured to perform coding performance analysis based on each second sample image, the fourth sample reconstructed image and the fourth sample block-level bitrate, and obtain the second coding performance data corresponding to each second sample image.

[0273] The second model training module is configured to perform quantization encoding loss on the second preset training model using the second encoding performance data, and train the second preset training model to obtain a preset differentiable codec.

[0274] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0275] Figure 11 This is a block diagram of an image encoding apparatus according to an exemplary embodiment. (Refer to...) Figure 11 The device includes:

[0276] The first image to be encoded acquisition module 1110 is configured to acquire at least one first image to be encoded.

[0277] The second quantization parameter prediction module 1120 is configured to perform spatial block-level quantization parameter prediction by inputting each first image to be encoded into the first prediction model, so as to obtain the sixth block-level quantization parameter corresponding to each first image to be encoded.

[0278] The first quantization encoding module 1130 is configured to perform quantization encoding on each first image to be encoded based on the sixth block-level quantization parameters to obtain image encoding information corresponding to each first image to be encoded.

[0279] The first prediction model is a quantized parameter prediction model obtained by training the first training model using spatial coding performance data as the quantized prediction loss of the first training model. The spatial coding performance data is obtained by analyzing the coding performance of each first sample image, the first sample reconstructed image of each first sample image, and the first sample block-level code rate corresponding to each first sample image in the first sample image set. The first sample reconstructed image and the first sample block-level code rate are obtained by encoding and decoding each first sample image using a preset doubly oriented codec.

[0280] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0281] Figure 12 This is a block diagram of another image encoding apparatus according to an exemplary embodiment. (Refer to...) Figure 12 The device includes:

[0282] The second image acquisition module 1210 is configured to acquire at least one second image to be encoded corresponding to the first video to be encoded.

[0283] The first target association image determination module 1220 is configured to determine a first target association image corresponding to each second image to be encoded; the first target association image includes at least one of the following: at least one adjacent frame image of each second image to be encoded, a prediction frame image corresponding to at least one adjacent frame image, and a residual frame image corresponding to at least one adjacent frame image.

[0284] The third quantization parameter prediction module 1230 is configured to perform temporal block-level quantization parameter prediction by inputting each second image to be encoded and the first target associated image corresponding to each second image to be encoded into the second prediction model, thereby obtaining the seventh block-level quantization parameter corresponding to each second image to be encoded.

[0285] The second quantization encoding module 1240 is configured to perform quantization encoding on each second image to be encoded based on the seventh block-level quantization parameters to obtain the image encoding information corresponding to each second image to be encoded.

[0286] The second prediction model is a quantized parameter prediction model obtained by training the second training model using temporal coding performance data as the quantized prediction loss. The temporal coding performance data is obtained by analyzing the coding performance of each first sample image, the first sample reconstructed image of each first sample image, the first sample associated image of each first sample image, the second sample reconstructed image of the first sample associated image, the first sample block-level code rate of each first sample image, and the second sample block-level code rate of the first sample associated image. The first sample reconstructed image, the first sample block-level code rate, the second sample reconstructed image, and the second sample block-level code rate are obtained by encoding and decoding each first sample image and the first sample associated image using a preset guided codec.

[0287] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0288] Figure 13This is a block diagram of another image encoding apparatus according to an exemplary embodiment. (Refer to...) Figure 13 The device includes:

[0289] The third image acquisition module 1310 is configured to acquire at least one third image corresponding to the second video to be encoded.

[0290] The second target association image determination module 1320 is configured to determine a second target association image corresponding to each third image to be encoded; the second target association image includes at least one of the following: at least one adjacent frame image of each third image to be encoded, a prediction frame image corresponding to at least one adjacent frame image, and a residual frame image corresponding to at least one adjacent frame image.

[0291] The fourth quantization parameter prediction module 1330 is configured to perform spatiotemporal joint block-level quantization parameter prediction by inputting each third image to be encoded and the second target associated image corresponding to each third image to be encoded into the third prediction model, thereby obtaining the eighth block-level quantization parameter and the ninth block-level quantization parameter corresponding to each third image to be encoded.

[0292] The third quantization encoding module 1340 is configured to perform quantization encoding on each third image to be encoded based on the eighth block level quantization parameters or the ninth block level quantization parameters, so as to obtain the image encoding information corresponding to each third image to be encoded.

[0293] The third prediction model is a quantized parameter prediction model obtained by training the third training model using joint coding performance data as the quantized prediction loss. The joint coding performance data is obtained by analyzing the coding performance of each first sample image, the first sample reconstructed image of each first sample image, the second sample associated image corresponding to each first sample image, the third sample reconstructed image of the second sample associated image, the first sample block-level code rate corresponding to each first sample image, and the third sample block-level code rate corresponding to the second sample associated image in the first sample image set. The first sample reconstructed image, the first sample block-level code rate, the point-multiple third sample reconstructed image, and the third sample block-level code rate are obtained by encoding and decoding each first sample image and the second sample associated image using a preset guided codec.

[0294] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0295] Figure 14 This is a block diagram illustrating an electronic device for image encoding according to an exemplary embodiment. Optionally, the electronic device may be a server, and its internal structure diagram may be as follows: Figure 14 As shown, the electronic device includes a processor, memory, and a model interface connected via a system bus. The processor provides computational and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The model interface is used to communicate with external terminals via a model connection. When the computer program is executed by the processor, it implements an image encoding method.

[0296] Figure 15 This is a block diagram illustrating an electronic device for training a quantization parameter prediction model according to an exemplary embodiment. Optionally, the electronic device may be a terminal, and its internal structure diagram may be as follows: Figure 15 As shown, the electronic device includes a processor, memory, model interface, display screen, and input devices connected via a system bus. The processor provides computational and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The model interface communicates with external terminals via model connection. When the computer program is executed by the processor, it implements a method for training a quantization parameter prediction model. The display screen can be a liquid crystal display (LCD) or an e-ink display. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad mounted on the device's casing, or an external keyboard, touchpad, or mouse.

[0297] Those skilled in the art will understand that Figure 14 or Figure 15 The structure shown is merely a block diagram of a portion of the structure related to the present disclosure and does not constitute a limitation on the electronic device to which the present disclosure is applied. A specific electronic device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0298] In an exemplary embodiment, an electronic device is also provided, including: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to execute the instructions to implement a quantization parameter prediction model training method or an image encoding method as described in the embodiments of this disclosure.

[0299] In an exemplary embodiment, a computer-readable storage medium is also provided, wherein when the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the quantization parameter prediction model training method or image encoding method of the present disclosure embodiments.

[0300] In an exemplary embodiment, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute the quantization parameter prediction model training method or image encoding method of the embodiments of this disclosure.

[0301] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.

[0302] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.

[0303] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. A method for training a quantized parameter prediction model, characterized in that, include: Obtain the first sample image set; Based on the first preset training model, block-level quantization parameters are predicted for each first sample image in the first sample image set to obtain the sample block-level quantization parameters corresponding to each first sample image. Using the sample block-level quantization parameters as the quantization parameters corresponding to each first sample image, each first sample image is encoded and decoded based on a preset guide codec to obtain the first sample reconstructed image of each first sample image and the first sample block-level bitrate corresponding to each first sample image; Based on each first sample image, the first sample reconstructed image, and the first sample block-level bitrate, coding performance analysis is performed to obtain the first coding performance data corresponding to each first sample image; Using the first encoding performance data as the quantization prediction loss of the first preset training model, the first preset training model is trained to obtain a quantization parameter prediction model.

2. The method for training a quantization parameter prediction model according to claim 1, characterized in that, The sample block-level quantization parameters include a first block-level quantization parameter, which is a spatial block-level quantization parameter; the first preset training model includes a first training model; the first coding performance data includes spatial coding performance data. The quantization parameter prediction model includes a first prediction model; The step of performing block-level quantization parameter prediction on each first sample image in the first sample image set based on the first preset training model to obtain the sample block-level quantization parameter corresponding to each first sample image includes: inputting each first sample image into the first training model to perform spatial block-level quantization parameter prediction to obtain the first block-level quantization parameter corresponding to each first sample image. The step of using the sample block-level quantization parameter as the quantization parameter corresponding to each first sample image, and performing encoding and decoding processing on each first sample image based on a preset guided codec to obtain the first sample reconstructed image of each first sample image and the first sample block-level bitrate corresponding to each first sample image includes: using the first block-level quantization parameter as the quantization parameter corresponding to each first sample image, and performing encoding and decoding processing on each first sample image based on the preset guided codec to obtain the first sample reconstructed image and the first sample block-level bitrate; The step of performing coding performance analysis based on each first sample image, the first sample reconstructed image, and the first sample block-level code rate to obtain the first coding performance data corresponding to each first sample image includes: performing coding performance analysis based on each first sample image, the first sample reconstructed image, and the first sample block-level code rate to obtain the spatial coding performance data corresponding to each first sample image; The step of using the first coding performance data as the quantization prediction loss of the first preset training model to train the first preset training model and obtaining the quantization parameter prediction model includes: using the spatial coding performance data as the quantization prediction loss of the first training model to train the first training model and obtaining the first prediction model.

3. The method for training a quantization parameter prediction model according to claim 2, characterized in that, The step of inputting each first sample image into the first model to be trained to predict spatial block-level quantization parameters, and obtaining the first block-level quantization parameters corresponding to each first sample image, includes: Each first sample image and the first preset block-level quantization parameter corresponding to each first sample image are input into the first model to be trained to predict the spatial block-level quantization parameters, thereby obtaining the first block-level quantization parameters.

4. The method for training a quantization parameter prediction model according to claim 1, characterized in that, The sample block-level quantization parameters include a second block-level quantization parameter, which is a time-domain block-level quantization parameter. The first preset model to be trained includes a second model to be trained; The first sample image set includes at least one sample video frame image corresponding to a sample video; the method further includes: Determine a first sample associated image corresponding to each first sample image; the first sample associated image includes at least one of the following: at least one neighboring frame image of each first sample image, a prediction frame image corresponding to the at least one neighboring frame image, and a residual frame image corresponding to the at least one neighboring frame image; The step of predicting block-level quantization parameters for each first sample image in the first sample image set based on the first preset training model includes: Each first sample image and the first sample associated image corresponding to each first sample image are input into the second model to be trained to predict temporal block-level quantization parameters, thereby obtaining the second block-level quantization parameters corresponding to each first sample image.

5. The method for training a quantization parameter prediction model according to claim 4, characterized in that, The step of inputting each first sample image and the first sample association image corresponding to each first sample image into the second model to be trained for temporal block-level quantization parameter prediction to obtain the second block-level quantization parameter corresponding to each first sample image includes: Each first sample image, the second preset block-level quantization parameter corresponding to each first sample image, and the first sample associated image corresponding to each first sample image are input into the second model to be trained to predict the temporal block-level quantization parameter, thereby obtaining the second block-level quantization parameter.

6. The method for training a quantization parameter prediction model according to claim 4, characterized in that, The first coding performance data includes time-domain coding performance data; the quantization parameter prediction model includes a second prediction model; the method further includes: The first sample reconstructed image corresponding to each first sample image is used as the reference frame image of the first sample associated image corresponding to each first sample image. The third preset block-level quantization parameter corresponding to the first sample associated image is used as the quantization parameter corresponding to the first sample associated image. The first sample associated image is encoded and decoded based on the preset guide codec to obtain the second sample reconstructed image of the first sample associated image and the second sample block-level bitrate corresponding to the first sample associated image. The step of performing coding performance analysis based on each first sample image, the first sample reconstructed image, and the first sample block-level code rate to obtain the first coding performance data corresponding to each first sample image includes: performing coding performance analysis based on each first sample image, the first sample associated image, the first sample reconstructed image, the second sample reconstructed image, the first sample block-level code rate, and the second sample block-level code rate to obtain the temporal coding performance data; The step of using the first coding performance data as the quantization prediction loss of the first preset training model to train the first preset training model and obtaining the quantization parameter prediction model includes: using the time-domain coding performance data as the quantization prediction loss of the second training model to train the second training model and obtaining the second prediction model.

7. The method for training a quantization parameter prediction model according to claim 6, characterized in that, The step of inputting each first sample image and the first sample association image corresponding to each first sample image into the second model to be trained for temporal block-level quantization parameter prediction to obtain the second block-level quantization parameter corresponding to each first sample image includes: Each first sample image, the second preset block-level quantization parameter corresponding to each first sample image, the first sample associated image corresponding to each first sample image, and the third preset block-level quantization parameter corresponding to the first sample associated image are input into the second model to be trained to predict the block-level quantization parameters, thereby obtaining the second block-level quantization parameters and the third block-level quantization parameters corresponding to the first sample associated image. The step of using the first sample reconstructed image corresponding to each first sample image as the reference frame image of the first sample associated image corresponding to each first sample image, and using the third preset block-level quantization parameter corresponding to the first sample associated image as the quantization parameter corresponding to the first sample associated image, and performing encoding and decoding processing on the first sample associated image based on the preset guideable codec to obtain the second sample reconstructed image of the first sample associated image and the second sample block-level bitrate corresponding to the first sample associated image includes: Using the first sample reconstructed image corresponding to each first sample image as the reference frame image of the first sample associated image corresponding to each first sample image, and using the third block-level quantization parameter corresponding to the first sample associated image as the quantization parameter corresponding to the first sample associated image, the first sample associated image is encoded and decoded based on the preset guided codec to obtain the second sample reconstructed image and the second sample block-level bitrate.

8. The method for training a quantization parameter prediction model according to claim 1, characterized in that, The sample block-level quantization parameters include a fourth block-level quantization parameter and a fifth block-level quantization parameter. The fourth block-level quantization parameter is a spatial block-level quantization parameter, and the fifth block-level quantization parameter is a temporal block-level quantization parameter. The first preset model to be trained includes a third model to be trained; The first sample image set includes at least one sample video frame image corresponding to a sample video; the method further includes: Determine a second sample associated image corresponding to each first sample image; the second sample associated image includes at least one of the following: at least one neighboring frame image of each first sample image, a prediction frame image corresponding to the at least one neighboring frame image, and a residual frame image corresponding to the at least one neighboring frame image; The step of predicting block-level quantization parameters for each first sample image in the first sample image set based on the first preset training model includes: Each first sample image and the corresponding second sample associated image are input into the third model to be trained for spatiotemporal joint block-level quantization parameter prediction, to obtain the fourth block-level quantization parameter and the fifth block-level quantization parameter corresponding to each first sample image.

9. The method for training a quantization parameter prediction model according to claim 8, characterized in that, The first coding performance data includes joint coding performance data; The quantization parameter prediction model includes a third prediction model; the method further includes: The first sample reconstructed image corresponding to each first sample image is used as the reference frame image of the second sample associated image corresponding to each first sample image. The fourth preset block-level quantization parameter corresponding to the second sample associated image is used as the quantization parameter corresponding to the second sample associated image. The second sample associated image is encoded and decoded based on the preset guide codec to obtain the third sample reconstructed image of the second sample associated image and the corresponding third sample block-level bitrate of the second sample associated image. The step of performing coding performance analysis based on each first sample image, the first sample reconstructed image, and the first sample block-level code rate to obtain the first coding performance data corresponding to each first sample image includes: performing coding performance analysis based on each first sample image, the second sample associated image, the first sample reconstructed image, the third sample reconstructed image, the first sample block-level code rate, and the third sample block-level code rate to obtain the joint coding performance data; The step of using the first coding performance data as the quantization prediction loss of the first preset training model to train the first preset training model and obtaining the quantization parameter prediction model includes: using the joint coding performance data as the quantization prediction loss of the third training model to train the third training model and obtaining the third prediction model.

10. The method for training a quantization parameter prediction model according to any one of claims 1 to 9, characterized in that, The method further includes: Obtain the second sample image set and the fifth preset block-level quantization parameters corresponding to each second sample image in the second sample image set; Using the fifth preset block-level quantization parameter as the quantization parameter corresponding to each second sample image, the second preset model to be trained is used to encode and decode each second sample image to obtain the fourth sample reconstructed image of each second sample image and the fourth sample block-level bitrate corresponding to each first sample image; Based on each second sample image, the fourth sample reconstructed image, and the fourth sample block-level bitrate, coding performance analysis is performed to obtain the second coding performance data corresponding to each second sample image; Using the second encoding performance data as the quantization encoding loss of the second preset training model, the second preset training model is trained to obtain the preset guideable codec.

11. An image encoding method, characterized in that, include: Obtain at least one first image to be encoded; Each first image to be encoded is input into the first prediction model to predict spatial block-level quantization parameters, thereby obtaining the sixth block-level quantization parameters corresponding to each first image to be encoded. Based on the sixth block-level quantization parameters, each first image to be encoded is quantized and encoded to obtain image encoding information corresponding to each first image to be encoded. Wherein, the first prediction model is a quantized parameter prediction model obtained by training the first training model with spatial domain coding performance data as the quantized prediction loss of the first training model. The spatial domain coding performance data is obtained by coding performance analysis based on each first sample image in the first sample image set, the first sample reconstructed image of each first sample image, and the first sample block-level code rate corresponding to each first sample image. The first sample reconstructed image and the first sample block-level code rate are obtained by encoding and decoding each first sample image based on a preset guided codec.

12. An image encoding method, characterized in that, include: Obtain at least one second image to be encoded corresponding to the first video to be encoded; Determine the first target associated image corresponding to each second image to be encoded; The first target associated image includes at least one of the following: at least one neighboring frame image of each second image to be encoded, a prediction frame image corresponding to the at least one neighboring frame image, and a residual frame image corresponding to the at least one neighboring frame image; Each second image to be encoded and the first target associated image corresponding to each second image to be encoded are input into the second prediction model to predict the temporal block-level quantization parameters, thereby obtaining the seventh block-level quantization parameters corresponding to each second image to be encoded. Based on the seventh block-level quantization parameters, each second image to be encoded is quantized and encoded to obtain the image encoding information corresponding to each second image to be encoded. The second prediction model is a quantized parameter prediction model obtained by training the second training model using temporal coding performance data as the quantized prediction loss of the second training model. The temporal coding performance data is obtained by analyzing the coding performance of each first sample image in the first sample image set, the first sample reconstructed image of each first sample image, the first sample associated image corresponding to each first sample image, the second sample reconstructed image of the first sample associated image, the first sample block-level code rate corresponding to each first sample image, and the second sample block-level code rate corresponding to the first sample associated image. The first sample reconstructed image, the first sample block-level code rate, the second sample reconstructed image, and the second sample block-level code rate are obtained by encoding and decoding each first sample image and the first sample associated image using a preset guided codec.

13. An image encoding method, characterized in that, include: Obtain at least one third image corresponding to the second video to be encoded; Determine the second target associated image corresponding to each third image to be encoded; The second target associated image includes at least one of the following: at least one neighboring frame image of each third image to be encoded, a predicted frame image corresponding to the at least one neighboring frame image, and a residual frame image corresponding to the at least one neighboring frame image; Each third image to be encoded and the second target associated image corresponding to each third image to be encoded are input into the third prediction model to perform spatiotemporal joint block-level quantization parameter prediction, so as to obtain the eighth block-level quantization parameter and the ninth block-level quantization parameter corresponding to each third image to be encoded. Based on the eighth block-level quantization parameter or the ninth block-level quantization parameter, each third image to be encoded is quantized and encoded to obtain the image encoding information corresponding to each third image to be encoded. The third prediction model is a quantized parameter prediction model obtained by training the third training model using joint coding performance data as the quantized prediction loss. The joint coding performance data is obtained by analyzing the coding performance of each first sample image in the first sample image set, the first sample reconstructed image of each first sample image, the second sample associated image corresponding to each first sample image, the third sample reconstructed image of the second sample associated image, the first sample block-level code rate corresponding to each first sample image, and the third sample block-level code rate corresponding to the second sample associated image. The first sample reconstructed image, the first sample block-level code rate, the point-multiple third sample reconstructed image, and the third sample block-level code rate are obtained by encoding and decoding each first sample image and the second sample associated image using a preset guided codec.

14. A training device for a quantization parameter prediction model, characterized in that, include: The first sample image set acquisition module is configured to acquire the first sample image set. The first quantization parameter prediction module is configured to perform block-level quantization parameter prediction on each first sample image in the first sample image set based on a first preset training model, so as to obtain the sample block-level quantization parameter corresponding to each first sample image. The first encoding / decoding processing module is configured to perform encoding / decoding processing on each first sample image based on a preset doubly oriented codec, using the sample block-level quantization parameters as the quantization parameters corresponding to each first sample image, to obtain the first sample reconstructed image of each first sample image and the first sample block-level bitrate corresponding to each first sample image; The first coding performance analysis module is configured to perform coding performance analysis based on each first sample image, the first sample reconstructed image and the first sample block-level bitrate, and obtain the first coding performance data corresponding to each first sample image. The first model training module is configured to execute a quantization prediction loss using the first encoding performance data as the first preset model to be trained, and train the first preset model to be trained to obtain a quantization parameter prediction model.

15. An image encoding device, characterized in that, include: The first image to be encoded acquisition module is configured to acquire at least one first image to be encoded. The second quantization parameter prediction module is configured to perform spatial block-level quantization parameter prediction by inputting each first image to be encoded into the first prediction model, thereby obtaining the sixth block-level quantization parameter corresponding to each first image to be encoded. The first quantization encoding module is configured to perform quantization encoding on each first image to be encoded based on the sixth block-level quantization parameters to obtain image encoding information corresponding to each first image to be encoded. Wherein, the first prediction model is a quantized parameter prediction model obtained by training the first training model with spatial domain coding performance data as the quantized prediction loss of the first training model. The spatial domain coding performance data is obtained by coding performance analysis based on each first sample image in the first sample image set, the first sample reconstructed image of each first sample image, and the first sample block-level code rate corresponding to each first sample image. The first sample reconstructed image and the first sample block-level code rate are obtained by encoding and decoding each first sample image based on a preset guided codec.

16. An image encoding device, characterized in that, include: The second image to be encoded acquisition module is configured to acquire at least one second image to be encoded corresponding to the first video to be encoded. The first target association image determination module is configured to determine the first target association image corresponding to each second image to be encoded; The first target associated image includes at least one of the following: at least one neighboring frame image of each second image to be encoded, a prediction frame image corresponding to the at least one neighboring frame image, and a residual frame image corresponding to the at least one neighboring frame image; The third quantization parameter prediction module is configured to perform temporal block-level quantization parameter prediction by inputting each second image to be encoded and the first target associated image corresponding to each second image to be encoded into the second prediction model to obtain the seventh block-level quantization parameter corresponding to each second image to be encoded. The second quantization encoding module is configured to perform quantization encoding on each second image to be encoded based on the seventh block-level quantization parameters, so as to obtain image encoding information corresponding to each second image to be encoded. The second prediction model is a quantized parameter prediction model obtained by training the second training model using temporal coding performance data as the quantized prediction loss of the second training model. The temporal coding performance data is obtained by analyzing the coding performance of each first sample image in the first sample image set, the first sample reconstructed image of each first sample image, the first sample associated image corresponding to each first sample image, the second sample reconstructed image of the first sample associated image, the first sample block-level code rate corresponding to each first sample image, and the second sample block-level code rate corresponding to the first sample associated image. The first sample reconstructed image, the first sample block-level code rate, the second sample reconstructed image, and the second sample block-level code rate are obtained by encoding and decoding each first sample image and the first sample associated image using a preset guided codec.

17. An image encoding device, characterized in that, include: The third image to be encoded acquisition module is configured to acquire at least one third image to be encoded corresponding to the second video to be encoded. The second target association image determination module is configured to determine the second target association image corresponding to each third image to be encoded. The second target associated image includes at least one of the following: at least one neighboring frame image of each third image to be encoded, a predicted frame image corresponding to the at least one neighboring frame image, and a residual frame image corresponding to the at least one neighboring frame image; The fourth quantization parameter prediction module is configured to perform spatiotemporal joint block-level quantization parameter prediction by inputting each third image to be encoded and the second target associated image corresponding to each third image to be encoded into the third prediction model, thereby obtaining the eighth block-level quantization parameter and the ninth block-level quantization parameter corresponding to each third image to be encoded. The third quantization encoding module is configured to perform quantization encoding on each third image to be encoded based on the eighth block-level quantization parameter or the ninth block-level quantization parameter, so as to obtain the image encoding information corresponding to each third image to be encoded. The third prediction model is a quantized parameter prediction model obtained by training the third training model using joint coding performance data as the quantized prediction loss. The joint coding performance data is obtained by analyzing the coding performance of each first sample image in the first sample image set, the first sample reconstructed image of each first sample image, the second sample associated image corresponding to each first sample image, the third sample reconstructed image of the second sample associated image, the first sample block-level code rate corresponding to each first sample image, and the third sample block-level code rate corresponding to the second sample associated image. The first sample reconstructed image, the first sample block-level code rate, the point-multiple third sample reconstructed image, and the third sample block-level code rate are obtained by encoding and decoding each first sample image and the second sample associated image using a preset guided codec.

18. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the instructions to implement the quantization parameter prediction model training method as described in any one of claims 1 to 10 or the image encoding method as described in any one of claims 11 to 13.

19. A computer-readable storage medium, characterized in that, When the instructions in the storage medium are executed by the processor of the electronic device, the electronic device is able to perform the quantization parameter prediction model training method as described in any one of claims 1 to 10 or the image encoding method as described in any one of claims 11 to 13.

20. A computer program product, characterized in that, When it is run on a computer, it causes the computer to perform the quantization parameter prediction model training method as described in any one of claims 1 to 10 or the image encoding method as described in any one of claims 11 to 13.

Citation Information

Patent Citations

  • Video encoding method and device, electronic equipment and storage medium

    CN112383777A

  • Image enhancement model training method, image enhancement method and related device

    CN112419219A