Image decoding method and related products
By employing joint training and differential parameter tuning, the problem of decoder incompatibility with bitstream files during the optimization and iteration process of traditional codecs is solved. This enables multiple decoding models to decode bitstream files generated by the same encoding model, thereby improving the optimization and iteration effect of the codec models.
Patent Information
- Application Number
- CN202411078791.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-07
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2044-08-07
AI Technical Summary
In traditional codec optimization iterations, multiple different decoders struggle to decode bitstream files encoded by the same encoder, leading to incompatibility of bitstream files during the optimization iteration process and affecting the optimization iteration effect of the codec.
By jointly training the first encoding model and the first decoding model, and adjusting the parameters based on the differences between the reconstructed image and the training image, an encoding model and a decoding model that meet the first objective are obtained. With the first encoding model unchanged, the second decoding model is trained to meet the second objective, thereby realizing the optimization and iteration of the encoding and decoding models and ensuring that multiple decoding models can decode the bitstream file generated by the same encoding model.
This enables multiple different decoding models to decode the same bitstream file generated by the same encoding model during the optimization and iteration process of the encoding and decoding model, thereby improving the compatibility of the encoding and decoding model and the optimization and iteration effect.
Smart Images

Figure CN119011861B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image encoding and decoding, and in particular to an image decoding method and related products. Background Technology
[0002] With the widespread application of image encoding and decoding technologies in various information systems, deep learning technology has also made significant progress in this field. As business scenarios continue to evolve, the demands for image encoding and decoding capabilities are also constantly changing. Utilizing deep learning technology to train encoders and decoders, achieving codec optimization and iteration, and expanding image encoding and decoding capabilities has become a technological trend. However, traditional codec optimization techniques suffer from the limitation that multiple different decoders struggle to decode bitstream files encoded by the same encoder. In other words, during the optimization and iteration of encoding and decoding models, multiple different decoding models find it difficult to decode bitstream files generated by the same encoding model. Summary of the Invention
[0003] Therefore, it is necessary to address the aforementioned technical problems by providing an image decoding method and related products that enable multiple different decoding models to decode the bitstream file generated by the same encoding model during the optimization and iteration process of the encoding and decoding model. These related products include video frame display methods, decoding model training methods, image decoding devices, video frame display devices, decoding model training devices, computer equipment, computer-readable storage media, and computer program products.
[0004] Firstly, this application provides an image decoding method, the method comprising:
[0005] A target bitstream file is obtained, which is obtained by encoding the image to be compressed using a first encoding model. The first encoding model and the first decoding model are trained in the following way: the first bitstream file output by the basic encoding model after encoding the first training image is used as the input of the first basic decoding model; the first basic decoding model is used to decode the first bitstream file to obtain the first reconstructed image; based on the difference between the first reconstructed image and the first training image, the basic encoding model and the first basic decoding model are tuned to obtain the first encoding model and the first decoding model that meets the first target.
[0006] The target bitstream file is decoded using a second decoding model to obtain a target image. The second decoding model is trained in the following way: the second bitstream file output by the first encoding model after encoding the second training image is used as the input of the second basic decoding model; the second basic decoding model is used to decode the second bitstream file to obtain a second reconstructed image; based on the difference between the second reconstructed image and the second training image, the parameters of the second basic decoding model are tuned to obtain a second decoding model that meets the second target.
[0007] The first objective of training the first basic decoding model is different from the second objective of training the second basic decoding model.
[0008] The image decoding method provided in the first aspect involves jointly training a first encoding model and a first decoding model. Specifically, the first bitstream file output by the basic encoding model after encoding the first training image is used as the input to the first basic decoding model. The first basic decoding model decodes the first bitstream file to obtain a first reconstructed image. Based on the difference between the first reconstructed image and the first training image, the basic encoding model and the first basic decoding model are tuned to obtain the first encoding model and a first decoding model that meets the first objective. The second decoding model is trained based on the first encoding model while keeping the first encoding model unchanged. Specifically, the second bitstream file output by the first encoding model after encoding the second training image is used as the input to the second basic decoding model. The second basic decoding model decodes the second bitstream file to obtain a second reconstructed image. Based on the difference between the second reconstructed image and the second training image, the second basic decoding model is tuned to obtain a second decoding model that meets the second objective.
[0009] Since the first objective of training the first basic decoding model differs from the second objective of training the second basic decoding model—that is, the first objective met by the first decoding model differs from the second objective met by the second decoding model—it indicates that the second decoding model and the first decoding model have achieved iterative optimization of the decoding model. This can also be understood as the first encoding model and the first decoding model, together with the first encoding model and the second decoding model, achieving iterative optimization of the encoding and decoding model. For example, when jointly training the first decoding model and the first encoding model, the first objective of the first decoding model is to decode the bitstream file output by the first encoding model after encoding the initial image into a decoded image similar to the initial image; when training the second decoding model based on the first encoding model, the second objective of the second decoding model is to decode the bitstream file output by the first encoding model after encoding the initial image into a decoded image similar to the initial image with the desired sharpness.
[0010] In addition, since the first decoding model is jointly trained with the first encoding model, it can decode the bitstream file output by the first encoding model during model inference. The second decoding model is trained by using the second bitstream file output by the first encoding model after encoding the second training image as input to the second basic decoding model. Therefore, the second decoding model can also decode the bitstream file output by the first encoding model during model inference.
[0011] The first encoding model, first decoding model, and second decoding model obtained through the above training process not only embody the optimization and iteration of the encoding and decoding models, but also demonstrate that multiple different decoding models can decode the bitstream file generated by the same encoding model during the optimization and iteration process. Based on this, by obtaining the target bitstream file obtained by encoding the image to be compressed using the first encoding model, and then decoding the target bitstream file using the second decoding model to obtain the target image, it is achieved that multiple different decoding models can decode the bitstream file generated by the same encoding model during the optimization and iteration process of the encoding and decoding models.
[0012] In one embodiment, the second target of training the second basic decoding model matches the display effect of the decoded image output by the second decoding model; before decoding the target bitstream file using the second decoding model to obtain the target image, the method further includes:
[0013] Obtain the display requirements for the display effect of the target image, wherein the target image refers to the image obtained by encoding the target bitstream file;
[0014] If the display effect of the decoded image output by the second decoding model meets the display requirements, the second decoding model is determined to be the model for decoding the target bitstream file.
[0015] In one embodiment, the second decoding model that meets the second objective is obtained by tuning the parameters of the second basic decoding model based on a loss, wherein the loss is calculated using a target loss function based on the difference between the second reconstructed image and the second training image.
[0016] The target training set, which includes the second training image, and the target loss function, are matched with the second target.
[0017] In one embodiment,
[0018] When the second objective is that the first target attribute of the decoded image generated by the second decoding model meets the image expectation requirement, and the first target attribute refers to any image attribute that affects the visual effect of the image, the objective loss function includes information to judge the difference between the first target attribute of the second reconstructed image generated by the second basic decoding model and the image expectation requirement.
[0019] In one embodiment,
[0020] When the second target is the region where the target object in the decoded image generated by the second decoding model meets the local expectation requirement, and the second target attribute refers to any image attribute that affects the visual effect of the image, the target loss function includes information on detecting the target object in the second reconstructed image generated by the second basic decoding model, and information on judging the difference between the second target attribute of the region where the target object in the second reconstructed image is located and the local expectation requirement. The number of second training images with target content in the target training set is greater than or equal to a first value, and having target content means having a preset number or more of the target objects.
[0021] In one embodiment,
[0022] When the second objective is that the third objective attribute of the initial image is lower than the preset requirement, the second decoding model decodes the bitstream file corresponding to the initial image, and the similarity between the decoded image and the initial image reaches the expected similarity value. The third objective attribute refers to any image attribute that affects the visual effect of the image. In the case of the third objective attribute being lower than the preset requirement, the number of second training images in the objective training set is greater than or equal to the second value.
[0023] Secondly, this application also provides a video frame display method, the method comprising:
[0024] A target bitstream file is obtained after video frames are encoded. The target bitstream file is obtained by encoding the video frames using a first encoding model. The first encoding model and the first decoding model are trained in the following way: the first bitstream file output by the basic encoding model after encoding the first training image is used as the input of the first basic decoding model; the first basic decoding model is used to decode the first bitstream file to obtain a first reconstructed image; based on the difference between the first reconstructed image and the first training image, the basic encoding model and the first basic decoding model are tuned to obtain the first encoding model and the first decoding model that meets the first target.
[0025] The target bitstream file is decoded using a second decoding model to obtain a target image. The second decoding model is trained as follows: the second bitstream file output by the first encoding model after encoding the second training image is used as the input of the second basic decoding model; the second basic decoding model is used to decode the second bitstream file to obtain a second reconstructed image; based on the difference between the second reconstructed image and the second training image, the parameters of the second basic decoding model are tuned to obtain a second decoding model that meets the second objective; wherein, the first objective for training the first basic decoding model is different from the second objective for training the second basic decoding model.
[0026] Display the target image.
[0027] Thirdly, this application also provides a decoding model training method, the method comprising:
[0028] Obtain the target training set;
[0029] The first encoding model is used to encode the second training image in the target training set to obtain a second bitstream file. The first encoding model and the first decoding model are trained in the following way: the first bitstream file output by the basic encoding model after encoding the first training image is used as the input of the first basic decoding model; the first basic decoding model is used to decode the first bitstream file to obtain a first reconstructed image; based on the difference between the first reconstructed image and the first training image, the basic encoding model and the first basic decoding model are tuned to obtain the first encoding model and the first decoding model that conforms to the first target.
[0030] The second bitstream file is input into the second basic decoding model to obtain the second reconstructed image;
[0031] Based on the difference between the second reconstructed image and the second training image, the parameters of the second basic decoding model are tuned to obtain a second decoding model that meets the second objective.
[0032] The first objective of training the first basic decoding model is different from the second objective of training the second basic decoding model.
[0033] In one embodiment, when the second objective is that the first target attribute of the decoded image generated by the second decoding model meets the image expectation requirement, where the first target attribute refers to any image attribute that affects the visual effect of the image, the step of tuning the parameters of the second basic decoding model based on the difference between the second reconstructed image and the second training image to obtain a second decoding model that meets the second objective includes:
[0034] The loss is calculated based on the difference between the second reconstructed image and the second training image, and the difference between the first target attribute of the second reconstructed image and the image expectation requirement.
[0035] Based on the loss, the parameters of the second basic decoding model are tuned to obtain the second decoding model that meets the second objective.
[0036] In one embodiment, when the second target is the region where the target object in the decoded image generated by the second decoding model meets the local expectation requirement, and the second target attribute refers to any image attribute that affects the visual effect of the image, the step of tuning the second basic decoding model based on the difference between the second reconstructed image and the second training image to obtain a second decoding model that meets the second target includes:
[0037] Based on the difference between the second reconstructed image and the second training image, and the difference between the second target attribute of the region where the target object is located in the second reconstructed image and the local expectation requirement, the loss is calculated;
[0038] Based on the loss, the parameters of the second basic decoding model are tuned to obtain the second decoding model that meets the second objective.
[0039] Fourthly, this application also provides an image decoding apparatus, the apparatus comprising:
[0040] The bitstream acquisition module is used to acquire a target bitstream file, which is obtained by encoding the image to be compressed using a first encoding model. The first encoding model and the first decoding model are trained in the following way: the first bitstream file output by the basic encoding model after encoding the first training image is used as the input of the first basic decoding model; the first basic decoding model is used to decode the first bitstream file to obtain a first reconstructed image; based on the difference between the first reconstructed image and the first training image, the basic encoding model and the first basic decoding model are tuned to obtain the first encoding model and the first decoding model that meets the first target.
[0041] The bitstream decoding module is used to decode the target bitstream file using a second decoding model to obtain a target image. The second decoding model is trained in the following way: the second bitstream file output by the first encoding model after encoding the second training image is used as the input of the second basic decoding model; the second basic decoding model is used to decode the second bitstream file to obtain a second reconstructed image; based on the difference between the second reconstructed image and the second training image, the parameters of the second basic decoding model are tuned to obtain a second decoding model that meets the second target.
[0042] The first objective of training the first basic decoding model is different from the second objective of training the second basic decoding model.
[0043] Fifthly, this application also provides a video frame display device, the device comprising:
[0044] The target bitstream acquisition module is used to acquire the target bitstream file formed after the video frames are encoded. The target bitstream file is obtained by encoding the video frames using a first encoding model. The first encoding model and the first decoding model are trained in the following way: the first bitstream file output by the basic encoding model after encoding the first training image is used as the input of the first basic decoding model; the first basic decoding model is used to decode the first bitstream file to obtain the first reconstructed image; based on the difference between the first reconstructed image and the first training image, the basic encoding model and the first basic decoding model are tuned to obtain the first encoding model and the first decoding model that meets the first target.
[0045] The target image generation module is used to decode the target bitstream file using a second decoding model to obtain a target image. The second decoding model is trained in the following way: the second bitstream file output by the first encoding model after encoding the second training image is used as the input of the second basic decoding model; the second basic decoding model is used to decode the second bitstream file to obtain a second reconstructed image; based on the difference between the second reconstructed image and the second training image, the parameters of the second basic decoding model are tuned to obtain a second decoding model that conforms to the second objective; wherein, the first objective for training the first basic decoding model is different from the second objective for training the second basic decoding model.
[0046] The target image display module is used to display the target image.
[0047] Sixthly, this application also provides a decoding model training apparatus, the apparatus comprising:
[0048] The training set acquisition module is used to acquire the target training set;
[0049] A bitstream generation module is used to encode a second training image in the target training set using a first encoding model to obtain a second bitstream file. The first encoding model and the first decoding model are trained in the following way: the first bitstream file output by the basic encoding model after encoding the first training image is used as the input of the first basic decoding model; the first basic decoding model is used to decode the first bitstream file to obtain a first reconstructed image; based on the difference between the first reconstructed image and the first training image, the basic encoding model and the first basic decoding model are tuned to obtain the first encoding model and the first decoding model that conforms to the first target.
[0050] The image reconstruction module is used to input the second bitstream file into the second basic decoding model to obtain the second reconstructed image;
[0051] The model generation module is used to tune the parameters of the second basic decoding model based on the difference between the second reconstructed image and the second training image to obtain a second decoding model that meets the second objective.
[0052] The first objective of training the first basic decoding model is different from the second objective of training the second basic decoding model.
[0053] In a seventh aspect, this application also provides a computer device, including: a memory and a processor, wherein the memory stores program instructions; when the program instructions are executed by the processor, the processor performs the method as shown in the first aspect or any embodiment of the first aspect, or performs the method as shown in the second aspect or any embodiment of the second aspect, or performs the method as shown in the third aspect or any embodiment of the third aspect.
[0054] Eighthly, this application also provides a computer-readable storage medium storing a computer program; when the computer program is run on one or more processors, it performs the method as shown in the first aspect or any embodiment of the first aspect, or performs the method as shown in the second aspect or any embodiment of the second aspect, or performs the method as shown in the third aspect or any embodiment of the third aspect.
[0055] Ninthly, this application also provides a computer program product comprising a computer program or instructions; wherein, when the computer program or instructions are executed on a computer, the computer performs the method as shown in the first aspect or any embodiment of the first aspect, or performs the method as shown in the second aspect or any embodiment of the second aspect, or performs the method as shown in the third aspect or any embodiment of the third aspect.
[0056] It is understood that the video frame display method provided in the second aspect, the decoding model training method provided in the third aspect, the image decoding apparatus provided in the fourth aspect, the video frame display apparatus provided in the fifth aspect, the decoding model training apparatus provided in the sixth aspect, the computer device provided in the seventh aspect, the computer-readable storage medium provided in the eighth aspect, and the computer program product provided in the ninth aspect can be used to execute the image decoding method shown in the first aspect or any embodiment of the first aspect of this application, or related to the image decoding method shown in the first aspect or any embodiment of the first aspect of this application. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects in the image decoding method, and will not be repeated here. Attached Figure Description
[0057] To clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the embodiments will be described below.
[0058] Figure 1 One of the schematic flowcharts of the image decoding method provided in the embodiments of this application;
[0059] Figure 2 A schematic diagram of the model structure of the basic encoding model and the first basic decoding model during joint training provided in the embodiments of this application;
[0060] Figure 3 A second schematic flowchart illustrating the image decoding method provided in this application embodiment;
[0061] Figure 4 A flowchart illustrating the video frame display method provided in an embodiment of this application;
[0062] Figure 5 A flowchart illustrating the decoding model training method provided in this application embodiment;
[0063] Figure 6 One of the flowcharts provided in this application illustrates how the second basic decoding model is tuned based on the difference between the second reconstructed image and the second training image to obtain a second decoding model that meets the second objective.
[0064] Figure 7The second flowchart illustrates how the second basic decoding model is tuned based on the difference between the second reconstructed image and the second training image to obtain a second decoding model that meets the second objective, as provided in the embodiments of this application.
[0065] Figure 8 This is a schematic diagram of the structure of the image decoding device provided in the embodiments of this application;
[0066] Figure 9 This is a schematic diagram of the structure of the video frame display device provided in the embodiments of this application;
[0067] Figure 10 This is a schematic diagram of the structure of the decoding model training device provided in the embodiments of this application;
[0068] Figure 11 A schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0069] To facilitate understanding of the embodiments of this application, a more comprehensive description of the embodiments of this application will be provided below with reference to the accompanying drawings. The drawings illustrate preferred embodiments of the embodiments of this application. However, the embodiments of this application can be implemented in many different forms and are not limited to the embodiments described herein. Rather, these embodiments are provided to make the disclosure of the embodiments of this application more thorough and complete.
[0070] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which embodiments of this application belong. The terminology used herein in the description of embodiments of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the embodiments of this application.
[0071] The terms "first," "second," etc., used in this document are used to distinguish different objects, not to describe a specific order. For the purpose of distinguishing different objects, a "first decoding model" can also be described as a "second decoding model," and vice versa. When used, the singular expressions "a," "an," "the," "the," "this," and "this" are intended to include the plural expressions as well, unless the context explicitly indicates otherwise. It should also be understood that the term "including / comprising" specifies the presence of the stated feature, whole, step, operation, part, or fusion thereof, but does not preclude the possibility of the presence or addition of one or more other features, wholes, steps, operations, parts, or fusions thereof.
[0072] Traditional iterative codec optimization techniques suffer from the drawback of having multiple different decoders struggling to decode a bitstream file encoded by the same encoder. Specifically, consider encoders E_a and D_a, and encoders E_b and D_b. In traditional techniques, the encoders and decoders are trained synchronously and jointly. That is, during the training phase, E_a and D_a are trained end-to-end, and E_b and D_b are trained end-to-end. The loss function used during training is typically the mean squared error function. However, during the inference phase, inference can be performed separately for E_a and D_a, and separately for E_b and D_b. E_a takes the initial image as input and outputs the encoded bitstream file S_a. Bitstream file S_a is used as input to D_a, and D_a outputs the decoded image. Similarly, E_b takes the initial image as input and outputs the encoded bitstream file S_b. Bitstream file S_b is used as input to D_a, and D_a outputs the decoded image. In this model, the bitstream file S_a encoded by E_a can be decoded by the decoder D_a and reconstructed into a decoded image, and the bitstream file S_b encoded by E_b can be decoded by the decoder D_b and reconstructed into a decoded image. However, D_b cannot decode the bitstream file S_a encoded by E_a, and similarly, D_a cannot decode the bitstream file S_b encoded by E_b. The codec needs continuous optimization and iteration. With each iteration, a problem arises where a second-version decoder struggles to decode bitstream files encoded by a first-version encoder. In other words, only a second-version decoder can decode bitstream files encoded by a second-version encoder, and only a first-version decoder can decode bitstream files encoded by a first-version encoder. Multiple different decoders struggle to decode bitstream files encoded by the same encoder. If we consider the codec as an encoding-decoding model, the encoder as an encoding model, and the decoder as a decoding model, the aforementioned technical problem can be understood as follows: during the optimization and iteration of the encoding-decoding model, bitstream files from different encoding-decoding models are incompatible, and multiple different decoding models struggle to decode bitstream files encoded by the same encoding model.
[0073] To meet the evolving demands for image encoding and decoding capabilities arising from changing business scenarios, codecs require continuous optimization and iteration. However, the aforementioned shortcomings in traditional codec optimization techniques not only diminish the significance of these iterations but also risk rendering existing bitstream files undecoded. Therefore, this application provides an image decoding method that ensures compatibility between bitstream files from different codec models during the codec optimization process—that is, during the optimization of the codec model. Different decoding models can decode bitstream files encoded by the same encoding model, and a second version of the decoding model can also decode bitstream files encoded by the first version of the encoding model.
[0074] This application provides an image decoding method. For example... Figure 1 As shown, the image decoding method includes the following steps S101 to S102.
[0075] S101, Obtain the target bitstream file. The target bitstream file is obtained by encoding the image to be compressed using the first encoding model. The first encoding model and the first decoding model are trained in the following way: the first bitstream file output by the basic encoding model after encoding the first training image is used as the input of the first basic decoding model; the first basic decoding model is used to decode the first bitstream file to obtain the first reconstructed image; based on the difference between the first reconstructed image and the first training image, the basic encoding model and the first basic decoding model are tuned to obtain the first encoding model and the first decoding model that meets the first target.
[0076] The first encoding model and the first decoding model are jointly trained as a whole. Based on the differences between the first reconstructed image and the first training image, the parameters of the basic encoding model and the first basic decoding model are tuned as a whole; that is, the parameters of the basic encoding model and the parameters of the first basic decoding model influence each other. The first objective refers to the decoding model training objective that the first decoding model conforms to, which is also the objective of training the first basic decoding model. It can also be understood as the encoding and decoding model training objective that the first encoding model and the first decoding model conform to.
[0077] Optionally, the base encoding model used to train the first encoding model can be Normalizer-Free ResNets (NFNet), Visiontransformer (ViT), Convolutional Neural Networks (CNN), etc. Similarly, the first base decoding model used to train the first decoding model can also be a Convolutional Neural Network (CNN), etc. Taking a scenario where both the base encoding model and the first base decoding model are CNNs, the model structures of the base encoding model and the first base decoding model during joint training are as follows: Figure 2 As shown. Figure 2 The meanings of the English symbols are as follows: Conv represents a convolutional layer; TConv represents a transposed convolutional layer; k5s2 indicates that the convolutional layer has a 5x5 kernel and a stride of 2; k3s1 indicates that the convolutional layer has a 3x3 kernel and a stride of 1; k3s2 indicates that the convolutional layer has a 3x3 kernel and a stride of 2; k4s2 indicates that the convolutional layer has a 4x4 kernel and a stride of 2; k4s1 indicates that the convolutional layer has a 4x4 kernel and a stride of 1; k4s4 indicates that the convolutional layer has a 4x4 kernel and a stride of 4; ReLU6 and ReLU represent activation functions; x represents the first training image; y represents the first feature of the image; z represents the second feature of the image. This represents the first reconstructed image; and These represent the data obtained after decoding the bitstream data; Q represents quantization; AE represents the arithmetic encoder; AD represents the arithmetic decoder; μ and σ both represent the parameters of the arithmetic encoder and the arithmetic decoder.
[0078] like Figure 2 As shown, the basic encoding model first comprises a series of convolutional layers and activation functions, designed to extract the first feature y of the first training image from the input first training image x. Specifically, the basic encoding model can sequentially employ four convolutional layers (Conv) and three activation functions (ReLU6), as shown below. Figure 2The four convolutional layers (Conv) are arranged alternately at the intervals shown. The kernels of all four Conv layers can be 5*5, and the stride is 2, to achieve a 16-fold downsampling of the first training image x. The ReLU6 activation function effectively improves the performance of the deep learning model, enabling the basic encoding model to better capture and transmit the linear relationships between features. Through the design of the above-mentioned Conv layers and ReLU6 activation function, rich image information can be captured while reducing the image size, and the first feature y of the image can be extracted. Since the bitstream data compressed based on the first feature y of the image may be too large, the basic encoding model also includes another series of convolutional layers and activation functions, aiming to extract the second feature z of the image from the first feature y corresponding to the input first training image x, and then compress the bitstream data based on the second feature z to obtain a smaller bitstream data. Specifically, the basic encoding model can sequentially use three Conv layers and three ReLU activation functions. First, a convolutional layer with a 3*3 kernel and a stride of 1 is used, followed by two convolutional layers with a 5*5 kernel and a stride of 2. The three Conv layers and three ReLU activation functions are as follows: Figure 2 The intervals shown are alternated. Through the design of the Conv convolutional layers and the ReLU activation function, it is possible to increase the depth and width of the network while further downsampling to reduce image size and capture the linear relationship between features, thereby improving the richness and diversity of the acquired second feature z. Then, the second feature z of the image is quantized by Q and input to the arithmetic encoder AE to obtain the first part of the bitstream data, such as 100111. Understandably, compressing the second feature z obtained by further downsampling results in a smaller first part of the bitstream data. Furthermore, the first feature y of the image is quantized by Q and input to the arithmetic encoder AE to obtain the second part of the bitstream data, such as 100111. Accordingly, the first part of the bitstream data 100111 and the second part of the bitstream data 100111 are concatenated, with the first part of the bitstream data first and the second part of the bitstream data second, to obtain the first bitstream file 100111100111 obtained by the basic encoding model encoding the first training image x.
[0079] It is important to note that, such as Figure 2 As shown, in the process of inputting the first feature y of the image into the arithmetic encoder AE after quantization Q to obtain the second part of the bitstream data, entropy coding is required, that is, parameters μ and σ need to be obtained to compress the first feature y more efficiently, so that the second part of the bitstream data output by the arithmetic encoder AE is shorter and smaller. At this time, the first part of the bitstream data is first decoded by the arithmetic decoder AD to obtain the data. Then through, such as Figure 2 The diagram shows three transposed convolutional layers (TConv) and two ReLU activation functions arranged in alternating intervals, which process the data... The data is converted into parameters μ and σ. First, it passes through a convolutional layer with a 3x3 kernel and a stride of 2, followed by two convolutional layers with 4x4 kernels and strides of 2 and 1 respectively. Then, the parameters μ and σ are input into the basic coding model, enabling the arithmetic encoder AE to output a smaller second part of the bitstream data using these parameters.
[0080] like Figure 2 As shown, the first basic decoding model contains a series of transposed convolutional layers and activation functions, aiming to reconstruct the first reconstructed image based on the second part of the input bitstream data. Specifically, the second part of the input bitstream data is first decoded using an arithmetic decoder (AD), and the output data is then processed. Correspondingly, the arithmetic decoder (AD) also needs to obtain parameters μ and σ during the decoding process of the second part of the bitstream data. Based on this, the first basic decoding model obtains the parameters μ and σ obtained in the aforementioned manner, enabling the arithmetic decoder (AD) to decode the second part of the bitstream data using these parameters. Furthermore, the first basic decoding model can sequentially use three transposed convolutional layers (TConv) and two activation functions (ReLU) to decode the data. Converted into the first reconstructed image This can be achieved by first using two transposed convolutional layers with 4x4 kernels and a stride of 2, followed by a single convolutional layer with 4x4 kernels and a stride of 4, to achieve a 16x upsampling. The three transposed convolutional layers (TConv) and two ReLU activation functions are as follows: Figure 2 The intervals shown are alternating. Understandably, to achieve a 16x upsampling corresponding to a 16x downsampling, the first basic decoding model could also contain four transposed convolutional layers, each with a stride of 2. Figure 2 The model structure shown, which includes three transposed convolutional layers (TConv), can accelerate the decoding and reconstruction speed. Through the design of the transposed convolutional layers (TConv) and the ReLU activation function, a first reconstructed image that is similar to the first training image x and of high quality can be reconstructed.
[0081] go through Figure 2 The model structures of the base encoding model and the first base decoding model shown during joint training are capable of encoding and decoding the first training image to obtain the first reconstructed image. It should be understood that... Figure 2The model structure shown is merely an optional example. Provided that the first training image can be encoded and decoded to obtain the first reconstructed image, this embodiment does not specifically limit the model structure of the basic encoding model and the first basic decoding model. Furthermore, based on the differences between the first reconstructed image and the first training image, the basic encoding model and the first basic decoding model can be tuned to obtain the first encoding model and the first decoding model that meets the first objective. Specifically, the tuning can involve adjusting the parameters of each convolutional layer and each loss function in the basic encoding model, and adjusting the parameters of each transposed convolutional layer and each loss function in the first basic decoding model.
[0082] S102, the target bitstream file is decoded using the second decoding model to obtain the target image. The second decoding model is trained as follows: the second bitstream file output by the first encoding model after encoding the second training image is used as the input of the second basic decoding model; the second basic decoding model is used to decode the second bitstream file to obtain the second reconstructed image; based on the difference between the second reconstructed image and the second training image, the parameters of the second basic decoding model are tuned to obtain a second decoding model that meets the second objective; wherein, the first objective of training the first basic decoding model is different from the second objective of training the second basic decoding model.
[0083] The second decoding model is trained based on the first encoding model while keeping the first encoding model unchanged. Based on the differences between the second reconstructed image and the second training image, the parameters of the second basic decoding model are tuned. At this point, the first encoding model is fixed, and adjusting the parameters of the second basic decoding model has no effect on the parameters of the first encoding model. The second objective refers to the decoding model training objective that the second decoding model conforms to, which is also the objective of training the second basic decoding model. It can also be understood as the encoding-decoding model training objective that the first and second encoding models conform to. Optionally, the second basic decoding model can also be a CNN.
[0084] Since the first objective of training the first basic decoding model differs from the second objective of training the second basic decoding model—that is, the first objective met by the first decoding model differs from the second objective met by the second decoding model—it indicates that the second decoding model and the first decoding model have achieved iterative optimization of the decoding model. This can also be understood as the first encoding model and the first decoding model, together with the first encoding model and the second decoding model, achieving iterative optimization of the encoding and decoding model. For example, when jointly training the first decoding model and the first encoding model, the first objective of the first decoding model is to decode the bitstream file output by the first encoding model after encoding the initial image into a decoded image similar to the initial image; when training the second decoding model based on the first encoding model, the second objective of the second decoding model is to decode the bitstream file output by the first encoding model after encoding the initial image into a decoded image similar to the initial image with the desired sharpness.
[0085] In addition, since the first decoding model is jointly trained with the first encoding model, it can decode the bitstream file output by the first encoding model during model inference. The second decoding model is trained by using the second bitstream file output by the first encoding model after encoding the second training image as input to the second basic decoding model. Therefore, the second decoding model can also decode the bitstream file output by the first encoding model during model inference.
[0086] Taking the first encoding model as E_a and the first decoding model as D_a as an example, E_a and D_a are obtained by jointly training the basic encoding model and the first basic decoding model. D_a meets the first goal of decoding the bitstream file output by E_a after encoding the initial image into a decoded image similar to the initial image.
[0087] Optionally, while keeping the first encoding model E_a unchanged, the encoding and decoding models are optimized iteratively, that is, the decoding model is optimized iteratively. Based on different decoding model training objectives, D_a1, D_a2, D_a3, etc., are trained respectively. The second decoding model is any one of D_a1, D_a2, D_a3, etc., and is trained from the second basic decoding model. Taking D_a1 as an example, D_a1 meets the second objective of decoding the bitstream file output by E_a after encoding the initial image into a decoded image that is similar to the initial image but of higher quality. Higher quality includes clearer details, richer details, and higher pixel count. During the optimization and iteration of the encoding and decoding models, both D_a1 and D_a can decode the bitstream file encoded by E_a. The first basic decoding model and the second basic decoding model can be the same or different; the second basic decoding model when the second decoding model is D_a1 can be the same or different from the second basic decoding model when the second decoding model is D_a2.
[0088] Optionally, while keeping the first encoding model E_a unchanged, the encoding and decoding models are optimized iteratively, that is, the decoding model is optimized iteratively, and D_aN is obtained by training based on the decoding model training objective. Further, while keeping the first encoding model E_a unchanged, the encoding and decoding models are optimized iteratively again, that is, the decoding model is optimized iteratively again. Based on different decoding model training objectives, D_aN1, D_aN2, D_aN3, etc., are obtained by training D_aN. The second decoding model is any one of D_a1, D_a2, D_a3, etc., and is obtained by training the second basic decoding model. Taking D_aN1 as an example, D_aN1 meets the second objective of decoding the bitstream file output after encoding the initial image with E_a into a decoded image that is similar to the initial image but of higher quality. Higher quality includes clearer details, richer details, higher pixel count, etc. It should be understood that in this optional example, the process of training D_aN based on the first decoding model D_a through optimization iteration, and then training the second decoding model through further optimization iteration, can still be regarded as an optimization iteration of the decoding model between the second and first decoding models. Similarly, the optimization iteration of the encoding and decoding models between the first and second decoding models is achieved between the first and second encoding models. During the optimization iteration of the encoding and decoding models, both D_aN1 and D_a can decode the bitstream file encoded by E_a. Unlike the aforementioned optional example where the second decoding model is obtained through only one optimization iteration of the decoding model, in this optional example, the second basic decoding model trained to obtain the second decoding model is D_aN. D_aN is obtained through optimization iteration relative to the first decoding model D_a; therefore, the second basic decoding model D_aN is different from the first basic decoding model trained to obtain the first decoding model D_a. Furthermore, the second basic decoding model when the second decoding model is D_aN1 is the same as the second basic decoding model when the second decoding model is D_aN2, both being D_aN.
[0089] The first encoding model, first decoding model, and second decoding model obtained through the above training process not only embody the optimization and iteration of the encoding and decoding model, but also demonstrate that multiple different decoding models can decode the bitstream file generated by the same encoding model during the optimization and iteration process. Based on this, the image decoding method in this embodiment obtains the target bitstream file obtained by encoding the image to be compressed using the first encoding model, and decodes the target bitstream file using the second decoding model to obtain the target image. This achieves that during the optimization and iteration process of the encoding and decoding model, multiple different decoding models can decode the bitstream file generated by the same encoding model, and the second version of the second decoding model can decode the target bitstream file generated by the first version of the first encoding model. This image decoding method is beneficial to promoting the development of deep learning technology in the field of image encoding and decoding, and also helps improve the efficiency of implementing deep learning-based image encoding and decoding technology in practical business applications.
[0090] In one embodiment, the second objective of training the second basic decoding model is to match the display effect of the decoded image output by the second decoding model. For example... Figure 3 As shown, the image decoding method further includes the following steps S301 to S304. Steps S301 and S304 correspond one-to-one with steps S101 and S102 in the aforementioned embodiments. For a detailed discussion of steps S301 and S304 in this embodiment, please refer to the detailed discussion of steps S101 and S102 in the aforementioned embodiments; these details will not be repeated here.
[0091] S301, Obtain the target bitstream file. The target bitstream file is obtained by encoding the image to be compressed using the first encoding model. The first encoding model and the first decoding model are trained in the following way: the first bitstream file output by the basic encoding model after encoding the first training image is used as the input of the first basic decoding model; the first basic decoding model is used to decode the first bitstream file to obtain the first reconstructed image; based on the difference between the first reconstructed image and the first training image, the basic encoding model and the first basic decoding model are tuned to obtain the first encoding model and the first decoding model that meets the first target.
[0092] S302, Obtain the display requirements for the target image. The target image refers to the image obtained by encoding the target bitstream file.
[0093] The display effect of the target image can be understood as its quality. For example, in terms of quality, besides being similar to the image to be compressed, the target image should have clearer details and higher saturation for subjects such as people or animals compared to the image to be compressed. Display requirements correspond to the display effect; for example, in terms of quality, besides needing to be similar to the image to be compressed, the target image should have clearer details and higher saturation for subjects such as people or animals compared to the image to be compressed.
[0094] The display requirements for the target image can be set manually, or they can be determined by modules, systems, or devices that identify basic image attributes such as brightness, saturation, and sharpness of the image to be compressed, and then use these identified attributes to determine the display requirements for the target image obtained after encoding and decoding. Alternatively, modules, systems, or devices can determine the display requirements based on the intended use of the target image obtained after encoding and decoding. This application does not limit the subject or method of determining the display requirements for the target image.
[0095] S303, if the display effect of the decoded image output by the second decoding model meets the display requirements, determine that the second decoding model is the model for decoding the target bitstream file.
[0096] Since both the second and first decoding models can decode the target bitstream file encoded by the first encoding model, and multiple optimized decoding models can be obtained through iterative optimization based on different training objectives, with the second decoding model being any one of these optimized models, it is possible to use multiple decoding models to decode the target bitstream file. Given that multiple decoding models can be used to decode the target bitstream file, the training objectives of each decoding model are known, and the display requirements for the target image are obtained, the target decoding model can be adaptively determined to decode the target bitstream file based on these display requirements. The decoded image output by this target decoding model will then meet the display requirements.
[0097] Based on this, if the display effect of the decoded image output by the second decoding model meets the display requirements, the second decoding model is determined to be the model for decoding the target bitstream file so that the target image has the required display effect.
[0098] Optionally, if the target image has multiple display requirements, such as needing clearer details for the target object (like a person or animal) compared to the image to be compressed, and also needing higher saturation, then the importance of each display requirement is determined, identifying the most important, second most important, and so on. The importance of each display requirement can be manually set or determined using preset rules. Once the importance of each display requirement is determined, an adaptive decoding model for the first target is used to decode the target bitstream file, ensuring the decoded image output by this model meets the most important display requirement. Then, an adaptive decoding model for the second target is used to decode the bitstream file corresponding to the target image decoded by the first target's decoding model, ensuring the decoded image output by this model meets the second most important display requirement. This process continues.
[0099] S304, the target bitstream file is decoded using the second decoding model to obtain the target image. The second decoding model is trained as follows: the second bitstream file output by the first encoding model after encoding the second training image is used as the input of the second basic decoding model; the second basic decoding model is used to decode the second bitstream file to obtain the second reconstructed image; based on the difference between the second reconstructed image and the second training image, the parameters of the second basic decoding model are tuned to obtain a second decoding model that meets the second objective; wherein, the first objective of training the first basic decoding model is different from the second objective of training the second basic decoding model.
[0100] In this embodiment, when multiple decoding models are available that can be used to decode the target bitstream file, the training objectives of each decoding model are known, and the display requirements for the target image are obtained, the target decoding model can be adaptively determined to decode the target bitstream file based on the display requirements for the target image. The display effect of the decoded image output by the target decoding model meets the display requirements.
[0101] In one embodiment, the second decoding model that meets the second objective is obtained by tuning the parameters of the second basic decoding model based on a loss, which is calculated using a target loss function based on the difference between the second reconstructed image and the second training image; wherein, at least one of the target training set and the target loss function, including the second training image, is matched with the second objective.
[0102] In the image decoding method of this embodiment, at least one of a specific training set and a specific loss function is set to meet the training objective required by the trained decoding model. A second basic decoding model is trained using the target training set and the target loss function, so that the trained second decoding model conforms to the specific training objective, i.e., it conforms to the second objective. Here, at least one of the target training set and the target loss function matches the second objective. It should be noted that "matching" can be understood as setting a specific target training set and / or a specific target loss function to distinguish the second objective conforming to the second decoding model from the training objectives conforming to other decoding models. A mismatch between the target training set and the second objective does not mean that a second decoding model conforming to the second objective cannot be trained using the target training set, but rather that the target training set used to train the second decoding model conforming to the second objective may be the same as the training set used to train other decoding models; this target training set is not a specific training set set set specifically set to meet the second objective.
[0103] Specifically, based on the first encoding model encoding the second training images in the target training set to obtain the target bitstream file, and the second basic decoding model decoding the target bitstream file to obtain the second reconstructed image, a loss is calculated using the target loss function based on the difference between the second reconstructed image and the second training image. Then, the parameters of the second basic decoding model are tuned based on the loss to obtain a second decoding model that meets the second objective. The second decoding model trained using the target training set and the target loss function thus meets the second objective.
[0104] In this implementation, the training objective required by the trained decoding model can be met by setting at least one of a specific training set and a specific loss function. Furthermore, different decoding models are trained using different training sets and / or different loss functions, resulting in different training objectives for each model, which facilitates iterative optimization among the decoding models.
[0105] In one embodiment,
[0106] If the first target attribute of the decoded image generated by the second decoding model meets the image expectation requirement, where the first target attribute refers to any image attribute that affects the visual effect of the image, the target loss function includes information to judge the difference between the first target attribute of the second reconstructed image generated by the second basic decoding model and the image expectation requirement.
[0107] Image attributes include, but are not limited to, pixels, resolution, size, color, bit depth, hue, saturation, brightness, sharpness, contrast, color channels, and image hierarchy. The first target attribute is any image attribute that can affect the visual effect of the image, such as saturation, brightness, sharpness, and resolution. The image expectation requirement refers to the pre-defined expected value of the first target attribute of the decoded image.
[0108] To determine which approach—specific training set, specific loss function, or a combination of both—is needed to meet the training objective of the trained decoding model, the type of training objective must first be identified. The first type of objective is to further optimize the primary target attribute of the decoded image generated by the decoding model. For example, the primary target attribute of the decoded image generated by the first decoding model before optimization iterations is an initial value. The goal is to achieve the desired primary target attribute of the decoded image generated by the decoding model through optimization iterations, i.e., to reach a preset expected value. Here, the expected value is not equal to the initial value. For this type of objective, setting a specific training set has little impact on the display effect of the reconstructed image. However, setting a specific loss function, so that the loss calculation focuses more on the loss of the reconstructed image regarding the primary target attribute, has a significant impact on the display effect of the reconstructed image. Therefore, when the second objective is for the primary target attribute of the decoded image generated by the second decoding model to meet the desired image requirement, a specific target loss function can be set to ensure that the second decoding model meets the requirements of the second objective. The target training set used to train the second decoding model may be the same as or different from the training set used to train other decoding models. In this embodiment, no special restrictions are imposed on the target training set.
[0109] Specifically, the target loss function is L = R + a * D + b_1 * D_1. Here, R represents the bitrate; D is used to judge the degree of difference between the second reconstructed image and the second training image, and can be represented by the mean-square error (MSE) function, multi-scale structural similarity (MS-SSIM) function, etc.; a is a coefficient greater than 0 and less than 1. Since the training objective of the decoding model is usually based on the premise that the decoded image is still similar to the initial image before encoding and decoding, the training objective of the decoding model must at least satisfy the requirement that the decoded image is similar to the initial image. Based on this, "R + a * D" in the target loss function is still used to calculate the first loss based on the difference between the second reconstructed image and the second training image. Specifically, D_1 is used to judge the degree of difference between the first target attribute of the second reconstructed image and the image expectation requirement; b_1 is a coefficient matching D_1. That is, "b_1 * D_1" in the target loss function is used to calculate the second loss based on the difference between the first target attribute of the second reconstructed image and the image expectation requirement. For example, if the first target attribute is sharpness, then D_1 in the target loss function is used to determine the degree of difference between the sharpness of the second reconstructed image and the expected sharpness value, and "b_1*D_1" in the target loss function is used to calculate the second loss based on the difference between the sharpness of the second reconstructed image and the expected sharpness value. Using the target loss function, the first loss is calculated based on the difference between the second reconstructed image and the second training image, and the second loss is calculated based on the difference between the first target attribute of the second reconstructed image and the image expectation requirement, resulting in the final comprehensive loss. Then, the parameters of the second basic decoding model are tuned using the final comprehensive loss to obtain the second decoding model. Since the second loss of the second reconstructed image with respect to the first target attribute can be considered when calculating the comprehensive loss using the target loss function, the parameters of the second basic decoding model are adjusted based on the first and second losses to make the first and second losses converge. The finally trained second decoding model can then meet the training objective of achieving the image expectation requirement for the first target attribute of the decoded image.
[0110] In this embodiment, when the first target attribute of the decoded image generated by the second decoding model meets the image expectation requirements, a specific target loss function is set to ensure that the second decoding model meets the second target. Specifically, the target loss function includes information on the difference between the second reconstructed image and the second training image, as well as information on the difference between the first target attribute of the second reconstructed image and the image expectation requirements. Through this specific target loss function, a second decoding model that meets the second target can be trained, achieving iterative optimization of the decoding model.
[0111] In one embodiment,
[0112] When the second target is the region where the target object in the decoded image generated by the second decoding model meets the local expectation requirement, the target loss function includes information on detecting the target object in the second reconstructed image generated by the second basic decoding model and information on judging the difference between the second target attribute and the local expectation requirement of the region where the target object in the second reconstructed image is located. The number of second training images with target content in the target training set is greater than or equal to the first value. Having target content means having a preset number or more target objects.
[0113] Here, the area where the target object is located can be understood as the area where the target object, such as a human figure or animal, is located. Image attributes can be referred to the explanation in the previous embodiments. The second target attribute is any image attribute that can affect the visual effect of the image, such as saturation, brightness, sharpness, resolution, etc. The local expectation requirement refers to the preset expected value of the second target attribute of the area where the target object is located in the decoded image.
[0114] To determine which of the following—setting a specific training set, setting a specific loss function, or setting both a specific training set and a specific loss function—is necessary to meet the training objective requirements of the trained decoding model, the type of training objective the decoding model needs to meet must first be determined. The second type of objective involves further optimizing the secondary target attribute of the region containing the target object in the decoded image generated by the decoding model, enhancing the detail focus of the target object region and making the details of the target object region clearer. For example, the secondary target attribute of the region containing the target object in the decoded image generated by the first decoding model before optimization iteration is the initial value. The goal is to achieve a local expectation requirement, i.e., a preset expected value, through decoding model optimization iteration. The expected value is not equal to the initial value. For this type of objective, a specific loss function can be set so that when calculating the loss using the loss function, more attention is paid to the loss of the secondary target attribute of the target object region in the decoded reconstructed image. A specific loss function has a significant impact on the display effect of the reconstructed image. Furthermore, to enhance the effect of the specific loss function and ensure that it better focuses on the loss of the second target attribute of the region where the target object is located in the reconstructed image during the training of the decoding model, the reconstructed image needs to contain the target object, and the more target objects, the better. Therefore, it is necessary to set a specific training set containing a large number of target objects to achieve the desired reconstructed image with more target objects after encoding and decoding. Thus, when the second target attribute of the region where the target object is located in the decoded image generated by the second decoding model meets the local expectation requirement, a specific target training set and a specific target loss function can be set to ensure that the second decoding model meets the requirements of the second target.
[0115] Specifically, the target loss function is L = R + a*D + b_2*D_2. Referring to the detailed discussion in the preceding embodiments, "R + a*D" in the target loss function is still used to calculate the first loss based on the difference between the second reconstructed image and the second training image. Specifically, D_2 is used to detect the target object in the second reconstructed image, and when a target object is detected, it determines the degree of difference between the second target attribute of the region where the target object is located in the second reconstructed image and the local expectation requirement; b_2 is a coefficient that matches D_2. That is, "b_2*D_2" in the target loss function is used to calculate the second loss based on the difference between the second target attribute of the region where the target object is located in the second reconstructed image and the local expectation requirement. For example, if the target object is a human image and the second target attribute is resolution, then D_2 in the target loss function is used to detect the human image in the second reconstructed image. If a human image is detected, the degree of difference between the resolution of the region where the human image is located in the second reconstructed image and the expected resolution value is determined. "b_2*D_2" in the target loss function is used to calculate the second loss based on the difference between the resolution of the region where the human image is located in the second reconstructed image and the expected resolution value. Using the target loss function, the first loss is calculated based on the difference between the second reconstructed image and the second training image, and the second loss is calculated based on the difference between the second target attribute of the region where the target object is located in the second reconstructed image and the local expected value, resulting in the final comprehensive loss. Then, the parameters of the second basic decoding model are tuned using the final comprehensive loss to obtain the second decoding model. Because the comprehensive loss is calculated using the target loss function, the second loss of the second target attribute of the region where the target object is located in the second reconstructed image can be considered, reducing the compression loss of the region where the target object is located. Furthermore, the parameters of the second basic decoding model are adjusted based on the first and second losses to make the first and second losses converge. The second decoding model, once trained, is able to meet the training objective of achieving the local expectation requirement for the second target attribute of the region where the target object is located in the decoded image.
[0116] Accordingly, the number of second training images containing target content in the target training set is greater than or equal to the first value. "Containing target content" refers to having a preset number or more target objects. To ensure that the second reconstructed image obtained after encoding and decoding the second training images contains target objects, and contains a large number of target objects, thereby strengthening the effect of the target loss function, both the first value and the preset number are more advantageous.
[0117] In this embodiment, when the second target attribute of the region where the target object in the decoded image generated by the second decoding model meets the local expectation requirement, a specific target loss function and a specific target training set are set to satisfy the requirement that the second decoding model meets the second target. Specifically, the target loss function includes information on the difference between the second reconstructed image and the second training image, as well as information on detecting the target object in the second reconstructed image and information on the difference between the second target attribute of the region where the target object in the second reconstructed image is located and the local expectation requirement. The number of second training images with target content in the target training set is greater than or equal to a first value, where "with target content" means having a preset number or more target objects. Through this specific target loss function, a second decoding model that meets the second target can be trained, realizing the optimization and iteration of the decoding model.
[0118] In one embodiment,
[0119] The second objective is that when the third objective attribute of the initial image is lower than the preset requirement, the second decoding model decodes the bitstream file corresponding to the initial image, and the similarity between the decoded image and the initial image reaches the expected similarity value. The third objective attribute refers to any image attribute that affects the visual effect of the image. In this case, the number of second training images in the target training set whose third objective attribute is lower than the preset requirement is greater than or equal to the second value.
[0120] The image attributes can be referred to the explanation in the previous embodiments. The third target attribute is any image attribute that can affect the visual effect of the image, such as saturation, brightness, sharpness, resolution, etc. The preset requirement refers to the preset base value of the third target attribute of the initial image.
[0121] To determine which of the following—setting a specific training set, setting a specific loss function, or setting both a specific training set and a specific loss function—is necessary to meet the training objective requirements of the trained decoding model, the type of training objective the decoding model needs to meet must first be determined. The third objective type occurs when the third objective attribute of the initial image is lower than a preset requirement. In this case, the second decoding model decodes the bitstream file corresponding to the initial image, and the resulting decoded image still achieves the expected similarity value to the initial image. For example, when the third objective attribute of the initial image is lower than the preset requirement, the decoded image generated by the first decoding model before optimization iteration fails to achieve the expected similarity value to the initial image, indicating a decoding performance deviation. Therefore, the goal is to optimize the decoding model iteratively so that, even when the third objective attribute of the initial image is lower than the preset requirement, the decoded image generated by the decoding model still achieves the expected similarity value to the initial image, thus maintaining a good decoding performance. For this objective type, setting a specific loss function has a relatively small impact on the display effect of the reconstructed image obtained from the decoding. However, by setting a specific training set that includes a large number of training images with third target attributes lower than the preset requirement, the basic decoding model can be trained to decode special bitstream files. Special bitstream files refer to the bitstream files corresponding to the training images with third target attributes lower than the preset requirement. Therefore, a decoding model can be trained that can better decode the bitstream files corresponding to the initial images with third target attributes lower than the preset requirement, generating decoded images similar to the initial images. Here, "similar to the initial image" means that the similarity to the initial image reaches the expected similarity value. Based on this, when the second objective is that when the third target attribute of the initial image is lower than the preset requirement, the second decoding model decodes the bitstream file corresponding to the initial image, and the resulting decoded image achieves the expected similarity value to the initial image, a specific target training set can be set to ensure that the second decoding model meets the requirements of the second objective. The target loss function used to train the second decoding model may be the same as or different from the loss function used to train other decoding models; this embodiment does not impose any special restrictions on the target loss function.
[0122] Specifically, the number of second training images in the target training set whose third target attribute is lower than a preset requirement is greater than or equal to a second value. A larger second value is more advantageous for better training the second basic decoding model's decoding ability for the bitstream files corresponding to the second training images whose third target attribute is lower than the preset requirement. For example, if the third target attribute is brightness, the number of second training images in the target training set whose brightness is lower than the preset requirement is greater than or equal to the second value to train the second basic decoding model's decoding ability for the bitstream files corresponding to the low-brightness second training images. Based on this, the trained second decoding model can perform better decoding of the bitstream files corresponding to the initial images whose third target attribute is lower than the preset requirement, generating decoded images similar to the initial images. That is, the second decoding model can meet the second target.
[0123] In this embodiment, when the third target attribute of the initial image (the second target) is lower than a preset requirement, the second decoding model decodes the bitstream file corresponding to the initial image. If the similarity between the resulting decoded image and the initial image reaches the expected similarity value, a specific target training set is set to ensure the second decoding model meets the second target requirement. Specifically, the number of second training images in the target training set whose third target attribute is lower than the preset requirement is greater than or equal to a second value. Through this specific target training set, a second decoding model that meets the second target can be trained, achieving iterative optimization of the decoding model.
[0124] Taking the initial encoding model E_a and the initial decoding model D_a as an example, E_a and D_a are obtained through joint training, and both E_a and D_a meet the initial objectives. Without needing to solve the technical problem of incompatibility between bitstream files encoded by E_a, i.e., without needing multiple different decoding models to decode bitstream files encoded by E_a, the encoding and decoding models can be optimized iteratively based on more and more complex training objectives. Specifically, to ensure the encoding and decoding models meet the iterative objectives, such as further improving the speed of encoding and decoding the initial image to obtain the decoded image, or further improving the detection accuracy of each object in the decoded image, the encoding and decoding models can be optimized iteratively to obtain the iterative encoding model E_c and the iterative decoding model D_c. E_c and D_c are obtained through joint training. Although D_c cannot decode the bitstream file encoded by E_a, E_c and D_c meet the iterative objectives. Furthermore, using E_c as the first encoding model and D_c as the first decoding model, based on the optimization iterations of the above encoding and decoding models, the decoding model can still be optimized and iterated based on D_c while keeping E_c unchanged, training D_c1, D_c2, D_c3, etc., with different training objectives. The second decoding model can be any one of D_c1, D_c2, D_c3, etc. Optimizing and iterating from E_a and D_a to obtain E_c and D_c, and then further optimizing E_c and D_c1, E_c and D_c2, E_c and D_c3, etc., may simultaneously satisfy both the iteration objective and the second objective, and D_c1, D_c2, D_c3, etc., can all decode the bitstream file encoded by E_c. In addition, compared to E_a before optimization iteration, the bitstream file generated by E_c encoding has higher performance. Based on this, when the optimized and iterated second decoding model decodes the bitstream file generated by E_c encoding, it helps to improve the performance of the obtained decoded image and helps the decoded image better achieve the image display effect corresponding to the second target.
[0125] Based on E_a and D_a, in the process of optimizing and iterating to obtain E_c and D_c that meet the iterative objective, at least one of the training set, loss function, and basic encoding / decoding model used must match the iterative objective. Specifically, training images from the training set are input into the basic encoding / decoding model to encode and decode them into reconstructed images. Then, the loss function is used to calculate the loss based on the difference between the reconstructed image and the training image, thereby tuning the parameters of the basic encoding / decoding model to obtain E_c and D_c that meet the iterative objective. For example, when the iterative objective is to further improve the detection accuracy of each object in the decoded image by the encoding / decoding model, a specific loss function is set to calculate the loss based on the detection accuracy of each object in the reconstructed image obtained by encoding and decoding; when the iterative objective is to further improve the speed of the encoding / decoding model in encoding and decoding the initial image to obtain the decoded image, the number of parameters in the basic encoding / decoding model is reduced to accelerate the training speed, thereby accelerating the image reconstruction speed.
[0126] This application also provides a video frame display method. This video frame display method is based on the image decoding method described in the foregoing embodiments, and uses video frames as images to be compressed. It decodes the target bitstream file formed after the video frames are encoded to obtain the target image, and then displays the target image. Therefore, for explanations of terms, detailed discussions of each step, derivation of technical effects, and examples of feasible embodiments in the video frame display method provided in this application, please refer to the descriptions of the corresponding content in the foregoing image decoding method embodiments. Further details are not repeated in this application.
[0127] This video frame display method can be applied to any application scenario that requires the display of video frames, such as live video streaming and video-on-demand. This application does not limit the application scenarios of this video frame display method. Figure 4 As shown, the video frame display method includes the following steps S401 to S403.
[0128] S401, Obtain the target bitstream file formed after the video frames are encoded. The target bitstream file is obtained by encoding the video frames using the first encoding model. The first encoding model and the first decoding model are trained in the following way: the first bitstream file output by the basic encoding model after encoding the first training image is used as the input of the first basic decoding model; the first basic decoding model is used to decode the first bitstream file to obtain the first reconstructed image; based on the difference between the first reconstructed image and the first training image, the basic encoding model and the first basic decoding model are tuned to obtain the first encoding model and the first decoding model that meets the first target.
[0129] S402, the target bitstream file is decoded using the second decoding model to obtain the target image. The second decoding model is trained as follows: the second bitstream file output by the first encoding model after encoding the second training image is used as the input of the second basic decoding model; the second basic decoding model is used to decode the second bitstream file to obtain the second reconstructed image; based on the difference between the second reconstructed image and the second training image, the parameters of the second basic decoding model are tuned to obtain a second decoding model that meets the second objective; wherein, the first objective of training the first basic decoding model is different from the second objective of training the second basic decoding model.
[0130] S403, Display target image.
[0131] In this embodiment, a target bitstream file obtained by encoding video frames using a first encoding model is acquired, and the target bitstream file is decoded using a second decoding model to obtain a target image. This enables multiple different decoding models to decode bitstream files generated by the same encoding model during the optimization and iteration process of the encoding and decoding models. The second version of the second decoding model can decode the target bitstream file generated by the first version of the first encoding model. Furthermore, this avoids situations where video frames and their corresponding target bitstream files become invalid due to bitstream file incompatibility, difficulty in decoding the target bitstream file generated by the first encoding model, or the inability to use the first decoding model. By decoding and displaying the target image, the display of the video frame is achieved.
[0132] This application also provides a decoding model training method. The second decoding model trained by this method, conforming to a second objective, is inferred and applied in the image decoding method described in the foregoing embodiments. That is, this decoding model training method describes the training process of the second decoding model, the image decoding method described in the foregoing embodiments describes the inference and application process of the second decoding model, and the training process of the second decoding model is also described when discussing the inference and application process. Therefore, for the explanation of terms, specific descriptions of each step, derivation of technical effects, and examples of feasible embodiments in the decoding model training method provided in this application, please refer to the descriptions of the corresponding content in the foregoing image decoding method embodiments. Further details are not repeated in this application.
[0133] like Figure 5 As shown, the decoding model training method includes the following steps S501 to S504.
[0134] S501, Obtain the target training set.
[0135] S502, the second training image in the target training set is encoded using the first encoding model to obtain a second bitstream file. The first encoding model and the first decoding model are trained as follows: the first bitstream file output by the basic encoding model after encoding the first training image is used as the input of the first basic decoding model; the first basic decoding model is used to decode the first bitstream file to obtain the first reconstructed image; based on the difference between the first reconstructed image and the first training image, the basic encoding model and the first basic decoding model are tuned to obtain the first encoding model and the first decoding model that conforms to the first target.
[0136] S503, input the second bitstream file into the second basic decoding model to obtain the second reconstructed image.
[0137] S504, based on the difference between the second reconstructed image and the second training image, the parameters of the second basic decoding model are tuned to obtain a second decoding model that meets the second objective. The first objective of training the first basic decoding model is different from the second objective of training the second basic decoding model.
[0138] In this embodiment, the first encoding model and the first decoding model are obtained through joint training. Then, while keeping the first encoding model unchanged, a second decoding model is trained based on the first encoding model. The first objective met by the first decoding model differs from the second objective met by the second decoding model. The decoding model training method in this embodiment not only achieves iterative optimization of the encoding and decoding models, but also enables multiple different decoding models to decode the bitstream file generated by the same encoding model during the optimization process, and allows the second version of the second decoding model to decode the target bitstream file generated by the first version of the first encoding model.
[0139] In one embodiment, when the first target attribute of the decoded image generated by the second decoding model meets the desired image requirements, the first target attribute refers to any image attribute that affects the visual effect of the image, such as... Figure 6 As shown, step S504 above, based on the difference between the second reconstructed image and the second training image, adjusts the parameters of the second basic decoding model to obtain a second decoding model that meets the second objective, including the following steps S601 to S602.
[0140] S601, calculate the loss based on the difference between the second reconstructed image and the second training image, and the difference between the first target attribute of the second reconstructed image and the image expectation requirement.
[0141] Referring to the specific discussion in the aforementioned image decoding method regarding the embodiment where "when the first target attribute of the decoded image generated by the second decoding model meets the image expectation requirement (where the first target attribute refers to any image attribute that affects the visual effect of the image), the target loss function includes information on the difference between the first target attribute of the second reconstructed image generated by the second base decoding model and the image expectation requirement," it can be seen that when the first target attribute of the decoded image generated by the second decoding model meets the image expectation requirement, a specific target loss function is set to satisfy the requirement that the second decoding model meets the second target. Using the target loss function, the loss is calculated based on the difference between the second reconstructed image and the second training image, and the difference between the first target attribute of the second reconstructed image and the image expectation requirement.
[0142] S602, based on the loss, the parameters of the second basic decoding model are tuned to obtain a second decoding model that meets the second objective.
[0143] Given that the first target attribute of the decoded image generated by the second decoding model meets the desired image requirements, the loss is calculated based on the two differences mentioned above. This calculation considers not only the similarity between the second reconstructed image and the second training image but also the loss of the second reconstructed image regarding the first target attribute. Furthermore, the second basic decoding model is tuned using the loss calculated based on these two differences until the loss converges. The resulting second decoding model then satisfies the second objective of ensuring that the first target attribute of the decoded image meets the desired image requirements.
[0144] In one embodiment, when the second target is the region where the target object in the decoded image generated by the second decoding model is located, and the second target attribute refers to any image attribute that affects the visual effect of the image, such as... Figure 7 As shown, step S504 above, based on the difference between the second reconstructed image and the second training image, tunes the parameters of the second basic decoding model to obtain a second decoding model that meets the second objective, including the following steps S701 to S702.
[0145] S701, calculate the loss based on the difference between the second reconstructed image and the second training image, and the difference between the second target attribute and the local expectation requirement of the region where the target object is located in the second reconstructed image.
[0146] Referring to the specific discussion in the aforementioned image decoding method regarding the embodiment where "the second target attribute of the region where the target object is located in the decoded image generated by the second decoding model meets the local expectation requirement, where the second target attribute refers to any image attribute that affects the visual effect of the image, the target loss function includes information on detecting the target object in the second reconstructed image generated by the second basic decoding model, and information on judging the difference between the second target attribute of the region where the target object is located in the second reconstructed image and the local expectation requirement, and the number of second training images with target content in the target training set is greater than or equal to a first value, where having target content means having a preset number or more target objects," it can be seen that in order to meet the requirement that the second decoding model meets the second target, a specific target loss function needs to be set. Using the target loss function, the loss is calculated based on the difference between the second reconstructed image and the second training image, and the difference between the second target attribute of the region where the target object is located in the second reconstructed image and the local expectation requirement.
[0147] S702, based on the loss, the second basic decoding model is tuned to obtain a second decoding model that meets the second objective.
[0148] When the second objective is that the second target attribute of the region where the target object is located in the decoded image generated by the second decoding model meets the local expectation requirement, the loss is calculated based on the two differences mentioned above. This not only considers the similarity between the second reconstructed image and the second training image, but also focuses on the loss of the second reconstructed image regarding the second target attribute of the region where the target object is located. Furthermore, the second basic decoding model is tuned using the loss calculated based on the two differences mentioned above until the loss converges. The final trained second decoding model can then satisfy the second objective that the second target attribute of the region where the target object is located in the decoded image meets the local expectation requirement.
[0149] like Figure 8As shown in the figure, this application embodiment also provides an image decoding device 800. The image decoding device 800 includes a bitstream acquisition module 801 and a bitstream decoding module 802. The bitstream acquisition module 801 is used to acquire a target bitstream file, which is obtained by encoding the image to be compressed using a first encoding model. The first encoding model and the first decoding model are trained in the following way: the first bitstream file output by the basic encoding model after encoding the first training image is used as the input of the first basic decoding model; the first basic decoding model is used to decode the first bitstream file to obtain a first reconstructed image; based on the difference between the first reconstructed image and the first training image, the basic encoding model and the first basic decoding model are tuned to obtain the first encoding model and the first decoding model that meets the first target. The bitstream decoding module 802 is used to decode the target bitstream file using a second decoding model to obtain the target image. The second decoding model is trained in the following way: the second bitstream file output by the first encoding model after encoding the second training image is used as the input of the second basic decoding model; the second basic decoding model is used to decode the second bitstream file to obtain the second reconstructed image; based on the difference between the second reconstructed image and the second training image, the parameters of the second basic decoding model are tuned to obtain a second decoding model that meets the second objective; wherein, the first objective of training the first basic decoding model is different from the second objective of training the second basic decoding model.
[0150] In one embodiment, the second objective of training the second basic decoding model matches the display effect of the decoded image output by the second decoding model. The image decoding apparatus 800 further includes a requirement acquisition module and a model determination module. The requirement acquisition module is used to acquire the display requirements for the display effect of the target image, where the target image refers to the image obtained by encoding the target bitstream file. The model determination module is used to determine, if the display effect of the decoded image output by the second decoding model meets the display requirements, to use the second decoding model as the model for decoding the target bitstream file.
[0151] In one embodiment, the second decoding model that meets the second objective is obtained by tuning the parameters of the second basic decoding model based on a loss, which is calculated using a target loss function based on the difference between the second reconstructed image and the second training image; wherein, at least one of the target training set and the target loss function, including the second training image, is matched with the second objective.
[0152] In one embodiment, when the first target attribute of the decoded image generated by the second decoding model meets the image expectation requirement, where the first target attribute refers to any image attribute that affects the visual effect of the image, the target loss function includes information to determine the difference between the first target attribute of the second reconstructed image generated by the second basic decoding model and the image expectation requirement.
[0153] In one embodiment, when the second target is the region where the target object in the decoded image generated by the second decoding model meets the local expectation requirement, and the second target attribute refers to any image attribute that affects the visual effect of the image, the target loss function includes information on detecting the target object in the second reconstructed image generated by the second basic decoding model, and information on judging the difference between the second target attribute of the region where the target object in the second reconstructed image is located and the local expectation requirement. The number of second training images with target content in the target training set is greater than or equal to a first value. Having target content means having a preset number or more target objects.
[0154] In one embodiment, when the second objective is that the third objective attribute of the initial image is lower than a preset requirement, the second decoding model decodes the bitstream file corresponding to the initial image, and the similarity between the decoded image and the initial image reaches the expected similarity value. The third objective attribute refers to any image attribute that affects the visual effect of the image. In this case, the number of second training images in the target training set whose third objective attribute is lower than the preset requirement is greater than or equal to the second value.
[0155] For explanations of the terms, please refer to the relevant descriptions in the aforementioned image decoding method embodiments, which will not be elaborated here.
[0156] It should be noted that the specific execution process of the aforementioned image decoding device 800 can be found in [reference needed]. Figures 1 to 3 The specific details of the embodiments shown are not elaborated here.
[0157] Each module in the aforementioned image decoding device 800 can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0158] like Figure 9As shown in the figure, this application embodiment also provides a video frame display device 900. The video frame display device 900 includes a target bitstream acquisition module 901, a target image generation module 902, and a target image display module 903. The target bitstream acquisition module 901 is used to acquire the target bitstream file formed after the video frame is encoded. The target bitstream file is obtained by encoding the video frame using a first encoding model. The first encoding model and the first decoding model are trained in the following way: the first bitstream file output by the basic encoding model after encoding the first training image is used as the input of the first basic decoding model; the first basic decoding model is used to decode the first bitstream file to obtain a first reconstructed image; based on the difference between the first reconstructed image and the first training image, the basic encoding model and the first basic decoding model are tuned to obtain the first encoding model and the first decoding model that meets the first target. The target image generation module 902 is used to decode the target bitstream file using a second decoding model to obtain the target image. The second decoding model is trained as follows: the second bitstream file output by the first encoding model after encoding the second training image is used as the input of the second basic decoding model; the second basic decoding model is used to decode the second bitstream file to obtain the second reconstructed image; based on the difference between the second reconstructed image and the second training image, the parameters of the second basic decoding model are tuned to obtain a second decoding model that meets the second objective; wherein, the first objective of training the first basic decoding model is different from the second objective of training the second basic decoding model. The target image display module 903 is used to display the target image.
[0159] For explanations of the terms, please refer to the relevant descriptions in the aforementioned video frame display method embodiments, which will not be elaborated here.
[0160] It should be noted that the specific execution process of the aforementioned video frame display device 900 can be found in [reference needed]. Figure 4 The specific details of the embodiments shown are not elaborated here.
[0161] Each module in the aforementioned video frame display device 900 can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.
[0162] like Figure 10As shown in the illustration, this application also provides a decoding model training device 1000. The decoding model training device 1000 includes a training set acquisition module 1001, a bitstream generation module 1002, an image reconstruction module 1003, and a model generation module 1004. The training set acquisition module 1001 is used to acquire a target training set. The bitstream generation module 1002 is used to encode a second training image in the target training set using a first encoding model to obtain a second bitstream file. The first encoding model and the first decoding model are trained in the following manner: the first bitstream file output by the basic encoding model after encoding the first training image is used as the input of a first basic decoding model; the first basic decoding model is used to decode the first bitstream file to obtain a first reconstructed image; based on the difference between the first reconstructed image and the first training image, the basic encoding model and the first basic decoding model are tuned to obtain a first encoding model and a first decoding model that meets the first target. The image reconstruction module 1003 is used to input the second bitstream file into the second basic decoding model to obtain a second reconstructed image. The model generation module 1004 is used to tune the parameters of the second basic decoding model based on the difference between the second reconstructed image and the second training image to obtain a second decoding model that meets the second objective; wherein, the first objective of training the first basic decoding model is different from the second objective of training the second basic decoding model.
[0163] In one embodiment, when the first target attribute of the decoded image generated by the second decoding model meets the image expectation requirement, where the first target attribute refers to any image attribute that affects the visual effect of the image, the model generation module 1004 is further used to calculate the loss based on the difference between the second reconstructed image and the second training image, and the difference between the first target attribute of the second reconstructed image and the image expectation requirement; and to tune the parameters of the second basic decoding model based on the loss to obtain a second decoding model that meets the second target.
[0164] In one embodiment, when the second target attribute of the region where the target object is located in the decoded image generated by the second decoding model meets the local expectation requirement, and the second target attribute refers to any image attribute that affects the visual effect of the image, the model generation module 1004 is further used to calculate the loss based on the difference between the second reconstructed image and the second training image, and the difference between the second target attribute of the region where the target object is located in the second reconstructed image and the local expectation requirement; and to tune the parameters of the second basic decoding model based on the loss to obtain a second decoding model that meets the second target.
[0165] For explanations of the terms, please refer to the relevant descriptions in the aforementioned examples of the decoding model training method, which will not be elaborated here.
[0166] It should be noted that the specific execution process of the aforementioned decoding model training device 1000 can be found in [reference needed]. Figures 5 to 7 The specific details of the embodiments shown are not elaborated here.
[0167] Each module in the aforementioned decoding model training device 1000 can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0168] like Figure 11 As shown in the illustration, this application also provides a computer device 1100. Exemplarily, the computer device 1100 may include a processor 1101, a communication interface 1102, a communication bus 1103, and a memory 1104. Specifically, the computer device 1100 may include:
[0169] The system includes at least one processor 1101, such as a CPU, at least one communication interface 1102, a memory 1104, and at least one communication bus 1103. The communication bus 1103 is used to enable communication between these components. The communication interface 1102 may optionally include a standard wired interface, a wireless interface (such as a Wi-Fi interface or a Bluetooth interface), etc. The memory 1104 may be high-speed RAM or non-volatile memory, such as at least one disk storage device. Optionally, the memory 1104 may also be at least one storage device located remotely from the aforementioned processor 1101. Figure 11 As shown, the memory 1104, which serves as a computer storage medium, may include an operating system and program instructions.
[0170] For example, processor 1101 can be used to implement the above. Figure 8 The steps or methods executed by the stream acquisition module 801 and the stream decoding module 802 in the middle.
[0171] Understandably, the above method is merely an example, and the above can also be executed by the processor 1101 and other modules in the aforementioned computer device 1100. Figure 8 The steps or methods executed by the stream acquisition module 801 and the stream decoding module 802 are not limited in this document.
[0172] exist Figure 11 In the computer device 1100 shown, the processor 1101 can be used to load program instructions stored in the memory 1104 and specifically perform the following operations:
[0173] The target bitstream file is obtained by encoding the image to be compressed using the first encoding model. The first encoding model and the first decoding model are trained in the following way: the first bitstream file output by the basic encoding model after encoding the first training image is used as the input of the first basic decoding model. The first basic decoding model is used to decode the first bitstream file to obtain the first reconstructed image. Based on the difference between the first reconstructed image and the first training image, the basic encoding model and the first basic decoding model are tuned to obtain the first encoding model and the first decoding model that meets the first target.
[0174] The target bitstream file is decoded using the second decoding model to obtain the target image. The second decoding model is trained in the following way: the second bitstream file output by the first encoding model after encoding the second training image is used as the input of the second basic decoding model. The second basic decoding model is used to decode the second bitstream file to obtain the second reconstructed image. Based on the difference between the second reconstructed image and the second training image, the parameters of the second basic decoding model are tuned to obtain a second decoding model that meets the second target.
[0175] The first objective of training the first basic decoding model is different from the second objective of training the second basic decoding model.
[0176] For explanations of the terms, please refer to the relevant descriptions in the aforementioned image decoding method embodiments, which will not be elaborated here.
[0177] It should be noted that the specific execution process can be found in [link to relevant documentation]. Figures 1 to 3 The specific details of the embodiments shown are not elaborated here.
[0178] For example, processor 1101 can be used to implement the above. Figure 9 The steps or methods executed by the target stream acquisition module 901, the target image generation module 902, and the target image display module 903.
[0179] Understandably, the above method is merely an example, and the above can also be executed by the processor 1101 and other modules in the aforementioned computer device 1100. Figure 9 The steps or methods executed by the target stream acquisition module 901, the target image generation module 902, and the target image display module 903 are not limited in this document.
[0180] exist Figure 11 In the computer device 1100 shown, the processor 1101 can be used to load program instructions stored in the memory 1104 and specifically perform the following operations:
[0181] The target bitstream file is obtained after the video frames are encoded. The target bitstream file is obtained by encoding the video frames using the first encoding model. The first encoding model and the first decoding model are trained in the following way: the first bitstream file output by the basic encoding model after encoding the first training image is used as the input of the first basic decoding model. The first basic decoding model is used to decode the first bitstream file to obtain the first reconstructed image. Based on the difference between the first reconstructed image and the first training image, the basic encoding model and the first basic decoding model are tuned to obtain the first encoding model and the first decoding model that meets the first target.
[0182] The target bitstream file is decoded using a second decoding model to obtain the target image. The second decoding model is trained as follows: the second bitstream file output by the first encoding model after encoding the second training image is used as the input of the second basic decoding model. The second basic decoding model is then used to decode the second bitstream file to obtain the second reconstructed image. Based on the difference between the second reconstructed image and the second training image, the parameters of the second basic decoding model are tuned to obtain a second decoding model that meets the second objective. The first objective of training the first basic decoding model is different from the second objective of training the second basic decoding model.
[0183] Display the target image.
[0184] For explanations of the terms, please refer to the relevant descriptions in the aforementioned video frame display method embodiments, which will not be elaborated here.
[0185] It should be noted that the specific execution process can be found in [link to relevant documentation]. Figure 4 The specific details of the embodiments shown are not elaborated here.
[0186] For example, processor 1101 can be used to implement the above. Figure 10 The steps or methods executed by the training set acquisition module 1001, the bitstream generation module 1002, the image reconstruction module 1003, and the model generation module 1004.
[0187] Understandably, the above method is merely an example, and the above can also be executed by the processor 1101 and other modules in the aforementioned computer device 1100. Figure 10 The steps or methods executed by the training set acquisition module 1001, the bitstream generation module 1002, the image reconstruction module 1003, and the model generation module 1004 are not limited in this paper.
[0188] exist Figure 11 In the computer device 1100 shown, the processor 1101 can be used to load program instructions stored in the memory 1104 and specifically perform the following operations:
[0189] Obtain the target training set;
[0190] The first encoding model is used to encode the second training image in the target training set to obtain the second bitstream file. The first encoding model and the first decoding model are trained in the following way: the first bitstream file output by the basic encoding model after encoding the first training image is used as the input of the first basic decoding model. The first basic decoding model is used to decode the first bitstream file to obtain the first reconstructed image. Based on the difference between the first reconstructed image and the first training image, the basic encoding model and the first basic decoding model are tuned to obtain the first encoding model and the first decoding model that meets the first target.
[0191] The second bitstream file is input into the second basic decoding model to obtain the second reconstructed image;
[0192] Based on the difference between the second reconstructed image and the second training image, the parameters of the second basic decoding model are tuned to obtain a second decoding model that meets the second objective.
[0193] The first objective of training the first basic decoding model is different from the second objective of training the second basic decoding model.
[0194] For explanations of the terms, please refer to the relevant descriptions in the aforementioned examples of the decoding model training method, which will not be elaborated here.
[0195] It should be noted that the specific execution process can be found in [link to relevant documentation]. Figures 5 to 7 The specific details of the embodiments shown are not elaborated here.
[0196] This application also provides a computer-readable storage medium that can store multiple instructions adapted to be loaded and executed by a processor as described above. Figures 1 to 7 The method steps of the illustrated embodiment can be found in the detailed execution process. Figures 1 to 7 The specific details of the illustrated embodiments will not be elaborated here.
[0197] As used in the above embodiments, depending on the context, the term "when..." can be interpreted as meaning "if...", "after...", "in response to determining...", or "in response to detecting...". Similarly, depending on the context, the phrase "when determining..." or "if (the stated condition or event) is detected" can be interpreted as meaning "if determining...", "in response to determining...", "when (the stated condition or event) is detected", or "in response to detecting (the stated condition or event)".
[0198] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. This computer program product includes one or more computer instructions. When these computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive), etc.
[0199] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.
[0200] The technical features of the above embodiments can be arbitrarily combined. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combination of these technical features does not contradict each other, they should be considered to be within the scope of this specification.
[0201] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. An image decoding method, characterized in that, The method includes: A target bitstream file is obtained, which is obtained by encoding the image to be compressed using a first encoding model. The first encoding model and the first decoding model are trained in the following way: the first bitstream file output by the basic encoding model after encoding the first training image is used as the input of the first basic decoding model; the first basic decoding model is used to decode the first bitstream file to obtain the first reconstructed image; based on the difference between the first reconstructed image and the first training image, the basic encoding model and the first basic decoding model are tuned to obtain the first encoding model and the first decoding model that meets the first target. The target bitstream file is decoded using a second decoding model to obtain a target image. The second decoding model is trained in the following way: the second bitstream file output by the first encoding model after encoding the second training image is used as the input of the second basic decoding model; the second basic decoding model is used to decode the second bitstream file to obtain a second reconstructed image; based on the difference between the second reconstructed image and the second training image, the parameters of the second basic decoding model are tuned to obtain a second decoding model that meets the second target. The first objective of training the first basic decoding model is different from the second objective of training the second basic decoding model.
2. The method according to claim 1, characterized in that, The second objective of training the second basic decoding model matches the display effect of the decoded image output by the second decoding model; Before decoding the target bitstream file using the second decoding model to obtain the target image, the method further includes: Obtain the display requirements for the display effect of the target image, wherein the target image refers to the image obtained by encoding the target bitstream file; If the display effect of the decoded image output by the second decoding model meets the display requirements, the second decoding model is determined to be the model for decoding the target bitstream file.
3. The method according to claim 1 or 2, characterized in that, The second decoding model that meets the second objective is obtained by tuning the parameters of the second basic decoding model based on a loss, wherein the loss is calculated using the target loss function based on the difference between the second reconstructed image and the second training image. The target training set, which includes the second training image, and the target loss function, are matched with the second target.
4. The method according to claim 3, characterized in that, When the second objective is that the first target attribute of the decoded image generated by the second decoding model meets the image expectation requirement, and the first target attribute refers to any image attribute that affects the visual effect of the image, the objective loss function includes information to judge the difference between the first target attribute of the second reconstructed image generated by the second basic decoding model and the image expectation requirement.
5. The method according to claim 3, characterized in that, When the second target is the region where the target object in the decoded image generated by the second decoding model meets the local expectation requirement, and the second target attribute refers to any image attribute that affects the visual effect of the image, the target loss function includes information on detecting the target object in the second reconstructed image generated by the second basic decoding model, and information on judging the difference between the second target attribute of the region where the target object in the second reconstructed image is located and the local expectation requirement. The number of second training images with target content in the target training set is greater than or equal to a first value, and having target content means having a preset number or more of the target objects.
6. The method according to claim 3, characterized in that, When the second objective is that the third objective attribute of the initial image is lower than the preset requirement, the second decoding model decodes the bitstream file corresponding to the initial image, and the similarity between the decoded image and the initial image reaches the expected similarity value. The third objective attribute refers to any image attribute that affects the visual effect of the image. In the case of the third objective attribute being lower than the preset requirement, the number of second training images in the objective training set is greater than or equal to the second value.
7. A method for displaying video frames, characterized in that, The method includes: A target bitstream file is obtained after video frames are encoded. The target bitstream file is obtained by encoding the video frames using a first encoding model. The first encoding model and the first decoding model are trained in the following way: the first bitstream file output by the basic encoding model after encoding the first training image is used as the input of the first basic decoding model; the first basic decoding model is used to decode the first bitstream file to obtain a first reconstructed image; based on the difference between the first reconstructed image and the first training image, the basic encoding model and the first basic decoding model are tuned to obtain the first encoding model and the first decoding model that meets the first target. The target bitstream file is decoded using a second decoding model to obtain a target image. The second decoding model is trained as follows: the second bitstream file output by the first encoding model after encoding the second training image is used as the input of the second basic decoding model; the second basic decoding model is used to decode the second bitstream file to obtain a second reconstructed image; based on the difference between the second reconstructed image and the second training image, the parameters of the second basic decoding model are tuned to obtain a second decoding model that meets the second objective; wherein, the first objective for training the first basic decoding model is different from the second objective for training the second basic decoding model. Display the target image.
8. A method for training a decoding model, characterized in that, The method includes: Obtain the target training set; The first encoding model is used to encode the second training image in the target training set to obtain a second bitstream file. The first encoding model and the first decoding model are trained in the following way: the first bitstream file output by the basic encoding model after encoding the first training image is used as the input of the first basic decoding model; the first basic decoding model is used to decode the first bitstream file to obtain a first reconstructed image; based on the difference between the first reconstructed image and the first training image, the basic encoding model and the first basic decoding model are tuned to obtain the first encoding model and the first decoding model that conforms to the first target. The second bitstream file is input into the second basic decoding model to obtain the second reconstructed image; Based on the difference between the second reconstructed image and the second training image, the parameters of the second basic decoding model are tuned to obtain a second decoding model that meets the second objective. The first objective of training the first basic decoding model is different from the second objective of training the second basic decoding model.
9. The method according to claim 8, characterized in that, When the first target attribute of the decoded image generated by the second decoding model meets the image expectation requirement (where the first target attribute refers to any image attribute that affects the visual effect of the image), the step of tuning the parameters of the second basic decoding model based on the difference between the second reconstructed image and the second training image to obtain a second decoding model that meets the second target includes: The loss is calculated based on the difference between the second reconstructed image and the second training image, and the difference between the first target attribute of the second reconstructed image and the image expectation requirement. Based on the loss, the parameters of the second basic decoding model are tuned to obtain the second decoding model that meets the second objective.
10. The method according to claim 8, characterized in that, When the second target attribute of the region where the target object is located in the decoded image generated by the second decoding model meets the local expectation requirement, and the second target attribute refers to any image attribute that affects the visual effect of the image, the step of tuning the parameters of the second basic decoding model based on the difference between the second reconstructed image and the second training image to obtain a second decoding model that meets the second target includes: Based on the difference between the second reconstructed image and the second training image, and the difference between the second target attribute of the region where the target object is located in the second reconstructed image and the local expectation requirement, the loss is calculated; Based on the loss, the parameters of the second basic decoding model are tuned to obtain the second decoding model that meets the second objective.
11. An image decoding device, characterized in that, The device includes: The bitstream acquisition module is used to acquire a target bitstream file, which is obtained by encoding the image to be compressed using a first encoding model. The first encoding model and the first decoding model are trained in the following way: the first bitstream file output by the basic encoding model after encoding the first training image is used as the input of the first basic decoding model; the first basic decoding model is used to decode the first bitstream file to obtain a first reconstructed image; based on the difference between the first reconstructed image and the first training image, the basic encoding model and the first basic decoding model are tuned to obtain the first encoding model and the first decoding model that meets the first target. The bitstream decoding module is used to decode the target bitstream file using a second decoding model to obtain a target image. The second decoding model is trained in the following way: the second bitstream file output by the first encoding model after encoding the second training image is used as the input of the second basic decoding model; the second basic decoding model is used to decode the second bitstream file to obtain a second reconstructed image; based on the difference between the second reconstructed image and the second training image, the parameters of the second basic decoding model are tuned to obtain a second decoding model that meets the second target. The first objective of training the first basic decoding model is different from the second objective of training the second basic decoding model.
12. A video frame display device, characterized in that, The device includes: The target bitstream acquisition module is used to acquire the target bitstream file formed after the video frames are encoded. The target bitstream file is obtained by encoding the video frames using a first encoding model. The first encoding model and the first decoding model are trained in the following way: the first bitstream file output by the basic encoding model after encoding the first training image is used as the input of the first basic decoding model; the first basic decoding model is used to decode the first bitstream file to obtain the first reconstructed image; based on the difference between the first reconstructed image and the first training image, the basic encoding model and the first basic decoding model are tuned to obtain the first encoding model and the first decoding model that meets the first target. The target image generation module is used to decode the target bitstream file using a second decoding model to obtain a target image. The second decoding model is trained in the following way: the second bitstream file output by the first encoding model after encoding the second training image is used as the input of the second basic decoding model; the second basic decoding model is used to decode the second bitstream file to obtain a second reconstructed image; based on the difference between the second reconstructed image and the second training image, the parameters of the second basic decoding model are tuned to obtain a second decoding model that conforms to the second objective; wherein, the first objective for training the first basic decoding model is different from the second objective for training the second basic decoding model. The target image display module is used to display the target image.
13. A decoding model training device, characterized in that, The device includes: The training set acquisition module is used to acquire the target training set; A bitstream generation module is used to encode a second training image in the target training set using a first encoding model to obtain a second bitstream file. The first encoding model and the first decoding model are trained in the following way: the first bitstream file output by the basic encoding model after encoding the first training image is used as the input of the first basic decoding model; the first basic decoding model is used to decode the first bitstream file to obtain a first reconstructed image; based on the difference between the first reconstructed image and the first training image, the basic encoding model and the first basic decoding model are tuned to obtain the first encoding model and the first decoding model that conforms to the first target. The image reconstruction module is used to input the second bitstream file into the second basic decoding model to obtain the second reconstructed image; The model generation module is used to tune the parameters of the second basic decoding model based on the difference between the second reconstructed image and the second training image to obtain a second decoding model that meets the second objective. The first objective of training the first basic decoding model is different from the second objective of training the second basic decoding model.
14. A computer device, characterized in that, include: A memory and a processor, wherein the memory stores program instructions; when the program instructions are executed by the processor, the processor performs the method as described in any one of claims 1 to 10.
15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program; when the computer program is run on one or more processors, it performs the method as described in any one of claims 1 to 10.
16. A computer program product, characterized in that, The computer program product includes a computer program or instructions; when the computer program or instructions are executed on a computer, the computer causes the computer to perform the method as described in any one of claims 1 to 10.
Citation Information
Patent Citations
Systems, methods, and bitstream structures for hybrid feature video bitstreams and decoders
CN117356092A
Processing method and device and storage medium
CN117546159A