Image decoding method, image encoding and decoding device and computer storage medium

By using an entropy model structure based on a spatial context model and combining spatial and spatiotemporal information for probability prediction, the problems of high computational complexity and insufficient information utilization in traditional image encoding and decoding technologies are solved, achieving efficient image decoding results.

CN121531133APending Publication Date: 2026-02-13ZHEJIANG DAHUA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411069624.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-08-05
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

Traditional image encoding and decoding technologies have high computational complexity, cannot effectively utilize spatial information, and are difficult to meet the subjective quality requirements of machine vision tasks. Furthermore, traditional video encoding technologies are unable to meet the needs of machine vision by optimizing objective indicators.

Method used

An entropy model structure based on a spatial context model is adopted. By using the spatial context model and a fusion network, the probability parameters of the undecoded bitstream are obtained. Decoding is performed using spatially decoded information and prior decoding information. Probability prediction is performed by combining spatiotemporal context information, thereby improving decoding efficiency and accuracy.

Benefits of technology

It improves the efficiency and effectiveness of image decoding. By making full use of spatial domain dependency information, it enhances the accuracy of probability prediction, reduces the coding bitrate, and improves the coding compression rate and decoding efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121531133A_ABST
    Figure CN121531133A_ABST
Patent Text Reader

Abstract

The invention discloses an image decoding method, an image encoding and decoding device and a computer storage medium. The image decoding method comprises the following steps: acquiring a code stream to be decoded; based on the to-be-decoded code stream, obtaining the spatial domain decoded information of the current frame corresponding to the to-be-decoded code stream; inputting the spatial domain decoded information into a spatial domain context model, and obtaining context information transmitted by the spatial domain decoded information; inputting the context information and the super-prior decoding information of the code streams to be decoded into a fusion network, and obtaining probability parameters of undecoded code streams in the code streams to be decoded; and decoding an undecoded code stream in the code stream to be decoded by using the probability parameter to obtain a decoding feature of the image. According to the invention, the entropy model structure based on the spatial domain context model is provided, the dependence information on the spatial domain is fully utilized, the accuracy of probability prediction is improved, and the image decoding effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image encoding and decoding technology, and in particular to an image decoding method, an image encoding and decoding device, and a computer storage medium. Background Technology

[0002] Traditional image encoding and decoding techniques are designed for human visual characteristics. However, with the superior performance of deep neural networks in various machine vision tasks, such as image classification, object detection, and semantic segmentation, a large number of AI applications based on machine vision have emerged. To ensure that the performance of machine vision tasks is not compromised by the image encoding process, an analysis-then-encode approach is adopted to meet the demands of machine vision. This involves directly extracting features from the lossless image using a neural network at the image acquisition end, then encoding and transmitting the extracted features. The decoding end directly uses the decoded features as input to subsequent network structures to complete different machine vision tasks. Therefore, to save transmission bandwidth resources, it is necessary to research image encoding methods specifically for machine vision.

[0003] However, traditional video encoding and decoding technologies are becoming increasingly computationally complex, making it difficult to make good use of decoding information in the spatial domain. Furthermore, traditional video coding technologies are generally optimized for objective indicators such as PSNR (Peak Signal-to-Noise Ratio), making it difficult to meet the requirements of subjective quality. Summary of the Invention

[0004] This application provides an image decoding method, an image encoding / decoding device, and a computer storage medium.

[0005] One technical solution adopted in this application is to provide an image decoding method, the image decoding method comprising:

[0006] Obtain the bitstream to be decoded;

[0007] Based on the code stream to be decoded, obtain the spatial domain decoded information of the current frame corresponding to the code stream to be decoded;

[0008] Input the decoded spatial information into the spatial context model to obtain the context information transmitted by the decoded spatial information;

[0009] The context information and the prior decoding information of the bitstream to be decoded are input into the fusion network to obtain the probability parameters of the undecoded bitstream in the bitstream to be decoded;

[0010] The undecoded bitstream in the bitstream to be decoded is decoded using the probability parameters to obtain the decoding features of the image.

[0011] The fusion network includes a first fusion network and a second fusion network;

[0012] The step of inputting the context information and the priori decoding information of the bitstream to be decoded into the fusion network to obtain the probability parameters of the undecoded bitstream in the bitstream to be decoded includes:

[0013] The context information and the prior decoding information of the bitstream to be decoded are input into the first fusion network to obtain the mean value of the probability parameters of the undecoded bitstream in the bitstream to be decoded;

[0014] The prior decoding information of the bitstream to be decoded is input into the second fusion network to obtain the variance of the probability parameters of the undecoded bitstream in the bitstream to be decoded.

[0015] The image decoding method further includes:

[0016] The current frame is divided into blocks according to a preset direction to obtain several feature blocks of the current frame;

[0017] The step of obtaining the spatial domain decoded information of the current frame corresponding to the bitstream to be decoded, based on the bitstream to be decoded, includes:

[0018] Based on the bitstream of the previous feature block of the current feature block in the bitstream to be decoded, obtain the decoded information of the previous feature block of the current feature block.

[0019] The preset direction includes the channel direction and / or the spatial direction.

[0020] Wherein, obtaining the spatial domain decoded information of the current frame corresponding to the bitstream to be decoded based on the bitstream to be decoded includes:

[0021] Based on the bitstreams of the neighboring feature blocks of the current feature block in the bitstream to be decoded, the spatial domain decoded information of the neighboring feature blocks of the current feature block is obtained. The neighboring feature blocks include the decoded feature blocks that are adjacent to the current feature block in the spatial domain on the top, left, right, bottom and / or diagonal lines, respectively. The decoded feature blocks have some channel feature blocks that have been decoded.

[0022] The step of inputting the decoded spatial information into the spatial context model to obtain the context information transmitted by the decoded spatial information includes:

[0023] Obtain the undecoded and decoded channels of the current feature block;

[0024] In the decoded channel of the current feature block, the first convolutional kernel of the spatial context model is selected to obtain the first decoded feature block in the spatial decoded information with the current feature block as the center;

[0025] In the undecoded channel of the current feature block, the second convolution kernel of the spatial context model is selected to obtain the second decoded feature block in the spatial decoded information with the current feature block as the center;

[0026] The first context information of the first decoded feature block is extracted from the decoded channel of the current feature block, and the second context information of the second decoded feature block is extracted from the undecoded channel of the current feature block.

[0027] The context information for transmitting the spatially decoded information is obtained by combining the first context information and the second context information according to the channel combination;

[0028] Wherein, the first decoded feature block includes the decoded feature blocks that are adjacent to the current feature block in the spatial domain above, to the left, to the right, and below, respectively; the first decoded feature block includes the decoded feature blocks that are adjacent to the current feature block on the diagonal in the spatial domain.

[0029] The image decoding method further includes:

[0030] Based on the bitstream to be decoded, obtain the previously decoded frames preceding the bitstream to be decoded;

[0031] The step of inputting the context information and the priori decoding information of the bitstream to be decoded into the fusion network to obtain the probability parameters of the undecoded bitstream in the bitstream to be decoded includes:

[0032] The cached information of the decoded frames, the context information, and the priori decoding information are input into the fusion network to obtain the probability parameters of the undecoded bitstream in the bitstream to be decoded;

[0033] The cached information refers to the decoding information of the decoded frame and / or intermediate variables generated during the decoding process.

[0034] The step of inputting the spatially decoded information into the spatial context model to obtain the decoding information of the bitstream to be decoded includes:

[0035] Input the decoded spatial information and the cached information into the spatial context model to obtain the context information transmitted by the decoded spatial information.

[0036] The cached information includes long-term information and short-term information;

[0037] The cached information is extracted through a recurrent neural network in the entropy model. The recurrent neural network is set in any one of the networks in the entropy model, or it is set before any one of the networks in the entropy model, or it is set after any one of the networks except the fusion network.

[0038] Another technical solution adopted in this application is to provide another image decoding method, which includes:

[0039] Obtain the bitstream to be decoded;

[0040] Based on the bitstream to be decoded, obtain the previously decoded frames preceding the bitstream to be decoded;

[0041] The buffer information of the decoded frame and the prior decoding information of the bitstream to be decoded are input into the fusion network to obtain the probability parameters of the undecoded bitstream in the bitstream to be decoded.

[0042] The undecoded bitstream in the bitstream to be decoded is decoded using the probability parameters to obtain the decoding features of the image;

[0043] The cached information refers to the decoding information of the decoded frame and / or intermediate variables generated during the decoding process.

[0044] The image decoding method further includes:

[0045] Based on the code stream to be decoded, obtain the spatial domain decoded information of the current frame corresponding to the code stream to be decoded;

[0046] Input the decoded spatial information into the spatial context model to obtain the context information transmitted by the decoded spatial information;

[0047] The step of inputting the buffer information of the decoded frames and the prior decoding information of the bitstream to be decoded into the fusion network to obtain the probability parameters of the undecoded bitstream in the bitstream to be decoded includes:

[0048] The buffer information of the decoded frames, the context information, and the prior decoding information are input into the fusion network to obtain the probability parameters of the undecoded bitstream in the bitstream to be decoded.

[0049] Another technical solution adopted in this application is to provide an image decoding device, which includes a memory and a processor coupled to the memory;

[0050] The memory is used to store program data, and the processor is used to execute the program data to implement the image decoding method described above.

[0051] Another technical solution adopted in this application is to provide a computer storage medium for storing program data, which, when executed by a computer, is used to implement the image decoding method described above.

[0052] The beneficial effects of this application are as follows: the image decoding device acquires the bitstream to be decoded; based on the bitstream to be decoded, it acquires the spatial domain decoded information of the current frame corresponding to the bitstream to be decoded; it inputs the spatial domain decoded information into a spatial domain context model to acquire the context information transmitted by the spatial domain decoded information; it inputs the context information and the prior decoding information of the bitstream to be decoded into a fusion network to acquire the probability parameters of the undecoded bitstream in the bitstream to be decoded; and it uses the probability parameters to decode the undecoded bitstream in the bitstream to be decoded to obtain the decoding features of the image. This application proposes an entropy model structure based on a spatial domain context model, which fully utilizes the dependency information in the spatial domain to improve the accuracy of probability prediction and enhance the image decoding effect. Attached Figure Description

[0053] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0054] Figure 1 This is a schematic diagram of the overall framework for end-to-end video encoding and decoding provided in this application;

[0055] Figure 2 This is a schematic diagram of the basic structure of the entropy model provided in this application;

[0056] Figure 3 This is a schematic diagram of the overall framework of the entropy model based on the spatiotemporal context model provided in this application;

[0057] Figure 4 This is a flowchart illustrating an embodiment of the image decoding method provided in this application;

[0058] Figure 5 This is a schematic diagram of the entropy model that incorporates the spatial context model provided in this application;

[0059] Figure 6 This is a schematic diagram of the decoupling entropy model that incorporates a spatial context model, as provided in this application.

[0060] Figure 7 This is a schematic diagram of the video frame segmentation results in the channel and spatial domain provided in this application;

[0061] Figure 8This is a flowchart illustrating the video frame channel segmentation process provided in this application;

[0062] Figure 9 This is a flowchart illustrating the video frame spatial segmentation process provided in this application;

[0063] Figure 10 This is a schematic diagram of the entropy decoding process that incorporates a context model, as provided in this application.

[0064] Figure 11 This is a schematic diagram of the decoding process of the channel-complementary chessboard context model provided in this application;

[0065] Figure 12 This is a schematic diagram of the decoding process of the spatial context model (based on blocks) provided in this application;

[0066] Figure 13 This is a flowchart illustrating another embodiment of the image decoding method provided in this application;

[0067] Figure 14 This is a schematic diagram of the structure of short-term features used in the entropy model provided in this application;

[0068] Figure 15 This is a schematic diagram of the structure of the long-term and short-term feature dependency module provided in this application;

[0069] Figure 16 This is a schematic diagram of the decoding process for short-term information dependence provided in this application;

[0070] Figure 17 This is a schematic diagram of the long-term and short-term feature dependency modules for long-term information dependency provided in this application;

[0071] Figure 18 This is a schematic diagram of the long-term and short-term feature dependency modules provided in this application, which consist of long-term information dependency and short-term information dependency.

[0072] Figure 19 This is a schematic diagram of the entropy model that incorporates a spatiotemporal context model, as provided in this application.

[0073] Figure 20 This application provides a schematic diagram of a spatiotemporal entropy model that incorporates a spatiotemporal context model.

[0074] Figure 21 This is a schematic diagram of an embodiment of the image decoding device provided in this application;

[0075] Figure 22 This is a schematic diagram of the structure of an embodiment of the computer storage medium provided in this application. Detailed Implementation

[0076] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0077] In recent years, deep neural networks have made some progress in video processing, such as video detection, video super-resolution, video denoising, and video enhancement. Due to their powerful non-linear expressive capabilities and the advantages of joint training, deep neural networks have demonstrated great potential in the image / video field.

[0078] Currently, deep learning is also beginning to develop in the field of video compression. Its applications are mainly divided into two categories. The first is as a deep learning tool applied to traditional video encoders. To date, many works have proven that combining traditional coding modules with deep learning is very effective. These modules include, but are not limited to, motion compensation and frame interpolation networks, intra-frame predictive coding modules, bit rate control modules, and post-processing modules. The second is an end-to-end deep video compression framework with deep neural networks as the core of video coding. This framework makes full use of the powerful nonlinear expression capabilities of neural networks and the advantages of end-to-end joint optimization, which can further improve compression efficiency and accuracy.

[0079] This application proposes a video end-to-end encoding and decoding method based on a spatiotemporal domain context entropy model, mainly considering the following points:

[0080] Spatial Context Model. A decoupled entropy model based on the spatial context model is designed. This means that all bitstreams are decoded simultaneously during entropy decoding, and then feature information is recovered sequentially based on the spatial context model. Simultaneously, the spatial context model considers all information within the receptive field, improving encoding / decoding performance while ensuring decoding efficiency.

[0081] Temporal Context Model. A temporal context model is designed, combining long-term and short-term information dependencies over time to fully predict probability parameters and improve encoding / decoding performance.

[0082] This application proposes an end-to-end video encoding and decoding method based on a spatiotemporal context entropy model. It mainly improves the accuracy of probability prediction by making full use of spatiotemporal context information in the entropy model, reduces the codeword length of the encoding, improves the encoding compression rate, and enhances the decoding efficiency and effect.

[0083] Please refer to details. Figure 1 , Figure 1 This is a schematic diagram of the overall framework for end-to-end video encoding and decoding provided in this application.

[0084] like Figure 1As shown, the overall framework for end-to-end video encoding and decoding in this application includes, but is not limited to, the following parts:

[0085] 1. Motion information compression and reconstruction:

[0086] (1) Motion estimation: The current frame and the reference frame obtain inter-frame motion information through the motion estimation network.

[0087] (2) Motion Information Encoder: Reduces the dimensionality of motion information (or motion information residual) and extracts compact information to obtain the motion information to be encoded.

[0088] (3) Motion Vector Entropy Model: Obtain the probability of each character appearing in the quantized motion information to be encoded, perform arithmetic encoding, and output the motion information bitstream.

[0089] (4) Motion information Decoder: Upgrades the low-dimensional motion features of the decoded data and reconstructs the motion information.

[0090] 2. Motion compensation and temporal prediction: Motion compensation operations such as warping are performed on the features of the reference frame using motion information, and relevant information is extracted from the reference frame to obtain the predicted frame of the current frame.

[0091] 3. Context compression and reconstruction:

[0092] (1) Context Encoder: Performs dimensionality reduction and compression on the context information between the current frame and the predicted frame. The context information includes, but is not limited to, the difference information or concatenation information between the two frames. The purpose is to remove temporal and intra-spatial correlations between frames.

[0093] (2) Context entropy model: Obtain the probability of each character appearing in the quantized context information to be encoded, perform arithmetic encoding, and output the context information bitstream.

[0094] (3) Context Decoder: The low-dimensional context features obtained from decoding the above bitstream are upscaled and reconstructed to obtain context information. In this process, the context information is combined with the predicted frame to obtain preliminary reconstruction information. For example, if the context at the Encoder end is a difference, the Decoder end will add the predicted frame information to obtain the preliminary reconstruction information of the current frame.

[0095] 4. Frame Reconstruction: Further processing of the preliminary reconstruction information yields the reconstructed image and reconstruction features of the current frame. Both are stored in the cache information and can be used as a reference frame for the next frame.

[0096] Figure 1 Please refer to the basic structure of the entropy model shown below. Figure 2 , Figure 2This is a schematic diagram of the basic structure of the entropy model provided in this application. For example... Figure 2 As shown, the basic structure of the entropy model sequentially performs super-prior encoding and super-prior decoding on the feature to be encoded, y, thereby inputting the super-prior decoded information into the fusion network, which then outputs the probability parameters required for encoding and decoding. Among these, Figure 2 The dashed branches shown represent optional branches, such as... Figure 1 The context entropy model shown does not have quantization parameter branches and dequantization parameter branches. Figure 1 The motion information entropy model shown has a quantization parameter branch and an inverse quantization parameter branch.

[0097] This application proposes to Figure 2 The improvement in the basic structure of the entropy model shown mainly lies in fully utilizing spatiotemporal information to enhance the accuracy of probabilistic predictions. Please refer to [link / reference needed] for details. Figure 3 , Figure 3 This is a schematic diagram of the overall framework of the entropy model based on the spatiotemporal context model provided in this application.

[0098] like Figure 3 As shown, this application is in Figure 2 The basic structure of the entropy model shown introduces a context model (spatial domain, spatiotemporal domain) and cached information (temporal domain), including but not limited to the following design considerations:

[0099] (1) Design an entropy model based on spatial context. Combine the spatially decoded information of the current frame with prior information to calculate the information needed for decoding the features to be decoded. In the context model, the encoding / decoding of features has a certain order, and there are multiple ways to partition the information, including but not limited to the following methods:

[0100] (A) Block-based context model.

[0101] (B) A chessboard context model with complementary channels.

[0102] Based on the spatial domain context model scheme, this application further proposes a decoupled entropy model to improve decoding efficiency. That is, the branch that generates the variance required for entropy decoding is independent of the branch that introduces spatial domain information (decoupling).

[0103] (a) The variance can be obtained at once so that entropy decoding can decode all the bitstreams at once.

[0104] (b) The context model branch is only used to calculate the mean in the probability parameters, and can be used to process the entropy decoding results to obtain the decoded features.

[0105] (2) Temporal Context Entropy Model. This model primarily combines information from already decoded frames (buffered information) and prior information to infer the entropy decoding information required for decoding the current frame. It includes, but is not limited to, the following methods:

[0106] (A) Short message dependency. Short messages include, but are not limited to, decoding information of decoded frames and buffered information generated during the decoding process.

[0107] (B) Long-term information dependency. Long-term dependency information between frames is captured using recurrent neural networks (RNN, LSTM, etc.) and passed to the current frame.

[0108] (3) Spatiotemporal context entropy model. When the entropy model simultaneously introduces the decoded information in the spatial domain and the decoded information in the temporal domain (buffered information) of the current frame, that is, the combination of the first scheme (1) and the second scheme (2), it is the spatiotemporal context entropy model.

[0109] For details regarding the entropy model of spatial context described above, please refer to [link / reference]. Figure 4 and Figure 5 , Figure 4 This is a flowchart illustrating an embodiment of the image decoding method provided in this application. Figure 5 This is a schematic diagram of the entropy model that incorporates the spatial context model provided in this application.

[0110] In existing technologies, the entropy model infrastructure typically combines prior information and complete decoding information of decoded frames in the time domain to predict and generate probability parameters (mean, variance), and / or quantization factors required for decoding the frame to be decoded.

[0111] To improve the accuracy of probability prediction and further reduce the bitrate of motion information, this application proposes the following... Figure 5 The entropy model shown aims to fully utilize the decoded context information by introducing a spatial context model between the decoding feature branch and the fusion network. For the current frame, some features are decoded first, and then the probability parameters of the next part of the features are generated based on the decoded features. The next part of the features is then solved, and so on.

[0112] like Figure 4 As shown, the image decoding method of this application embodiment includes the following steps:

[0113] Step S11: Obtain the bitstream to be decoded.

[0114] In this embodiment of the application, the bitstream to be decoded is the decoded bitstream of the current feature block among several feature blocks divided in the current frame, and the feature blocks in the current frame before the current feature block are the decoded feature blocks.

[0115] Step S12: Based on the bitstream to be decoded, obtain the spatial domain decoded information of the current frame corresponding to the bitstream to be decoded.

[0116] In this embodiment, the image encoding / decoding device determines the current feature block of the current frame and the decoded feature blocks of the current frame based on the bitstream to be decoded. The image encoding / decoding device uses the decoding information of the decoded feature blocks of the current frame as the spatial domain decoded information.

[0117] Step S13: Input the decoded spatial information into the spatial context model to obtain the context information passed by the decoded spatial information.

[0118] In this embodiment of the application, the image encoding and decoding device inputs the spatial domain decoded information. Figure 5 The spatial context model shown obtains the context information for the transmission of spatially decoded information. By introducing a spatial context model, this application fully utilizes the spatially decoded information, enabling the context information of the decoded frame to be applied to the decoding process of each feature block, thereby improving the accuracy of the probability parameters output by the fusion network.

[0119] Step S14: Input the context information and the prior decoding information of the bitstream to be decoded into the fusion network to obtain the probability parameters of the undecoded bitstream in the bitstream to be decoded.

[0120] In this embodiment of the application, the image encoding and decoding device simultaneously inputs the decoding information of the decoded feature blocks and the prior decoding information of the bitstream to be decoded into the fusion network to generate probability parameters of the bitstream to be decoded.

[0121] The process of generating the priori decoding information of the bitstream to be decoded is as follows:

[0122] The image encoding and decoding device performs a priori encoding on the features to be encoded corresponding to the bitstream to be decoded, then obtains the binary bitstream through a quantization and entropy encoding model, and finally inputs the binary bitstream into the entropy decoding model for decoding to obtain the a priori decoding information.

[0123] Furthermore, the fusion network of this application can also output quantization parameters for dequantizing the bitstream to be decoded and recovering the decoding features.

[0124] Step S15: Use probability parameters to decode the undecoded bitstream in the bitstream to be decoded to obtain the decoding features of the image.

[0125] In the embodiments of this application, the image encoding and decoding device uses probability parameters to predict the feature values ​​of each feature point of the bitstream to be decoded, thereby decoding the decoding features or the decoded image.

[0126] In this embodiment, an image encoding / decoding device acquires a bitstream to be decoded; based on the bitstream to be decoded, it acquires the spatially decoded information of the current frame corresponding to the bitstream to be decoded; it inputs the spatially decoded information into a spatial context model to acquire the context information transmitted by the spatially decoded information; it inputs the context information and the prior decoding information of the bitstream to be decoded into a fusion network to acquire the probability parameters of the undecoded bitstream in the bitstream to be decoded; and it uses the probability parameters to decode the undecoded bitstream in the bitstream to be decoded to obtain the decoding features of the image. This application proposes an entropy model structure based on a spatial context model, which fully utilizes the dependency information in the spatial domain to improve the accuracy of probability prediction and enhance the image decoding effect.

[0127] In one specific implementation, to accelerate decoding speed, this application further proposes a decoupled entropy model, specifically in... Figure 5 Please refer to the following: Figure 6 , Figure 6 This is a schematic diagram of the decoupling entropy model that incorporates a spatial context model, as provided in this application.

[0128] like Figure 6 As shown, Figure 5 The fusion network in the middle can be divided into the first fusion network, that is Figure 6 The fusion network and the second fusion network in the middle, namely Figure 6 The image encoding / decoding device inputs the decoding information output by the spatial context model and the prior decoding information into the first fusion network to obtain the mean of the probability parameters output by the first fusion network. The image encoding / decoding device inputs the prior decoding information into the second fusion network to obtain the variance of the probability parameters output by the second fusion network. Furthermore, the second fusion network can also output quantization parameters.

[0129] pass Figure 6 In the decoupled entropy model, the variance (or variance + quantization parameter) required by the image encoder / decoder during the entropy decoding process (using a probability model with a mean of 0) is directly generated by the super-prior branch, without referencing the spatial context model branch. In other words, the parameters related to entropy decoding need to be obtained all at once, and the spatial context model branch is only used to calculate the mean of the probability parameters.

[0130] In one specific implementation, Figure 5 The spatial context model shown can be a block-based spatial context model. Specifically, the block-based spatial context model divides the current frame into blocks according to a preset direction, obtaining several feature blocks of the current frame. The spatial decoded information of the current frame is the decoded information of the previous feature block of the current feature block to be decoded. The decoded information of the previous feature block is obtained based on the bitstream of the previous feature block of the current feature block in the bitstream to be decoded.

[0131] The preset direction segmentation of this application can be segmented in the channel dimension, the spatial dimension, or both the channel dimension and the spatial dimension simultaneously.

[0132] An image encoding / decoding device can divide the current frame into blocks in the channel dimension and / or spatial dimension, first decode the information of some feature blocks, and then sequentially decode the features of subsequent blocks based on the decoded feature blocks. Let the dimension of the feature to be encoded / decoded X be [H, W, C], where H, W, and C are the length, width, and number of channels, respectively.

[0133] Please refer to details. Figure 7 , Figure 7 This is a schematic diagram illustrating the video frame segmentation results in the channel and spatial domain provided in this application. For example... Figure 7 As shown, the video frame is divided into channel-oriented blocks and spatial-oriented blocks, and the size of each feature block can be different. To facilitate subsequent dimensionality processing and avoid information redundancy, the block size can be designed to be a multiple of 2, so that the desired feature dimensions can be generated in the context model using upsampling / downsampling methods. Figure 7 The specific process for block segmentation is as follows:

[0134] (1) Channel segmentation. Please refer to [link / reference]. Figure 8 , Figure 8 This is a schematic diagram of the video frame channel segmentation process provided in this application. The image encoding and decoding device divides channel C into m parts, and the features of each part are denoted as A1, A2, ..., A... m The corresponding channel numbers are C1, C2, ..., C m .

[0135] (2) Spatial partitioning. Please refer to [link / reference]. Figure 9 , Figure 9 This is a flowchart illustrating the spatial segmentation of video frames provided in this application. The image encoding / decoding device performs non-overlapping segmentation of features in the spatial dimension, with the number of segments denoted as N(1, 2, ..., H*W). Specifically, the above m features A1, A2, ..., A... m The number of blocks obtained by dividing the space dimension is N1, N2, ..., N. m 1 <= N1, N2, ..., N m <=H*W.

[0136] N(N1, N2, ..., N) m When ) = 1, it means that there is no block division in the spatial direction and the spatial dimension remains unchanged.

[0137] N(N1, N2, ..., N) m When ) = H*W, it means that the space is divided into H*W blocks, that is, each feature point is one block.

[0138] Please continue reading. Figure 10 , Figure 10 This is a schematic diagram of the entropy decoding process that incorporates a context model, as provided in this application. Figure 10 As shown, the image encoding / decoding device decodes the feature blocks obtained from the channel segmentation and / or spatial segmentation in sequence. The image encoding / decoding device first decodes a portion of the feature blocks, then combines the decoded feature blocks with the prior information of the current frame to generate the probability parameters or quantization parameters required for the next block to be decoded. Specifically, the entropy decoding process incorporating a context model is as follows:

[0139] (1) Context Module. Decoded block information is input into the context model to generate context information. The context model network includes, but is not limited to, residual networks, attention structures, and transformers.

[0140] (2) Fusion. Contextual information, together with prior decoding information, is input into the fusion network to obtain the probability parameters and quantization parameters required for decoding the current feature to be decoded.

[0141] (3) Entropy decoding (inverse quantization optional). The decoding information of the current feature block is obtained based on the probability parameters (quantization parameters).

[0142] The decoded information is then used as reference information for the next block to be decoded. The above operations (1), (2) and (3) are repeated until all feature blocks are decoded, and the complete decoded information of the current frame is obtained.

[0143] Generally, the closer the blocks are, the stronger their correlation. Therefore, when decoding, image encoding and decoding devices can prioritize features that are spatially or channel-level adjacent to each other to infer the probability parameters (and / or quantization parameters) of the current block.

[0144] In one specific implementation, Figure 5 The spatial context model shown can be a channel-complementary chessboard context model. Specifically, the channel-complementary chessboard context model obtains the spatially decoded information of the neighboring feature blocks of the current feature block based on the bitstream of the current feature block in the bitstream to be decoded. The neighboring feature blocks include the decoded feature blocks that are spatially adjacent to the current feature block above, to the left, to the right, below, and / or diagonally. The decoded feature blocks contain some channel feature blocks that have been decoded.

[0145] The image encoding / decoding device acquires the undecoded channel and the decoded channel of the current feature block; in the decoded channel of the current feature block, a first convolutional kernel of the spatial context model is selected to acquire the first decoded feature block in the spatial decoded information centered on the current feature block; in the undecoded channel of the current feature block, a second convolutional kernel of the spatial context model is selected to acquire the second decoded feature block in the spatial decoded information centered on the current feature block; first decoding information of the first decoded feature block is extracted from the decoded channel of the current feature block, and second decoding information of the second decoded feature block is extracted from the undecoded channel of the current feature block; the first decoding information and the second decoding information are combined according to the channel to obtain the decoding information of the bitstream to be decoded.

[0146] Wherein, the first decoded feature block includes the decoded feature blocks that are adjacent to the current feature block in the spatial domain above, to the left, to the right, and below, respectively; the first decoded feature block includes the decoded feature blocks that are adjacent to the current feature block on the diagonal in the spatial domain.

[0147] To ensure that each point to be decoded can utilize information from its surrounding receptive field, the image encoding / decoding device can first decode some points with intervals, and then decode the remaining points in the next stage. Methods include, but are not limited to, the following:

[0148] (1) Dual-channel chessboard context model. In the first stage, the first half of the channel solves the points at the same spatial location (which can be denoted as odd-numbered points, and the decoding point starts from the first point). In the second stage, the subsequent points are solved based on the decoded points (denoted as even-numbered points, and the decoding point starts from the second point). Here, in the context model, the activation position of the convolution (the effective weight is 1) is consistent with the position of the decoded point.

[0149] (2) Complementary Channel Chessboard Context Model. In the first stage, when decoding the features of the first channel, all points within the receptive field are solved according to the channel (i.e., odd and even points are solved). In this way, during the second stage of decoding, the information in the receptive field around the point to be decoded can be used to calculate the probability parameters of the point to be decoded.

[0150] The following describes the decoding process of the channel-complementary chessboard context model. Please refer to [link / reference needed] for details. Figure 11 , Figure 11 This is a schematic diagram of the decoding process of the channel-complementary chessboard context model provided in this application.

[0151] like Figure 11 As shown, suppose the feature X to be decoded has C channels. The decoding process of the complementary channel chessboard context model is as follows:

[0152] (1) First stage. Initialize a feature with all zeros as context information, and input it along with the super-prior decoding information into fusion network 1 to obtain probability parameters, and further decode odd point 1 and even point 1 (darker areas represent decoded points). That is, C1 channels decode odd point 1 and C2 channels decode even point 2. Here,

[0153] C = C1 + C2 (C1 and C2 can be different)

[0154] Fusion Network 1 includes, but is not limited to, residual blocks, convolutional neural networks, etc.

[0155] It should be noted that this application employs multiple methods for dividing the information into C1 and C2 channels, including but not limited to the following methods:

[0156] Method 1: Odd-Even Interleaving Method. Solve for odd-numbered points 1 using the odd-numbered channels, and solve for even-numbered points 1 using the even-numbered channels.

[0157] Method 2: Sequential method. Solve odd-numbered points (and even-numbered points) using the first C1 channels, and then solve even-numbered points (and odd-numbered points) using the last C2 channels.

[0158] (2) In the second stage, the remaining even and odd points are solved based on the above odd and even points.

[0159] (a) Context Model. Decoded points are input into the context model to predict contextual information. The context model mainly consists of convolutions. The number of input channels of the convolution kernels is consistent with the number of feature channels to be decoded, and they maintain a one-to-one correspondence. Furthermore, the effective weights (darker dots) of each convolution kernel channel are consistent with the positions of the decoded points in the first stage. Figure 11 Taking the center point as an example,

[0160] (i) For odd-numbered points in channel 1 (C1), the effective weights of the corresponding convolutional kernel channel are the four points above, below, left, and right of the center point, and the weights of the remaining points are 0, because the other points of this branch have not been decoded.

[0161] (ii) For even-numbered points in channel 1 (C2), the effective weights of the corresponding convolutional kernel channel are the four points diagonally opposite to the center point and the center point, and the weights of the remaining points are 0, because the other points of this branch have not been decoded.

[0162] Note: The convolution kernel is calculated as a weighted sum of the regions covered by all convolution kernel channels in the input features. Therefore, the location points of the effective weights of these channels can constitute a complete receptive field. Thus, in the second stage, the information around each point to be decoded is included as much as possible.

[0163] (b) Fusion Network 2. The context information obtained in step (a) is combined with the prior decoding information and input into fusion network 2 to obtain the probability parameters (quantization parameters) of the remaining even and odd points. The even point information of C1 channels and the odd point information of C2 channels are further decoded.

[0164] (3) Merging. The odd-numbered points 1 of channel C1 and even-numbered points 1 of channel C2 decoded in the first stage are merged with the even-numbered points 2 of channel C1 and odd-numbered points 2 of channel C2 decoded in the second stage to obtain the decoding information of all position points of channel C.

[0165] In implementation example 1, with Figure 5 Taking the entropy model shown as an example, the image encoding / decoding device selects a block-based context model. Please refer to [link / reference] for details. Figure 12 , Figure 12 This is a schematic diagram of the decoding process of the spatial context model (based on blocks) provided in this application.

[0166] like Figure 12 As shown, let the feature dimensions to be encoded / decoded be [H, W, C]. Taking decoding as an example, the encoding process is similar.

[0167] Channels are divided into blocks. The C channels are not evenly divided into 3 blocks, m = 2.

[0168] Spatial domain is divided into blocks. Without blocks, N=1. Therefore, the feature is divided into 3 parts: A1, A2, and A3.

[0169] Initialize the decoded feature block to 0, and decode the first feature block using context model 1 (which has all 0s as input, so it can be deleted) and combined with the prior information.

[0170] The first feature is used to obtain contextual information through Context Model 2 (Residual Block) and combined with prior information to decode the second feature.

[0171] The second feature is obtained by acquiring contextual information through context model 3 (residual block) and combined with prior information to decode the third feature.

[0172] By concatenating the three decoded feature channels, all decoded features are obtained.

[0173] In implementation example 2, with Figure 5 Taking the entropy model shown as an example, the image encoding and decoding device selects a block-based context model. Let the dimension of the feature to be encoded / decoded be [H, W, C]. Taking decoding as an example, the encoding process is similar.

[0174] Channels are segmented. Channel C is not segmented.

[0175] The spatial domain is divided into 4 equal blocks, N=4.

[0176] The decoding process is the same as in Implementation Example 1 above. First, the first feature block is decoded. Then, based on the first feature block, the parameters required for the entropy decoding of the second feature block are generated, and the second feature block is decoded. This process is repeated until four feature blocks are decoded. The blocks are then spatially concatenated to obtain a complete decoded feature map.

[0177] Here, since the closer the blocks are, the greater their correlation, the reference blocks are adjacent during the decoding process.

[0178] In implementation example 3, the image encoding / decoding device is selected. Figure 6 The decoupled entropy model, where the context model branch is used only to generate the mean, and the super-prior branch is used only to generate the variance (or variance + quantization parameters) related to entropy decoding. This means the variance can be obtained all at once, allowing entropy decoding to decode all bitstreams at once without waiting for all previous bitstreams to be decoded, resulting in faster speed.

[0179] The image encoding / decoding device selects a chessboard-like, context-dependent model with complementary channels. Let the dimensions of the features to be encoded / decoded be [H, W, C], where C = 64. Taking decoding as an example, the encoding process is similar.

[0180] Phase 1:

[0181] Information partitioning. The odd-even crossover method is chosen. Odd channels solve for odd-numbered points 1, and even channels solve for even-numbered points 1. Therefore, there are C1 = 32 channels solving for odd-numbered points 1 and C2 = 32 channels solving for even-numbered points 1.

[0182] Initialize all-zero features and input them along with prior information into a fusion network (convolutional network) to obtain parameters and decode them. Odd channels decode information for odd-numbered points (1), and even channels decode information for even-numbered points (1).

[0183] Phase Two:

[0184] All the above channel decoding points are input into the context model (single convolution) to obtain context information. The position of the effective weight of each channel in the convolution kernel is consistent with the effective position of the input information, that is:

[0185] The odd-numbered channels of the convolution kernel have only odd-numbered points with a weight of 1, while the rest have a weight of 0 (which has no effect).

[0186] The even-numbered channels of the convolution kernel have a weight of 1 only at even-numbered points, while the rest have a weight of 0 (which has no effect).

[0187] Contextual information and prior information are input into fusion network 2 (convolutional network) to obtain parameters and decode them. The even-numbered points 2 of odd-numbered channels and the odd-numbered points 2 of even-numbered channels are then solved.

[0188] The decoding features of the odd / even channels in the first and second stages are added together, and then the channels are concatenated to obtain the information of all channels.

[0189] Regarding the temporal context entropy model described above, in order to accurately predict the probability of encoded features and reduce codeword length, this application proposes to introduce decoded information (buffered information) from previous frames into the entropy model, including short-term and long-term information. The buffered information includes, but is not limited to, the decoding information of decoded frames and intermediate variables generated during the decoding process, such as reconstructed images, reconstructed features, decoded features, and other feature information.

[0190] Please refer to details. Figure 13 , Figure 13 This is a flowchart illustrating another embodiment of the image decoding method provided in this application.

[0191] like Figure 13 As shown, the image decoding method of this application embodiment includes the following steps:

[0192] Step S21: Obtain the bitstream to be decoded.

[0193] Step S22: Based on the bitstream to be decoded, obtain the previously decoded frames before the bitstream to be decoded.

[0194] Step S23: Input the buffer information of the decoded frame and the prior decoding information of the bitstream to be decoded into the fusion network to obtain the probability parameters of the undecoded bitstream in the bitstream to be decoded.

[0195] In the embodiments of this application, the short-term information in the cache information of the decoded frame includes, but is not limited to, the decoding information of the decoded frame and the intermediate variables generated during the decoding process (all of which can be used as cache information).

[0196] As the number of frames increases, the amount of information transmitted between frames decreases, and the amount of reliable, effective information also decreases. Therefore, the information directly transmitted can be called short-term information. The short-term characteristics of the cache can be passed to any module of the entropy model. Figure 2 Taking the entropy model shown as an example, the specific locations where the cache information can be embedded in this application are as follows: Figure 14 As shown, Figure 14 This is a schematic diagram of the structure using short-term features in the entropy model provided in this application. For example... Figure 14 As shown, the entropy model can pass cached short-term features to the super-prior encoder, fusion network, and / or super-prior decoder.

[0197] Furthermore, in order to mine long-term inter-frame information, this application can also introduce recurrent neural networks (RNN or LSTM, etc.) into the entropy model to make full use of time-domain information and improve the accuracy of probability prediction.

[0198] Please refer to details. Figure 15 , Figure 15 This is a schematic diagram of the structure of the long-term and short-term feature dependency module provided in this application. For example... Figure 15As shown, the long-short feature dependency module includes, but is not limited to, RNN, LSTM and ConvLSTM. Due to the gating mechanism, this module can pass the long-short information of the previous frame to the current frame.

[0199] It should be noted that, Figure 15 The long-term and short-term feature dependency modules shown can be set in any network of the entropy model, before any network of the entropy model, or after any network of the entropy model except for the fusion network.

[0200] In implementation example 4, the buffer information is assumed to be intermediate variables generated during the decoding of previous frames. Based on Figure 2 The entropy model shown in this application incorporates cached information into the fusion network; please refer to [link / reference] for details. Figure 16 , Figure 16 This is a schematic diagram of the decoding process for short-term information dependence provided in this application.

[0201] In implementation example 5, this application employs long-short-term information dependency. Figure 2 The entropy model shown introduces a long-short-term feature dependency module between the fusion network and the super-prior decoding, enabling the fusion network to utilize long-term information for probability / quantization parameter prediction. Please refer to [link to details]. Figure 17 , Figure 17 This is a schematic diagram of the long-term and short-term feature dependency modules for long-term information dependency provided in this application. For example... Figure 17 As shown, this application selects an LSTM structure, so it will pass two references: a long-term reference and a short-term reference.

[0202] In implementation example 6, this application combines short-term and long-term information dependencies in the cache. Figure 2 The entropy model shown in this application introduces a long-short-term feature dependency module to pass long-short-term information to the cached information (short-term information), enabling each network of the entropy model to use long-term information for probability / quantization parameter prediction. Please refer to [link to details]. Figure 18 , Figure 18 This is a schematic diagram of the long-term and short-term feature dependency modules provided in this application, consisting of long-term information dependency and short-term information dependency. For example... Figure 18 As shown, this application selects the ConvLSTM structure, so it will pass two references, a long-term reference and a short-term reference.

[0203] Step S24: Use probability parameters to decode the undecoded bitstream in the bitstream to be decoded to obtain the decoding features of the image.

[0204] For details regarding the entropy model of spatiotemporal context described above, please refer to [link / reference]. Figure 19 , Figure 19 This is a schematic diagram of the entropy model that incorporates a spatiotemporal context model, as provided in this application.

[0205] To improve the accuracy of entropy model prediction parameters, the entropy model incorporates both decoded information in the spatial domain and decoded information (buffered information) in the temporal domain of the current frame. This combines the first temporal context entropy model scheme with the second temporal entropy model scheme, resulting in a spatiotemporal context entropy model. Therefore, this application will... Figure 5 The spatial context model in the original text has been changed to a spatiotemporal context model.

[0206] In Implementation Example 7, this application adopts a spatiotemporal entropy model structure: spatial context model ( Figure 5 + Short-term information dependence. Please refer to the following for details. Figure 20 , Figure 20 This application provides a schematic diagram of a spatiotemporal entropy model that incorporates a spatiotemporal context model.

[0207] For the spatial context branch (decoded information branch), the method for selecting the chessboard context.

[0208] For the temporal information branch, cached information is introduced into the spatiotemporal context model and fusion network. The cached information is the cached features generated during the decoding of previous frames.

[0209] In the spatiotemporal context model, the information from two branches can be fused using methods such as concatenation, addition, or concatenation plus convolution.

[0210] This application proposes an entropy model structure based on a spatiotemporal context model, which makes full use of dependency information in the spatial domain and long-term and short-term information in the temporal domain to improve the accuracy of probabilistic prediction.

[0211] This application proposes a decoupled entropy model that further improves decoding speed while utilizing spatiotemporal information. The branch that generates entropy decoding parameters is independent of the branch that introduces spatiotemporal information (decoupled), so that entropy decoding can decode all bitstreams at once. Then, the spatiotemporal information branch is used to further process the entropy decoding results to obtain decoding features.

[0212] This application proposes a channel-complementary chessboard spatial context model, which makes the most of the information in the complete receptive field around the current point in the spatial domain to predict the information of the current point.

[0213] This application designs a block-based context model that uses neighboring blocks to infer information about the current block, which can utilize spatial information and control decoding speed.

[0214] This application designs a time-domain context model, which can incorporate long-term and short-term information in the time domain.

[0215] The networks and modules proposed in this application can be combined to form multiple solutions.

[0216] The above embodiments are merely one common example of this application and do not constitute any limitation on the technical scope of this application. Therefore, any minor modifications, equivalent changes, or alterations made to the above content based on the substance of the solution of this application shall still fall within the scope of the technical solution of this application.

[0217] Please continue reading Figure 21 , Figure 21 This is a schematic diagram of an embodiment of the image encoding / decoding apparatus provided in this application. The image encoding / decoding apparatus 500 of this application embodiment includes a processor 51, a memory 52, an input / output device 53, and a bus 54.

[0218] The processor 51, memory 52, and input / output device 53 are respectively connected to the bus 54. The memory 52 stores program data, and the processor 51 is used to execute the program data to implement the image decoding method described in the above embodiments.

[0219] In this embodiment, processor 51 can also be referred to as a CPU (Central Processing Unit). Processor 51 may be an integrated circuit chip with signal processing capabilities. Processor 51 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor can be a microprocessor, or processor 51 can be any conventional processor.

[0220] This application also provides a computer storage medium; please refer to the following: Figure 22 , Figure 22 This is a schematic diagram of a computer storage medium according to an embodiment of the present application. The computer storage medium 700 stores program data 71, which is used to implement the image decoding method of the above embodiment when executed by the processor.

[0221] When embodiments of this application are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, etc.) to...

[0222] The storage medium may be a network device or a processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0223] The above description is merely an embodiment of this application and does not limit the patent scope of this application. Equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. An image decoding method, characterized in that, The image decoding method includes: Obtain the bitstream to be decoded; Based on the code stream to be decoded, obtain the spatial domain decoded information of the current frame corresponding to the code stream to be decoded; Input the decoded spatial information into the spatial context model to obtain the context information transmitted by the decoded spatial information; The context information and the prior decoding information of the bitstream to be decoded are input into the fusion network to obtain the probability parameters of the undecoded bitstream in the bitstream to be decoded; The undecoded bitstream in the bitstream to be decoded is decoded using the probability parameters to obtain the decoding features of the image.

2. The image decoding method according to claim 1, characterized in that, The converged network includes a first converged network and a second converged network; The step of inputting the context information and the priori decoding information of the bitstream to be decoded into the fusion network to obtain the probability parameters of the undecoded bitstream in the bitstream to be decoded includes: The context information and the prior decoding information of the bitstream to be decoded are input into the first fusion network to obtain the mean value of the probability parameters of the undecoded bitstream in the bitstream to be decoded; The prior decoding information of the bitstream to be decoded is input into the second fusion network to obtain the variance of the probability parameters of the undecoded bitstream in the bitstream to be decoded.

3. The image decoding method according to claim 1, characterized in that, The image decoding method further includes: The current frame is divided into blocks according to a preset direction to obtain several feature blocks of the current frame; The step of obtaining the spatial domain decoded information of the current frame corresponding to the bitstream to be decoded, based on the bitstream to be decoded, includes: Based on the bitstream of the previous feature block of the current feature block in the bitstream to be decoded, obtain the decoded information of the previous feature block of the current feature block.

4. The image decoding method according to claim 3, characterized in that, The preset direction includes the channel direction and / or the airspace direction.

5. The image decoding method according to claim 1, characterized in that, The step of obtaining the spatial domain decoded information of the current frame corresponding to the bitstream to be decoded, based on the bitstream to be decoded, includes: Based on the bitstreams of the neighboring feature blocks of the current feature block in the bitstream to be decoded, the spatial domain decoded information of the neighboring feature blocks of the current feature block is obtained. The neighboring feature blocks include the decoded feature blocks that are adjacent to the current feature block in the spatial domain on the top, left, right, bottom and / or diagonal lines, respectively. The decoded feature blocks have some channel feature blocks that have been decoded.

6. The image decoding method according to claim 5, characterized in that, The step of inputting the decoded spatial information into the spatial context model to obtain the context information transmitted by the decoded spatial information includes: Obtain the undecoded and decoded channels of the current feature block; In the decoded channel of the current feature block, the first convolutional kernel of the spatial context model is selected to obtain the first decoded feature block in the spatial decoded information with the current feature block as the center; In the undecoded channel of the current feature block, the second convolution kernel of the spatial context model is selected to obtain the second decoded feature block in the spatial decoded information with the current feature block as the center; The first context information of the first decoded feature block is extracted from the decoded channel of the current feature block, and the second context information of the second decoded feature block is extracted from the undecoded channel of the current feature block. The context information for transmitting the spatially decoded information is obtained by combining the first context information and the second context information according to the channel combination; Wherein, the first decoded feature block includes the decoded feature blocks that are adjacent to the current feature block in the spatial domain above, to the left, to the right, and below, respectively; the first decoded feature block includes the decoded feature blocks that are adjacent to the current feature block on the diagonal in the spatial domain.

7. The image decoding method according to claim 1, characterized in that, The image decoding method further includes: Based on the bitstream to be decoded, obtain the previously decoded frames preceding the bitstream to be decoded; The step of inputting the context information and the priori decoding information of the bitstream to be decoded into the fusion network to obtain the probability parameters of the undecoded bitstream in the bitstream to be decoded includes: The cached information of the decoded frames, the context information, and the priori decoding information are input into the fusion network to obtain the probability parameters of the undecoded bitstream in the bitstream to be decoded; The cached information refers to the decoding information of the decoded frame and / or intermediate variables generated during the decoding process.

8. The image decoding method according to claim 7, characterized in that, The step of inputting the spatially decoded information into the spatial context model to obtain the decoding information of the bitstream to be decoded includes: Input the decoded spatial information and the cached information into the spatial context model to obtain the context information transmitted by the decoded spatial information.

9. The image decoding method according to claim 7, characterized in that, The cached information includes long-term information and short-term information; The cached information is extracted through a recurrent neural network in the entropy model. The recurrent neural network is set in any one of the networks in the entropy model, or it is set before any one of the networks in the entropy model, or it is set after any one of the networks except the fusion network.

10. An image decoding method, characterized in that, The image decoding method includes: Obtain the bitstream to be decoded; Based on the bitstream to be decoded, obtain the previously decoded frames preceding the bitstream to be decoded; The buffer information of the decoded frame and the prior decoding information of the bitstream to be decoded are input into the fusion network to obtain the probability parameters of the undecoded bitstream in the bitstream to be decoded. The undecoded bitstream in the bitstream to be decoded is decoded using the probability parameters to obtain the decoding features of the image; The cached information refers to the decoding information of the decoded frame and / or intermediate variables generated during the decoding process.

11. The image decoding method according to claim 10, characterized in that, The image decoding method further includes: Based on the code stream to be decoded, obtain the spatial domain decoded information of the current frame corresponding to the code stream to be decoded; Input the decoded spatial information into the spatial context model to obtain the context information transmitted by the decoded spatial information; The step of inputting the buffer information of the decoded frames and the prior decoding information of the bitstream to be decoded into the fusion network to obtain the probability parameters of the undecoded bitstream in the bitstream to be decoded includes: The buffer information of the decoded frames, the context information, and the prior decoding information are input into the fusion network to obtain the probability parameters of the undecoded bitstream in the bitstream to be decoded.

12. An image encoding and decoding apparatus, characterized in that, The image encoding / decoding device includes a memory and a processor coupled to the memory; The memory is used to store program data, and the processor is used to execute the program data to implement the image decoding method as described in any one of claims 1 to 11.

13. A computer storage medium, characterized in that, The computer storage medium is used to store program data, which, when executed by the computer, is used to implement the image decoding method as described in any one of claims 1 to 11.