Dimension self-derivation image coding and decoding method and device and storage equipment

By self-derived tensor dimension information in the image encoding and decoding method, the problem of redundancy in the transmission of tensor dimension information in the prior art is solved, and more efficient image encoding is achieved.

CN120017852APending Publication Date: 2025-05-16ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411657770.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-16
Filing Date
2024-11-18
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing intelligent image encoding and decoding methods based on neural networks have redundancy when transmitting tensor dimension information, resulting in additional code rate overhead.

Method used

A dimensional self-deduction image encoding and decoding method is proposed. By obtaining the tensor dimension information required for other decoding in the decoder, the tensor dimension information that needs to be transmitted during encoding is reduced.

Benefits of technology

Without affecting the decoding process, the amount of data consumed by the encoded image is reduced and the encoding efficiency is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120017852A_ABST
    Figure CN120017852A_ABST
Patent Text Reader

Abstract

The invention discloses a dimension self-derivation image encoding and decoding method and device, and the encoding method comprises the steps: an encoder encodes an input original image tensor and first tensor dimension information, and obtains a code stream, and the first tensor dimension information comprises the width and height of the tensor; the decoding method comprises the following steps: a decoder analyzes syntax elements from an input code stream to obtain first tensor dimension information; according to the obtained first tensor dimension information, tensor dimension information required by other decoding is obtained through self-derivation based on the correlation between the tensor dimension information; and decoding the input code stream based on the obtained information of each tensor dimension. According to the encoding and decoding method and device, tensor dimension information needing to be encoded and transmitted during encoding can be reduced, the data size consumed by image encoding is reduced while the decoding process is not affected, and the encoding efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image coding and decoding technology, and more specifically, to a dimension self-deriving image coding and decoding method and device. Background Art

[0002] The existing encoding and decoding processing of image features includes an intelligent encoding and decoding method based on a neural network and a hybrid encoding and decoding method based on a combination of a neural network and a traditional encoding and decoding method. The intelligent encoding and decoding method based on a neural network is to use a neural network to transform the image, convert the image into a feature tensor, and perform operations such as downsampling, quantization, and entropy coding to achieve compression coding. The hybrid encoding and decoding method based on a combination of a neural network and a traditional encoding and decoding method first uses a neural network to transform the image, convert the image into a feature tensor, and perform downsampling, then fix and concatenate the feature tensor, convert it into a data type that can be input by the traditional encoding and decoding method, and then use the traditional encoding and decoding method to achieve compression coding.

[0003] Intelligent coding and decoding methods based on neural networks have achieved good performance in the field of image compression. The network structures used in intelligent coding and decoding methods can be mainly divided into structures based on recurrent neural networks (RNNs) and autoencoder structures based on convolutional neural networks (CNNs). Among them, the intelligent coding and decoding method based on the autoencoder structure of convolutional neural networks (CNNs) has a higher coding efficiency. In recent years, an intelligent coding and decoding method based on the autoencoder structure of convolutional neural networks (CNNs) combined with super prior information has been widely studied and applied, and has become the current mainstream intelligent coding and decoding method. Among them, super prior information is used to further eliminate the spatial redundancy of the image feature information tensor obtained by the neural network transformation of the image, so as to improve the coding efficiency. The image feature information tensor is usually called the Y feature tensor, and the super prior information tensor is called the Z feature tensor.

[0004] The existing intelligent coding and decoding method based on neural network needs to encode the image feature tensor, super prior information tensor, reconstructed feature tensor output by the decoder and the dimension information of the original image tensor required by the decoder in the decoding process and transmit them in the bitstream. The problem with this method is that the dimensions of the image feature tensor, super prior information tensor, output feature tensor and original image tensor required by the decoder are correlated, resulting in redundancy between the tensor dimension information currently encoded and transmitted, which in turn increases the additional bit rate overhead. Summary of the invention

[0005] In order to solve the above-mentioned shortcomings of the prior art, the present invention proposes a dimension self-deriving image encoding and decoding method and device, including the following aspects:

[0006] A first aspect of the present invention provides a dimension self-deriving image encoding method, comprising:

[0007] The encoder encodes the input original image tensor and first tensor dimension information to obtain a code stream, where the first tensor dimension information includes the width and height of the tensor.

[0008] A second aspect of the present invention provides a dimension self-deriving image decoding method, comprising:

[0009] The decoder parses the syntax elements from the input bitstream to obtain the first tensor dimension information;

[0010] According to the first tensor dimension information obtained, based on the correlation between the tensor dimension information, other tensor dimension information required for decoding is self-derived;

[0011] Based on the acquired information of each tensor dimension, the decoder decodes the input bitstream.

[0012] Furthermore, the other tensor dimension information required for decoding includes information used to determine the dimension of the tensor filled with the element values ​​parsed from the bitstream and the dimension of the decoder output reconstructed feature tensor or reconstructed image tensor.

[0013] Furthermore, the correlation between the tensor dimension information is pre-defined offline and shared by the encoder and decoder.

[0014] Further, the first tensor dimension information is Z feature tensor dimension information, and the self-deriving of other tensor dimension information required for decoding based on the correlation between the tensor dimension information according to the acquired first tensor dimension information includes:

[0015] Initialize the Z feature initialization tensor based on the dimension information of the Z feature tensor;

[0016] Performing a first entropy decoding process on the code stream to obtain element values, and filling the obtained element values ​​into the Z feature initialization tensor based on a first filling method to obtain a Z feature tensor;

[0017] Processing the Z feature tensor through a first decoding network to obtain a Z feature representation tensor;

[0018] The dimension information of the Z feature tensor is processed by the first self-derivation to obtain the dimension information of the Y feature tensor;

[0019] Initialize the Y feature initialization tensor based on the Y feature tensor dimension information;

[0020] Performing a second entropy decoding process on the code stream to obtain element values, and filling the obtained element values ​​into the Y feature initialization tensor based on a second filling method to obtain a Y feature tensor;

[0021] The Y feature tensor is processed by the second decoding network to obtain the reconstructed feature tensor.

[0022] Further, the first tensor dimension information is Y feature tensor dimension information, and the tensor dimension information required for decoding is derived based on the correlation between the tensor dimension information according to the acquired first tensor dimension information, including:

[0023] Based on the Y feature tensor dimension information, the Z feature tensor dimension information is obtained through the first self-derivation process;

[0024] Initialize the Z feature initialization tensor based on the dimension information of the Z feature tensor;

[0025] Performing a first entropy decoding process on the code stream to obtain element values, and filling the obtained element values ​​into the Z feature initialization tensor based on a first filling method to obtain a Z feature tensor;

[0026] Processing the Z feature tensor through a first decoding network to obtain a Z feature representation tensor;

[0027] Initialize the Y feature initialization tensor based on the Y feature tensor dimension information;

[0028] Performing a second entropy decoding process on the code stream to obtain element values, and filling the obtained element values ​​into the Y feature initialization tensor based on a second filling method to obtain a Y feature tensor;

[0029] The Y feature tensor is processed by the second decoding network to obtain the reconstructed feature tensor.

[0030] Furthermore, the first tensor dimension information is the reconstructed feature tensor dimension information, and the tensor dimension information required for decoding is derived based on the correlation between the tensor dimension information according to the acquired first tensor dimension information, including:

[0031] Based on the reconstructed feature tensor dimension information, the Z feature tensor dimension information is obtained through a first self-derivation process;

[0032] Initialize the Z feature initialization tensor based on the dimension information of the Z feature tensor;

[0033] Performing a first entropy decoding process on the code stream to obtain element values, and filling the obtained element values ​​into the Z feature initialization tensor based on a first filling method to obtain a Z feature tensor;

[0034] Processing the Z feature tensor through a first decoding network to obtain a Z feature representation tensor;

[0035] Based on the reconstructed feature tensor dimension information, the Y feature tensor dimension information is obtained through the first self-derivation process;

[0036] Initialize the Y feature initialization tensor based on the Y feature tensor dimension information;

[0037] Performing a second entropy decoding process on the code stream to obtain element values, and filling the obtained element values ​​into the Y feature initialization tensor based on a second filling method to obtain a Y feature tensor;

[0038] The Y feature tensor is processed by the second decoding network to obtain the reconstructed feature tensor.

[0039] Furthermore, the first self-derivation processing is defined as scaling the width and height in the tensor dimension information to S times the original value, where S is an integer power of 2.

[0040] Furthermore, the first tensor dimension information is the tensor dimension information of the original image, and the tensor dimension information required for decoding is derived based on the correlation between the tensor dimension information according to the acquired first tensor dimension information, including:

[0041] Based on the original image tensor dimension information, the Z feature tensor dimension information is obtained through the second self-derivation process;

[0042] Initialize the Z feature initialization tensor based on the dimension information of the Z feature tensor;

[0043] Performing a first entropy decoding process on the code stream to obtain element values, and filling the obtained element values ​​into the Z feature initialization tensor based on a first filling method to obtain a Z feature tensor;

[0044] Processing the Z feature tensor through a first decoding network to obtain a Z feature representation tensor;

[0045] Based on the original image tensor dimension information, the Y feature tensor dimension information is obtained through the second self-derivation process;

[0046] Initialize the Y feature initialization tensor based on the Y feature tensor dimension information;

[0047] Performing a second entropy decoding process on the code stream to obtain element values, and filling the obtained element values ​​into the Y feature initialization tensor based on a second filling method to obtain a Y feature tensor;

[0048] Processing the Y feature tensor through a second decoding network to obtain a processed reconstructed feature tensor;

[0049] Processing the processed reconstructed feature tensor through a third decoding network to obtain a processed reconstructed image tensor;

[0050] Based on the dimensional information of the original image tensor, the processed reconstructed image tensor is inversely processed in a preset defined manner to obtain a reconstructed image tensor.

[0051] Further, the second self-derivation process includes:

[0052] First, the original image tensor dimension information is processed based on a preset definition method to obtain the processed image tensor dimension information;

[0053] The width and height in the processed image tensor dimension information are scaled to S times the original values, where S is an integer power of 2.

[0054] A third aspect of the present invention provides a dimensionality self-deriving image encoding device, comprising:

[0055] The encoding unit is used to encode the input original image tensor and first tensor dimension information to obtain a code stream, where the first tensor dimension information includes the width and height of the tensor.

[0056] A fourth aspect of the present invention provides a dimensional self-deriving image decoding device, comprising:

[0057] A syntax element parsing unit, configured to parse syntax elements from an input bitstream to obtain first tensor dimension information;

[0058] A self-deriving unit, which is used to self-derive other tensor dimension information required for decoding based on the correlation between the tensor dimension information according to the acquired first tensor dimension information;

[0059] A decoding unit is used to decode the input code stream based on the acquired tensor dimension information and fill and transform the decoded element values ​​to obtain the reconstructed feature / image tensor.

[0060] A fifth aspect of the present invention provides a storage device, which includes a code stream, which is encoded using the dimensional self-derived image encoding method as described in the first aspect above, and is decoded using the dimensional self-derived image decoding method as described in the second aspect above.

[0061] The image encoding and decoding method and device for self-deriving dimension of the present invention can only encode and decode the dimension information of the first tensor required, obtain the dimension information of the first tensor during the decoding process, and self-derive the dimension information of other tensors based on the obtained tensor dimension information. It can reduce the tensor dimension information that needs to be transmitted during encoding, reduce the amount of data consumed by the encoded image without affecting the decoding process, and improve the encoding efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0063] Figure 1 It is a schematic diagram of the overall framework of the image encoding and decoding device of the present invention;

[0064] Figure 2 Schematic diagram of the framework of the image encoding and decoding method based on self-derivation of Z feature tensor dimension information in an embodiment of the present invention;

[0065] Figure 3 Schematic diagram of the framework of the image encoding and decoding method based on the self-derivation of the Y feature tensor dimension information in an embodiment of the present invention;

[0066] Figure 4 Schematic diagram of the framework of the image encoding and decoding method based on self-deriving of reconstructed feature tensor dimension information in an embodiment of the present invention;

[0067] Figure 5 Schematic diagram of the framework of the image encoding and decoding method based on self-derivation of original image tensor dimensional information in an embodiment of the present invention. DETAILED DESCRIPTION

[0068] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention is further described in detail below with reference to the accompanying drawings and embodiments.

[0069] This embodiment discloses a dimension self-deriving image encoding and decoding method, wherein the encoding method and the decoding method are shown in the attached Figure 1 The implementation subject of the method can be a set of encoders and decoders, wherein the encoder includes the following steps:

[0070] The encoder encodes the input original picture tensor and the first tensor dimension information to obtain a code stream, where the first tensor dimension information includes the width and height of the tensor;

[0071] The decoder includes the following steps:

[0072] The decoder parses the syntax element from the input bitstream to obtain first tensor dimension information, where the first tensor dimension information includes the width and height of the tensor;

[0073] According to the first tensor dimension information obtained, based on the correlation between the tensor dimension information, other tensor dimension information required for decoding is self-derived;

[0074] Based on the acquired information of each tensor dimension, the decoder decodes the input bitstream.

[0075] The encoding method is described in detail below.

[0076] Specifically, the encoder encodes the input original image tensor and the first tensor dimension information to obtain a code stream.

[0077] The decoding method is described in detail below.

[0078] For the input bitstream, the decoder obtains the first tensor dimension information from the bitstream. The tensor dimension information includes the width and height of the tensor. This information directly contains the dimension information of the first tensor in the decoder and indirectly contains the dimension information of other tensors in the decoder. In addition, the number of channels of the tensor is defined offline, and this information is shared by the encoder and decoder, and does not need to be included in the bitstream. The decoder can use the obtained tensor dimension information to self-derive the tensor dimension information required by other decoders. These tensor dimension information are used to determine the dimensions of the tensor filled with the element values ​​parsed from the bitstream and the dimensions of the decoder output reconstructed feature tensor or reconstructed image tensor. The self-derived dimension method uses the correlation between the tensor dimension information in the decoder to define offline. Usually, the tensor dimension information contained in the bitstream is only the dimension information of the first tensor. The advantage of this is that the dimension information of other tensors in the decoder is self-derived through the parsed tensor dimension information, which can reduce the amount of data in the bitstream.

[0079] The above is an overview of the entire encoding method and decoding method process. The following will describe the details of each operation in detail with reference to specific examples.

[0080] In one embodiment, if Figure 2 As shown, for the encoding method, the encoder encodes the input original image tensor and Z feature tensor dimension information to obtain a code stream.

[0081] For the decoding method, the decoder obtains the Z feature tensor dimension information from the syntax elements of the code stream, including the width W and height H of the Z feature tensor. In this embodiment, the number of channels of the Z feature tensor is defined as N (for example, N=128). According to the dimension requirements of the Z feature tensor, a Z feature initialization tensor with all element values ​​​​being 0 is initialized, and then the code stream is processed by the first entropy decoding. In this embodiment, the first entropy decoding uses the probability model defined offline for entropy decoding. In the entropy decoding process, all element values ​​in a single channel of the Z feature initialization tensor are obtained in turn, and then all element values ​​in the obtained single channel are filled into the corresponding channel of the Z feature initialization tensor according to the scanning order of the raster based on the first filling method. In this way, each channel of the Z feature initialization tensor is filled in turn until all channels of the Z feature initialization tensor are filled to obtain the Z feature tensor. Then the Z feature tensor is processed by the first decoding network to obtain the Z feature representation tensor.

[0082] The dimension information of the Z feature tensor is processed by the first self-derivation to obtain the dimension information of the Y feature tensor. In this embodiment, the number of channels of the Y feature tensor is defined as N (for example, N=128), and the first self-derivation process is to scale the width and height in the tensor dimension information to S times the original value (for example, S is 2 to the power of 2). Then, according to the dimension requirements of the Y feature tensor, a Y feature initialization tensor with all element values ​​​​being 0 is initialized, and then the code stream is processed by the second entropy decoding. In this embodiment, the second entropy decoding utilizes a probability estimation model, a Y feature initialization tensor, and a Z feature characterization tensor for entropy decoding, and the probability estimation model uses an autoregressive probability estimation model. In the entropy decoding process, the element value of each channel and each position of the Y feature initialization tensor is obtained in turn, and the element value is filled into the corresponding position of the Y feature initialization tensor based on the second filling method. Specifically, first obtain the element value of the upper left corner position of the first channel of the Y feature initialization tensor and fill it, then obtain the element value of the same position of the second channel of the Y feature initialization tensor and fill it, then obtain the element value of the same position of the third channel of the Y feature initialization tensor and fill it, and so on, until the upper left corner position of all channels of the Y feature initialization tensor is filled. Then, in the order of raster scanning, follow the above process to fill other positions in turn until all channels and positions of the Y feature initialization tensor are filled to obtain the Y feature tensor. Then, the Y feature tensor is processed by the second decoding network to obtain the reconstructed feature tensor.

[0083] In this embodiment, a grammatical structure is shown in Table 1.

[0084] Table 1 An example of a grammatical structure

[0085]

[0086]

[0087] Among them, the grammatical structure image_header() is the image header data, image_feature_data() is the image feature encoding data, HyperFeatureWidth and HyperFeatureHeight are the width and height in the Z feature tensor dimension information, LowResImageFeatureWidth and LowResImageFeatureHeight are the width and height in the Y feature tensor dimension information, and NumChannel is the number of channels of the Z feature tensor and the Y feature tensor. Among them, the grammatical semantics are as follows:

[0088] icm_header_start_code: marks the beginning of the image header data;

[0089] first_feature_width_minus1: represents the width in the first tensor dimension information, an 8-bit unsigned integer. In this embodiment, it represents the width in the Z feature tensor dimension information;

[0090] first_feature_height_minus1: represents the height in the first tensor dimension information, an 8-bit unsigned integer. In this embodiment, it represents the height in the Z feature tensor dimension information;

[0091] icm_feature_start_code: marks the beginning of image feature data;

[0092] model_id: indicates the label of the network model used for decoding operation, a 4-bit unsigned integer;

[0093] hyper_feature_sub_bitstream_byte_length[i]: indicates the length of the sub-bitstream corresponding to all Z feature element values ​​of the i-th channel, in bytes, a 16-bit unsigned integer;

[0094] hyper_feature_value[i][j][k]: represents the element value of the i-th channel, j-th row, and k-th column in the Z feature tensor. The parsing process is predefined by ne(v).

[0095] image_feature_sub_bitstream_byte_length[j][k]: indicates the length of the sub-bitstream corresponding to all Y feature element values ​​in the j-th row and k-th column, in bytes, a 16-bit unsigned integer.

[0096] image_feature_value[i][j][k]: represents the Y feature element value of the i-th channel, j-th row, k-th column in the image feature tensor. The parsing process is predefined by ne(v);

[0097] In another embodiment, if Figure 3 As shown, for the encoding method, the encoder encodes the dimension information of the input original image tensor and the Y feature tensor to obtain a code stream.

[0098] For the decoding method, the decoder obtains the dimension information of the Y feature tensor from the syntax elements of the code stream, including the width W and height H of the Y feature tensor. In this embodiment, the number of channels of the Y feature tensor is defined as N (for example, N=128). The dimension information of the Y feature tensor is self-derived to obtain the dimension information of the Z feature tensor. In this embodiment, the number of channels of the Z feature tensor is defined as N (for example, N=128), and the first self-derived process is to scale the width and height in the tensor dimension information to S times the original value (for example, S is 2 to the power of -2). According to the dimension requirements of the Z feature tensor, a Z feature initialization tensor with all element values ​​​​being 0 is initialized. Then the code stream is processed by the first entropy decoding, and the Z feature tensor is obtained based on the first filling method. In this embodiment, the first entropy decoding uses a probability model defined offline for entropy decoding. Then the Z feature tensor is processed by the first decoding network to obtain the Z feature representation tensor.

[0099] According to the dimensionality requirement of the Y feature tensor, a Y feature initialization tensor whose element values ​​are all 0 is initialized. Then the code stream is processed by the second entropy decoding, and the Y feature tensor is obtained based on the second filling method. In this embodiment, the second entropy decoding uses a probability estimation model, a Y feature initialization tensor, and a Z feature characterization tensor for entropy decoding, and the probability estimation model uses an autoregressive probability estimation model. Then the Y feature tensor is processed by the second decoding network to obtain a reconstructed feature tensor.

[0100] In this embodiment, a syntax structure is the same as Table 1. In the syntax semantics, first_feature_width_minus1 in this embodiment indicates the width in the Y feature tensor dimension information; first_feature_height_minus1 in this embodiment indicates the height in the Y feature tensor dimension information.

[0101] In another embodiment, if Figure 4 As shown, for the encoding method, the encoder encodes the input original image tensor and the reconstructed feature tensor dimension information to obtain a code stream.

[0102] For the decoding method, the decoder obtains the reconstructed feature tensor dimension information from the syntax elements of the code stream, including the width W and height H of the reconstructed feature tensor. In this embodiment, the number of channels of the reconstructed feature tensor is defined as N (for example, N=128). The dimension information of the reconstructed feature tensor is processed by the first self-derivation to obtain the dimension information of the Z feature tensor. In this embodiment, the number of channels of the Z feature tensor is defined as N (for example, N=128), and the first self-derivation process is to scale the width and height in the tensor dimension information to S times the original value (for example, S is 2 to the power of -4). According to the dimension requirements of the Z feature tensor, a Z feature initialization tensor with all element values ​​​​being 0 is initialized, and then the code stream is processed by the first entropy decoding, and the Z feature tensor is obtained based on the first filling method. In this embodiment, the first entropy decoding uses a probability model defined offline for entropy decoding. Then the Z feature tensor is processed by the first decoding network to obtain the Z feature representation tensor.

[0103] The dimensional information of the reconstructed feature tensor is processed by the first self-derivation to obtain the dimensional information of the Y feature tensor. In this embodiment, the number of channels of the Y feature tensor is defined as N (for example, N=128), and the first self-derivation process is to scale the width and height in the tensor dimensional information to S times the original value (for example, S is 2 to the power of -2). According to the dimensional requirements of the Y feature tensor, a Y feature initialization tensor with all element values ​​​​being 0 is initialized, and then the code stream is processed by the second entropy decoding, and the Y feature tensor is obtained based on the second filling method. In this embodiment, the second entropy decoding utilizes a probability estimation model, a Y feature initialization tensor, and a Z feature characterization tensor for entropy decoding, and the probability estimation model uses an autoregressive probability estimation model. Then the Y feature tensor is processed by the second decoding network to obtain a reconstructed feature tensor.

[0104] In this embodiment, a syntax structure is the same as Table 1. In the syntax semantics, first_feature_width_minus1 in this embodiment indicates the width in the dimension information of the reconstructed feature tensor; first_feature_height_minus1 in this embodiment indicates the height in the dimension information of the reconstructed feature tensor.

[0105] In another embodiment, if Figure 5 As shown, for the encoding method, the encoder processes the original image tensor based on a preset definition method to obtain a processed image tensor, and then encodes the processed image tensor and the original image tensor dimensional information to obtain a code stream, and the tensor dimensional information contained in the code stream is only the original image tensor dimensional information.

[0106] For the decoding method, the decoder obtains the original image tensor dimension information from the syntax elements of the code stream, including the width W and height H of the original image tensor. In this embodiment, the number of channels of the original image tensor is defined as 3. The original image tensor dimension information is processed by the second self-derivation to obtain the dimension information of the Z feature tensor. In this embodiment, the number of channels of the Z feature tensor is defined as N (for example, N=128), and the second self-derivation is to process the original image tensor dimension information based on the preset definition method to obtain the processed image tensor dimension information, and then scale the width and height in the processed image tensor dimension information to S times the original value (for example, S is 2 to the power of -6). In this embodiment, the preset definition method is shared by the encoder and the decoder, and the information does not need to be included in the code stream. According to the dimension requirements of the Z feature tensor, a Z feature initialization tensor with all element values ​​​​being 0 is initialized, and then the code stream is processed by the first entropy decoding, and the Z feature tensor is obtained based on the first filling method. In this embodiment, the first entropy decoding uses the probability model defined offline for entropy decoding. Then the Z feature tensor is processed by the first decoding network to obtain the Z feature representation tensor.

[0107] The original image tensor dimension information is processed by the second self-derivation to obtain the dimension information of the Y feature tensor. In this embodiment, the number of channels of the Y feature tensor is defined as N (for example, N=128), and the second self-derivation is to process the original image tensor dimension information based on a preset definition method to obtain the processed image tensor dimension information, and then scale the width and height in the processed image tensor dimension information to S times the original value (for example, S is 2 to the power of -4). According to the dimension requirements of the Y feature tensor, a Y feature initialization tensor with all element values ​​​​are initialized to 0, and then the code stream is processed by the second entropy decoding, and the Y feature tensor is obtained based on the second filling method. In this embodiment, the second entropy decoding uses a probability estimation model, a Y feature initialization tensor and a Z feature characterization tensor for entropy decoding, and the probability estimation model uses an autoregressive probability estimation model. Then the Y feature tensor is processed by the second decoding network to obtain a processed reconstructed feature tensor. Then the processed reconstructed feature tensor is processed by the third decoding network to obtain a processed reconstructed image tensor. Based on the dimensional information of the original image tensor, the processed reconstructed image tensor is subjected to a preset defined inverse processing to obtain a reconstructed image tensor.

[0108] In this embodiment, a syntax structure is the same as Table 1. In the syntax semantics, first_feature_width_minus1 in this embodiment indicates the width in the original image tensor dimension information; first_feature_height_minus1 in this embodiment indicates the height in the original image tensor dimension information.

[0109] This embodiment also discloses a dimensional self-deriving image encoding device, including:

[0110] An encoding unit, configured to encode an input original image tensor and first tensor dimension information to obtain a code stream, wherein the first tensor dimension information includes a width and a height of the tensor;

[0111] This embodiment also discloses a dimensional self-deriving image decoding device, including:

[0112] A syntax element parsing unit, configured to parse syntax elements from an input bitstream to obtain first tensor dimension information;

[0113] A self-deriving unit, which is used to self-derive other tensor dimension information required for decoding based on the correlation between the tensor dimension information according to the acquired first tensor dimension information;

[0114] A decoding unit is used to decode the input code stream based on the acquired tensor dimension information and fill and transform the decoded element values ​​to obtain the reconstructed feature / image tensor.

[0115] This embodiment also discloses a coding device, which includes a processor and a memory, and is used to execute the coding method disclosed in the present invention.

[0116] This embodiment also discloses a decoding device, which includes a processor and a memory, and is used to execute the decoding method disclosed in the present invention.

[0117] This embodiment also discloses a storage device, which includes a code stream, which is encoded using the dimension self-deriving image encoding method disclosed in the present invention and decoded using the image decoding method.

[0118] The above embodiments are only used to help understand the method and core idea of ​​the present invention. It should be noted that, for those skilled in the art, several improvements and modifications can be made to the present invention without departing from the principles of the present invention, and these improvements and modifications also fall within the scope of protection of the claims of the present invention.

Claims

1. A dimensionally self-derived image coding method, characterized in that: include: The encoder encodes the input original image tensor and first tensor dimension information to obtain a code stream, where the first tensor dimension information includes the width and height of the tensor.

2. A dimensional self-deriving image decoding method, characterized in that: include: The decoder parses the syntax elements from the input bitstream to obtain the first tensor dimension information; According to the first tensor dimension information obtained, based on the correlation between the tensor dimension information, other tensor dimension information required for decoding is self-derived; Based on the obtained information of each tensor dimension, the input code stream is decoded.

3. The image decoding method of claim 2, wherein: The other tensor dimension information required for decoding includes the dimension used to determine the dimension of the tensor filled with the element value parsed from the code stream and the dimension of the decoder output reconstructed feature tensor or reconstructed image tensor.

4. The image decoding method of claim 2, wherein: The correlation between the tensor dimension information is pre-defined offline and shared by the encoder and decoder.

5. The image decoding method of claim 2, wherein: The first tensor dimension information is the Z feature tensor dimension information, and the self-deriving of other tensor dimension information required for decoding based on the correlation between the tensor dimension information according to the acquired first tensor dimension information includes: Initialize the Z feature initialization tensor based on the dimension information of the Z feature tensor; Performing a first entropy decoding process on the code stream to obtain element values, and filling the obtained element values ​​into the Z feature initialization tensor based on a first filling method to obtain a Z feature tensor; Processing the Z feature tensor through a first decoding network to obtain a Z feature representation tensor; The dimension information of the Z feature tensor is processed by the first self-derivation to obtain the dimension information of the Y feature tensor; Initialize the Y feature initialization tensor based on the Y feature tensor dimension information; Performing a second entropy decoding process on the code stream to obtain element values, and filling the obtained element values ​​into the Y feature initialization tensor based on a second filling method to obtain a Y feature tensor; The Y feature tensor is processed by the second decoding network to obtain the reconstructed feature tensor.

6. The dimensional self-deriving image decoding method according to claim 5, characterized in that: The first self-derivation process is defined as scaling the width and height in the tensor dimension information to S times the original value, where S is an integer power of 2.

7. The dimensional self-deriving image decoding method according to claim 2, characterized in that: The first tensor dimension information is the tensor dimension information of the original image, and the tensor dimension information required for decoding is derived based on the correlation between the tensor dimension information according to the acquired first tensor dimension information, including: Based on the original image tensor dimension information, the Z feature tensor dimension information is obtained through the second self-derivation process; Initialize the Z feature initialization tensor based on the dimension information of the Z feature tensor; Performing a first entropy decoding process on the code stream to obtain element values, and filling the obtained element values ​​into the Z feature initialization tensor based on a first filling method to obtain a Z feature tensor; Processing the Z feature tensor through a first decoding network to obtain a Z feature representation tensor; Based on the original image tensor dimension information, the Y feature tensor dimension information is obtained through the second self-derivation process; Initialize the Y feature initialization tensor based on the Y feature tensor dimension information; Performing a second entropy decoding process on the code stream to obtain element values, and filling the obtained element values ​​into the Y feature initialization tensor based on a second filling method to obtain a Y feature tensor; Processing the Y feature tensor through a second decoding network to obtain a processed reconstructed feature tensor; Processing the processed reconstructed feature tensor through a third decoding network to obtain a processed reconstructed image tensor; Based on the dimensional information of the original image tensor, the processed reconstructed image tensor is inversely processed in a preset defined manner to obtain a reconstructed image tensor.

8. The image decoding method of claim 7, wherein: The second self-derivation process comprises: First, the original image tensor dimension information is processed based on a preset definition method to obtain the processed image tensor dimension information; The width and height in the processed image tensor dimension information are scaled to S times the original values, where S is an integer power of 2.

9. A dimensional self-deriving image coding device, characterized in that: include: The encoding unit is used to encode the input original image tensor and first tensor dimension information to obtain a code stream, where the first tensor dimension information includes the width and height of the tensor.

10. A dimensional self-deriving image decoding device, characterized in that: include: A syntax element parsing unit, configured to parse syntax elements from an input bitstream to obtain first tensor dimension information; A self-deriving unit, which is used to self-derive other tensor dimension information required for decoding based on the correlation between the tensor dimension information according to the acquired first tensor dimension information; A decoding unit is used to decode the input code stream based on the acquired tensor dimension information and fill and transform the decoded element values ​​to obtain the reconstructed feature / image tensor.

11. A storage device, characterized in that: The storage device contains a code stream, which is encoded using the dimensional self-deriving image encoding method as claimed in claim 1 and is decoded using the image decoding method as claimed in any one of claims 2-8.