Coding and decoding device, method and system and storage medium

Through feature space alignment, enhancement, dimensionality reduction and coding, the problem that existing encoding and codec technology is not suitable for machine vision tasks is solved, and the effective encoding and codec of multi-size features in machine vision tasks is realized, which improves task performance.

CN120236156APending Publication Date: 2025-07-01CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311829608.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-28
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The existing image or video encoding and decoding technology is aimed at human visual tasks and is not completely suitable for machine vision tasks. It requires a codec solution suitable for machine vision tasks.

Method used

Through feature space alignment, feature enhancement, channel dimensionality reduction and encoding, the coded code stream is obtained; and through decoding, channel dimensionality up, feature size reduction and feature enhancement, the encoding and decoding of multi-size output features is realized.

Benefits of technology

Effective encoding, transmission and decoding of multi-size features of the intermediate layer of machine vision task neural network is realized, and the performance of machine vision tasks is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236156A_ABST
    Figure CN120236156A_ABST
Patent Text Reader

Abstract

The invention provides a coding and decoding device, method and system and a storage medium, and relates to the field of artificial intelligence. For multi-size input features, a coding code stream is obtained through feature space alignment, feature enhancement, channel dimension reduction and coding, and for the coding code stream, multi-size output features are obtained through decoding, channel dimension raising, feature size reduction and feature enhancement, and a coding and decoding scheme suitable for machine vision tasks is provided. The multi-size features of the middle layer of the machine vision task neural network are coded, transmitted and decoded.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence, particularly to the field of machine vision encoding and decoding of images or videos, and more particularly to an encoding and decoding device, method, system, and storage medium. Background Art

[0002] Traditional image or video encoding and decoding technologies are oriented towards human vision tasks. With the development of artificial intelligence, more and more image or video encoding and decoding requirements are oriented towards machine vision tasks rather than human vision tasks.

[0003] Image or video encoding and decoding technologies oriented towards human vision tasks are not entirely suitable for machine vision tasks. In many cases, the encoding and decoding algorithms oriented towards human vision tasks are not the optimal solutions in machine vision application scenarios. Therefore, an encoding and decoding scheme suitable for machine vision tasks is needed. Summary of the Invention

[0004] In embodiments of the present disclosure, for multi-size input features, an encoded bitstream is obtained through feature space alignment, feature enhancement, channel dimensionality reduction, and encoding. For the encoded bitstream, multi-size output features are obtained through decoding, channel dimensionality increase, feature size restoration, and feature enhancement, providing an encoding and decoding scheme suitable for machine vision tasks, enabling the encoding, transmission, and decoding of multi-size features in the middle layer of a machine vision task neural network.

[0005] Some embodiments of the present disclosure propose an encoding device, including:

[0006] A feature space alignment module configured to spatially align N input features with different sizes to obtain N first features with the same width and height but different numbers of channels, and concatenate the N first features along the channel dimension to form a second feature, where N is greater than or equal to 2;

[0007] A first feature enhancement module configured to enhance the representation ability of the second feature to obtain a third feature, where the second feature and the third feature have the same size;

[0008] A feature channel dimensionality reduction module configured to reduce the dimensionality of the third feature along the channel dimension to obtain a fourth feature;

[0009] A feature encoding module configured to quantize and encode the fourth feature to obtain an encoded bitstream.

[0010] In some embodiments, the feature space alignment module is configured to perform 2-fold downsampling on the i-th input feature with a size of (C i , 2 i-1 ×H, 2 i-1 ×W), where C i-1 times downsampling, where C idenotes the number of channels of the i-th input feature, where i ∈ [1, N], and H and W respectively denote the height and width of the input feature with the smallest size.

[0011] In some embodiments, the feature space alignment module is: a nearest neighbor downsampling module, a bilinear interpolation downsampling module, a pooling downsampling module, a convolutional downsampling module, or a first residual network module, where the first residual network module is composed of the addition of a nearest neighbor downsampling module or a bilinear interpolation downsampling module and a pooling downsampling module or a convolutional downsampling module.

[0012] In some embodiments, the first feature enhancement module is: a first attention mechanism module or a second residual network module for enhancing the feature representation ability.

[0013] In some embodiments, the feature channel dimensionality reduction module is composed of a first convolutional module with a stride of one and a number of channels of the first channel number C′ or a third residual network module, or is composed of a channel truncation module, where C i denotes the number of channels of the i-th input feature, and C′ denotes the number of channels after dimensionality reduction during encoding.

[0014] In some embodiments, the feature encoding module is configured to perform quantization using at least one of the following: linear quantization, clustering quantization, principal component analysis (PCA) quantization, discrete cosine transform (DCT) quantization; or / and the feature encoding module is configured to perform encoding using at least one of the following: a non-artificial intelligence encoder or an artificial intelligence encoder.

[0015] Some embodiments of the present disclosure propose a decoding device, including:

[0016] A feature decoding module, configured to decode and dequantize the encoded bitstream to obtain a fifth feature, where the methods of decoding and dequantization are symmetric to the methods of encoding and quantization;

[0017] A feature channel dimensionality increase module, configured to increase the dimensionality of the fifth feature along the channel dimension to obtain a sixth feature;

[0018] A feature size restoration module, configured to restore the size of the sixth feature to obtain N seventh features, and the sizes of the N seventh features are the same as the sizes of the N input features;

[0019] A second feature enhancement module, configured to enhance the representation ability of the N seventh features to obtain N eighth features.

[0020] In some embodiments, the feature decoding module is configured to perform inverse quantization using at least one of the following: linear inverse quantization, clustering inverse quantization, principal component analysis (PCA) inverse quantization, discrete cosine transform (DCT) inverse quantization; or / and the feature decoding module is configured to perform decoding using at least one of the following: a non-artificial intelligence decoder or an artificial intelligence decoder.

[0021] In some embodiments, the feature channel upsampling module is composed of a second convolutional module or a fourth residual network module with a stride of one and a channel number of a second channel number C″, or is composed of a channel padding module, where C i represents the channel number of the i-th input feature, C′ represents the channel number after dimensionality reduction during the encoding process, and C″ represents the channel number after dimensionality upsampling during the decoding process.

[0022] In some embodiments, the feature size reduction module is composed of N parallel feature size reduction sub-modules. The i-th feature size reduction sub-module corresponding to the i-th input feature is implemented by a transposed convolution or a sub-pixel convolution with a channel number of C i and a stride of 2 i-1 , where i ∈ [1, N].

[0023] In some embodiments, the second feature enhancement module is: a second attention mechanism module or a fifth residual network module for enhancing the feature representation ability.

[0024] Some embodiments of the present disclosure propose an encoding method, including:

[0025] Spatially aligning N input features with different sizes to obtain N first features with the same height and width but different channel numbers, and concatenating the N first features along the channel dimension to form a second feature, where N is greater than or equal to 2;

[0026] Enhancing the representation ability of the second feature to obtain a third feature, where the second feature and the third feature have the same size;

[0027] Reducing the dimension of the third feature along the channel dimension to obtain a fourth feature;

[0028] Quantizing and encoding the fourth feature to obtain an encoded bitstream.

[0029] In some embodiments, the performing spatial alignment includes: downsampling the i-th input feature with a size of (C i , 2 i-1 ×H, 2 i-1 ×W) by a factor of 2 i-1 , where C i represents the channel number of the i-th input feature, i ∈ [1, N], and H and W respectively represent the height and width of the smallest-sized input feature.

[0030] In some embodiments, the downsampling includes:

[0031] Performing downsampling by using a nearest neighbor downsampling module, a bilinear interpolation downsampling module, a pooling downsampling module, a convolutional downsampling module, or a first residual network module.

[0032] Wherein, the first residual network module is composed of the sum of a nearest neighbor downsampling module or a bilinear interpolation downsampling module and a pooling downsampling module or a convolutional downsampling module.

[0033] In some embodiments, enhancing the representation ability of the second feature to obtain a third feature includes: enhancing the representation ability of the second feature to obtain a third feature by using a first attention mechanism module or a second residual network module for enhancing the feature representation ability.

[0034] In some embodiments, reducing the third feature in the channel dimension to obtain a fourth feature includes:

[0035] Reducing the third feature in the channel dimension to obtain a fourth feature by using a first convolutional module or a third residual network module with a stride of one and a number of channels of a first number of channels C', where C i represents the number of channels of the i-th input feature, and C' represents the number of channels after dimensionality reduction during encoding; or

[0036] Reducing the third feature in the channel dimension to obtain a fourth feature by using channel truncation.

[0037] In some embodiments, at least one of the following is used for quantization: linear quantization, clustering quantization, principal component analysis (PCA) quantization, discrete cosine transform (DCT) quantization; and / or at least one of the following is used for encoding: a non-artificial intelligence encoder or an artificial intelligence encoder.

[0038] Some embodiments of the present disclosure propose a decoding method, including:

[0039] Decoding and inverse quantizing the encoded code stream to obtain a fifth feature, where the methods of decoding and inverse quantization are symmetric to the methods of encoding and quantization;

[0040] Increasing the fifth feature in the channel dimension to obtain a sixth feature;

[0041] Restoring the size of the sixth feature to obtain N seventh features, and the sizes of the N seventh features are the same as the sizes of the N input features;

[0042] Enhancing the representation ability of the N seventh features to obtain N eighth features.

[0043] In some embodiments, at least one of the following is used for dequantization: linear dequantization, clustering dequantization, principal component analysis (PCA) dequantization, discrete cosine transform (DCT) dequantization; or / and at least one of the following is used for decoding: a non-artificial intelligence decoder or an artificial intelligence decoder.

[0044] In some embodiments, the process of upsampling the fifth feature along the channel dimension to obtain the sixth feature includes:

[0045] Using a second convolutional module or a fourth residual network module with a stride of one and a number of channels of the second number of channels C″ to upsample the fifth feature along the channel dimension to obtain the sixth feature, where C i represents the number of channels of the i-th input feature, C′ represents the number of channels after dimension reduction during the encoding process, and C″ represents the number of channels after dimension increase during the decoding process; or

[0046] Using channel padding to upsample the fifth feature along the channel dimension to obtain the sixth feature.

[0047] In some embodiments, the process of reducing the size of the sixth feature to obtain N seventh features includes: using N parallel feature size reduction sub-modules to reduce the size of the sixth feature to obtain N seventh features, and the i-th feature size reduction sub-module corresponding to the i-th input feature is implemented by a transposed convolution or a sub-pixel convolution with a number of channels of C i , a stride of 2 i-1 , where i ∈ [1, N].

[0048] In some embodiments, the process of enhancing the representation ability of the N seventh features to obtain N eighth features includes: using a second attention mechanism module or a fifth residual network module for enhancing the feature representation ability to enhance the representation ability of the N seventh features to obtain N eighth features.

[0049] Some embodiments of the present disclosure propose an encoding device, including: a memory; and a processor coupled to the memory, the processor being configured to execute an encoding method based on instructions stored in the memory.

[0050] Some embodiments of the present disclosure propose a decoding device, including: a memory; and a processor coupled to the memory, the processor being configured to execute a decoding method based on instructions stored in the memory.

[0051] Some embodiments of the present disclosure propose an encoding and decoding system, including:

[0052] An encoding device configured to execute an encoding method; and

[0053] A decoding device configured to execute a decoding method.

[0054] Some embodiments of the present disclosure provide a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, an encoding method, and / or a decoding method is implemented. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] The following will briefly introduce the drawings required for use in the embodiments or the related art descriptions. The present disclosure can be more clearly understood according to the following detailed description with reference to the drawings.

[0056] Obviously, the drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0057] Figure 1 Schematic diagram showing an encoding / decoding system according to some embodiments of the present disclosure.

[0058] Figure 2 Schematic diagram showing an encoding / decoding method according to some embodiments of the present disclosure.

[0059] Figure 3 Schematic diagram showing an encoding / decoding scheme according to some embodiments of the present disclosure.

[0060] Figure 4 Schematic diagram showing an encoding / decoding scheme according to some embodiments of the present disclosure.

[0061] Figure 5 Schematic diagram showing an encoding / decoding scheme according to some embodiments of the present disclosure.

[0062] Figure 6 Schematic diagram showing the structure of an encoding / decoding device according to some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0063] It should be noted that: Unless otherwise specifically stated, the relative arrangements, numerical expressions, and values of the components and steps described in these embodiments do not limit the scope of the present disclosure.

[0064] Those skilled in the art can understand that the terms "first", "second", etc. in the embodiments of the present disclosure are only used to distinguish different steps, devices, or modules, etc., and neither represent any specific technical meaning nor indicate an inevitable logical order between them.

[0065] It should also be understood that in the embodiments of the present disclosure, "a plurality of" may refer to two or more, and "at least one" may refer to one, two, or more.

[0066] It should also be understood that for any component, data, or structure mentioned in the embodiments of the present disclosure, in the absence of clear limitations or contrary implications in the context, it is generally understood to be one or more.

[0067] In addition, the term "and / or" in the present disclosure is merely a description of the association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B may represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in the present disclosure generally represents an "or" relationship between the associated objects before and after.

[0068] It should also be understood that the descriptions of the various embodiments in the present disclosure emphasize the differences between the various embodiments, and their similarities or similarities can be referred to each other. For the sake of brevity, they will not be elaborated one by one.

[0069] At the same time, it should be understood that for the sake of description, the dimensions of the various parts shown in the drawings are not drawn according to the actual proportional relationship.

[0070] The following description of at least one exemplary embodiment is actually merely illustrative and in no way restricts the present disclosure or its application or use.

[0071] The technologies, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but in appropriate cases, the said technologies, methods, and devices should be regarded as part of the specification.

[0072] It should be noted that similar reference numerals and letters indicate similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further discussed in subsequent drawings.

[0073] In addition, in order to avoid obscuring the present disclosure with unnecessary details, only the processing steps and / or device structures closely related to at least the solution of the present disclosure are shown in the drawings, while other details less related to the present disclosure are omitted. It should also be noted that similar reference numerals and letters in the drawings indicate similar items, and therefore once an item is defined in one drawing, it does not need to be further discussed for subsequent drawings.

[0074] Figure 1 Schematic diagram of an encoding and decoding system showing some embodiments of the present disclosure.

[0075] As Figure 1 shown, the encoding and decoding system 100 of this embodiment includes: an encoding device 110 configured to perform an encoding method on multi-size input features; and a decoding device 120 configured to perform a decoding method on the encoded bitstream output by the encoding device.

[0076] The encoding device 110 includes: a feature space alignment module 111, a first feature enhancement module 112, a feature channel dimensionality reduction module 113, and a feature encoding module 114.

[0077] The decoding device 120 includes: a feature decoding module 121, a feature channel dimensionality increase module 122, a feature size reduction module 123, and a second feature enhancement module 124.

[0078] In a machine vision task, image data in an image or video is input into a machine vision task neural network, and the multi-size features output by the middle layer of the neural network can be used as the multi-size input features in the embodiments of the present disclosure.

[0079] For the multi-size input features, an encoded bitstream is obtained through feature space alignment, feature enhancement, channel dimensionality reduction, and encoding. For the encoded bitstream, multi-size output features are obtained through decoding, channel dimensionality increase, feature size reduction, and feature enhancement, providing an encoding and decoding scheme applicable to machine vision tasks, enabling the multi-size features in the middle layer of the machine vision task neural network to be encoded, transmitted, and decoded.

[0080] In addition, at the encoding end, feature space alignment is performed first and then channel dimensionality reduction, so that the machine vision task can have better performance.

[0081] In addition, at the decoding end, feature enhancement is performed once again after feature size reduction, so that the machine vision task can have better performance.

[0082] Data compression can usually be achieved through encoding, and data decompression can usually be achieved through decoding.

[0083] The following describes each module at the encoding end and the decoding end one by one.

[0084] The following describes the feature space alignment module 111, the first feature enhancement module 112, the feature channel dimensionality reduction module 113, and the feature encoding module 114 of the encoding device 110.

[0085] The feature space alignment module 111 is configured to perform spatial alignment on N input features with different sizes to obtain N first features with the same width and height but different numbers of channels, and concatenate the N first features along the channel dimension to form a second feature, where N is greater than or equal to 2.

[0086] The feature space alignment module 111 is configured to perform 2 i ,2 i-1 ×H,2 i-1 ×W) of the i-th input feature to perform spatial alignment, where C i-1 times downsampling, i ​Denote the number of channels of the \(i\)-th input feature, which is usually a power of 2, such as 128, 256, 512, etc., but not limited to the examples given. \(i\in[1,N]\), and \(H\) and \(W\) respectively represent the height and width of the smallest-size input feature.

[0087] The feature space alignment module 111 aligns \(N\) input features with sizes \((C1, H, W)\), \((C2, 2\times H, 2\times W)\), \((C3, 4\times H, 4\times W)\), …, \((C N , 2 N-1 \times H, 2 N-1 \times W)\). Since the stride of a neural network is generally 1 or 2, by default, the height and width of all feature layers are \(2\) times the height and width of \(H\) and \(W\) (i.e., the height and width of the smallest-size input feature). After alignment, \(N\) first features with the same height and width but different numbers of channels are obtained. The sizes of the \(N\) first features are respectively \((C1, H, W)\), \((C2, H, W)\), \((C3, H, W)\), …, \((C N , H, W)\). Connect these \(N\) first features along the channel dimension to form a second feature with a size of \((C1 + C2+\cdots + C N , H, W)\).

[0088] Downsampling can be implemented using different methods. Correspondingly, the feature space alignment module 111 can be: a nearest neighbor downsampling module, a bilinear interpolation downsampling module, a pooling downsampling module, a convolutional downsampling module, or a first residual network module. Among them, the nearest neighbor downsampling module is used to implement nearest neighbor downsampling, and the algorithm of nearest neighbor downsampling can refer to related technologies. The bilinear interpolation downsampling module is used to implement bilinear interpolation downsampling, and the algorithm of bilinear interpolation downsampling can refer to related technologies. The pooling downsampling module is used to implement downsampling based on a pooling layer. The convolutional downsampling module is used to implement downsampling based on a convolutional layer. Among them, the first residual network module is composed of the addition of a nearest neighbor downsampling module or a bilinear interpolation downsampling module and a pooling downsampling module or a convolutional downsampling module. The first residual network module can fuse multiple downsampling results, enabling better performance in machine vision tasks.

[0089] If the feature space alignment module 111 is implemented based on the pooling downsampling module, the convolutional downsampling module, or the first residual network module: for the input feature with a size of \((C1, H, W)\), a pooling downsampling module, a convolutional downsampling module, or the first residual network module with a stride of 1 can be used to obtain a first feature with a size of \((C1, H, W)\); for the input feature with a size of \((C2, 2\times H, 2\times W)\), a pooling downsampling module, a convolutional downsampling module, or the first residual network module with a stride of 2 can be used to obtain a first feature with a size of \((C2, H, W)\); for the input feature with a size of \((C i , 2 i-1 \times H, 2i-1 The input feature of (C×W) can adopt a pooling downsampling module, a convolutional downsampling module, or a first residual network module with a stride of 2 i-1 to obtain a first feature with a size of (C2, H, W); for the input feature of (C N , 2 N-1 ×H, 2 N-1 ×W), a pooling downsampling module, a convolutional downsampling module, or a first residual network module with a stride of 2 N-1 can be adopted to obtain a first feature with a size of (C N , H, W).

[0090] The first feature enhancement module 112 is configured to enhance the representation ability of the second feature to obtain a third feature, and the second feature and the third feature have the same size. Thus, the representation ability of the network is strengthened and the compression performance is improved.

[0091] The first feature enhancement module 112 can be composed of any enhancement module that does not change the size of the feature layer. For example, it can be a first attention mechanism module, a second residual network module, or a convolutional module for enhancing the feature representation ability. The first attention mechanism module is used to implement the attention mechanism of the neural network. The second residual network module is used to implement the residual network function. The convolutional module is used to implement the convolutional network function. The stride of the first attention mechanism module, the second residual network module, or the convolutional module can be set to one.

[0092] The feature channel dimensionality reduction module 113 is configured to reduce the dimensionality of the third feature along the channel dimension to obtain a fourth feature.

[0093] The feature channel dimensionality reduction module 113 is composed of a first convolutional module or a third residual network module with a stride of one and a channel number of the first channel number C′, or is composed of a channel truncation module, where C i represents the channel number of the i-th input feature, and C′ represents the channel number after dimensionality reduction during the encoding process. Thus, the third feature with a size of (C1 + C2 + … + C N , H, W) is reduced in dimensionality along the channel dimension to obtain a fourth feature with a size of (C′, H, W). C′ is usually a very small number, such as 64, 128, etc.

[0094] Among them, the first convolutional module or the third residual network module is a parameterized dimensionality reduction method. The first convolutional module is used to implement the convolutional network function, and the third residual network module is used to implement the residual network function. The channel truncation module is a non-parameterized dimensionality reduction method, and the channel truncation module is used to achieve dimensionality reduction through channel truncation.

[0095] The feature encoding module 114 is configured to quantize and encode the fourth feature to obtain an encoded bitstream. The encoded bitstream can be stored and transmitted.

[0096] The feature encoding module 114 is configured to perform quantization using at least one of the following: linear quantization, clustering quantization, principal component analysis (PCA) quantization, discrete cosine transform (DCT) quantization. After quantization, rounding operations such as rounding up or rounding to the nearest integer can also be performed.

[0097] The feature encoding module 114 is configured to perform encoding using at least one of the following: a non-artificial intelligence encoder or an artificial intelligence (AI) encoder. Non-artificial intelligence encoders are, for example, traditional encoders such as H.264 / H.265 / H.266. AI encoders are, for example, encoders such as Cheng2020, Scale-Space-Flow, or end-to-end encoders composed of a neural network and an entropy encoder.

[0098] The following describes the feature decoding module 121, the feature channel dimension increasing module 122, the feature size reduction module 123, and the second feature enhancement module 124 of the decoding device 120.

[0099] The feature decoding module 121 is configured to decode and inverse-quantize the encoded bitstream to obtain a fifth feature. Among them, the methods of decoding and inverse-quantization are symmetric to the methods of encoding and quantization. The size of the fifth feature is the same as that of the fourth feature, both being (C′, H, W).

[0100] The feature decoding module 121 is configured to perform inverse-quantization using at least one of the following: linear inverse-quantization, clustering inverse-quantization, principal component analysis PCA inverse-quantization, discrete cosine transform DCT inverse-quantization.

[0101] The feature decoding module 121 is configured to perform decoding using at least one of the following: a non-artificial intelligence decoder or an artificial intelligence decoder. Non-artificial intelligence decoders are, for example, traditional decoders such as H.264 / H.265 / H.266. AI decoders are, for example, decoders such as Cheng2020, Scale-Space-Flow, or end-to-end decoders composed of a neural network and an entropy decoder.

[0102] The feature channel dimension increasing module 122 is configured to increase the dimension of the fifth feature along the channel dimension to obtain a sixth feature. That is, the fifth feature with a size of (C′, H, W) is dimension-increased to obtain a sixth feature with a size of (C″, H, W).

[0103] The feature channel upsampling module 122 is composed of a second convolutional module or a fourth residual network module with a stride of one and a channel number of the second channel number C″, or is composed of a channel padding module, where C i represents the channel number of the i-th input feature, C′ represents the channel number after dimensionality reduction during the encoding process, and C″ represents the channel number after dimensionality increase during the decoding process.

[0104] Among them, the second convolutional module is used to implement the convolutional network function, the fourth residual network module is used to implement the residual network function, and the channel padding module is used to implement the channel padding function.

[0105] The feature size reduction module 123 is configured to reduce the size of the sixth feature to obtain N seventh features, and the sizes of the N seventh features are the same as those of the N input features, all being (C1, H, W), (C2, 2×H, 2×W), (C3, 4×H, 4×W), …, (C N , 2 N-1 ×H, 2 N-1 ×W).

[0106] The size reduction can be performed in parallel, enabling better performance in machine vision tasks. That is, the feature size reduction module 123 is composed of N parallel feature size reduction sub-modules. The sixth feature is used as the input of each feature size reduction sub-module. The i-th feature size reduction sub-module corresponding to the i-th input feature is implemented by a transposed convolution or sub-pixel convolution with a channel number of C i , a stride of 2 i-1 , where i ∈ [1, N]. Thus, the feature size reduction sub-module can adjust the channel number and the feature height and width simultaneously.

[0107] The second feature enhancement module 124 is configured to enhance the representation ability of the N seventh features to obtain N eighth features. The sizes of the seventh features of each branch are the same as those of the eighth features of that branch. Thus, the representation ability of the network is strengthened and the compression performance is improved.

[0108] The second feature enhancement module 124 can be composed of any enhancement module that does not change the size of the feature layer. For example, it can be a second attention mechanism module or a fifth residual network module for enhancing the feature representation ability, or a convolutional module. The second attention mechanism module is used to implement the attention mechanism of the neural network. The fifth residual network module is used to implement the residual network function. The convolutional module is used to implement the convolutional network function. The stride of the second attention mechanism module or the fifth residual network module or the convolutional module can be set to one.

[0109] Figure 2 A schematic diagram showing the encoding and decoding method of some embodiments of the present disclosure.

[0110] Such asFigure 2 As shown, the encoding method in this embodiment includes steps S210 - S240, and the corresponding decoding method includes steps S250 - S280.

[0111] In step S210, N input features with different sizes are spatially aligned to obtain N first features with the same width and height but different numbers of channels, and the N first features are concatenated along the channel dimension to form a second feature, where N is greater than or equal to 2.

[0112] Among them, the spatial alignment includes: downsampling the i-th input feature with a size of (C i , 2 i-1 ×H, 2 i-1 ×W) by a factor of 2, where C i-1 represents the number of channels of the i-th input feature, i ∈ [1, N], and H and W respectively represent the height and width of the input feature with the smallest size. i

[0113] Among them, the downsampling includes: using a nearest neighbor downsampling module, a bilinear interpolation downsampling module, a pooling downsampling module, a convolutional downsampling module, or a first residual network module for downsampling, where the first residual network module is composed of the addition of a nearest neighbor downsampling module or a bilinear interpolation downsampling module and a pooling downsampling module or a convolutional downsampling module.

[0114] In step S220, the representational ability of the second feature is enhanced to obtain a third feature, and the second feature and the third feature have the same size.

[0115] In some embodiments, the representational ability of the second feature is enhanced to obtain a third feature by using a first attention mechanism module or a second residual network module for enhancing the feature representational ability.

[0116] In step S230, the third feature is dimensionally reduced along the channel dimension to obtain a fourth feature.

[0117] In some embodiments, the third feature is dimensionally reduced along the channel dimension to obtain a fourth feature by using a first convolutional module with a stride of one and a number of channels of the first channel number C′ or a third residual network module, where C i represents the number of channels of the i-th input feature, and C′ represents the number of channels after dimensional reduction in the encoding process.

[0118] In some embodiments, the third feature is dimensionally reduced along the channel dimension to obtain a fourth feature by using channel truncation.

[0119] In step S240, the fourth feature is quantized and encoded to obtain an encoded bitstream.

[0120] Among them, at least one of the following is adopted for quantization: linear quantization, clustering quantization, principal component analysis (PCA) quantization, discrete cosine transform (DCT) quantization.

[0121] Among them, at least one of the following is adopted for encoding: non-artificial intelligence encoder or artificial intelligence encoder.

[0122] In step S250, the encoded bitstream is decoded and inverse-quantized to obtain the fifth feature, where the methods of decoding and inverse-quantization are symmetric to the methods of encoding and quantization.

[0123] At least one of the following is adopted for inverse-quantization: linear inverse-quantization, clustering inverse-quantization, principal component analysis (PCA) inverse-quantization, discrete cosine transform (DCT) inverse-quantization.

[0124] At least one of the following is adopted for decoding: non-artificial intelligence decoder or artificial intelligence decoder.

[0125] In step S260, the fifth feature is dimensionally upsampled along the channel dimension to obtain the sixth feature.

[0126] In some embodiments, a second convolutional module or a fourth residual network module with a stride of one and a number of channels equal to the second number of channels C″ is used to dimensionally upsample the fifth feature along the channel dimension to obtain the sixth feature, where C i represents the number of channels of the i-th input feature, C′ represents the number of channels after dimensionality reduction during the encoding process, and C″ represents the number of channels after dimensionality upsampling during the decoding process.

[0127] In some embodiments, channel padding is used to dimensionally upsample the fifth feature along the channel dimension to obtain the sixth feature.

[0128] In step S270, the sixth feature is size-reduced to obtain N seventh features, and the sizes of the N seventh features are the same as the sizes of the N input features.

[0129] Using N parallel feature size reduction sub-modules, the sixth feature is size-reduced to obtain N seventh features, and the i-th feature size reduction sub-module corresponding to the i-th input feature consists of a transposed convolution or sub-pixel convolution with a number of channels of C i , a stride of 2 i-1 to implement, where i ∈ [1, N].

[0130] In step S280, the representational ability of the N seventh features is enhanced to obtain N eighth features.

[0131] Using a second attention mechanism module or a fifth residual network module for enhancing the representational ability of features, the representational ability of the N seventh features is enhanced to obtain N eighth features.

[0132] For multi - size input features, an encoded bitstream is obtained through feature space alignment, feature enhancement, channel dimensionality reduction, and encoding. For the encoded bitstream, multi - size output features are obtained through decoding, channel dimensionality increase, feature size restoration, and feature enhancement, providing an encoding - decoding scheme suitable for machine vision tasks, enabling the encoding, transmission, and decoding of multi - size features in the intermediate layer of a machine vision task neural network.

[0133] In addition, at the encoding end, feature space alignment is performed first and then channel dimensionality reduction is carried out, enabling better performance in machine vision tasks.

[0134] In addition, at the decoding end, feature enhancement is performed once again after feature size restoration, enabling better performance in machine vision tasks.

[0135] Figure 3 A schematic diagram showing the encoding - decoding scheme of some embodiments of the present disclosure is presented. The following combines Figure 3 to describe the encoding - decoding scheme of this embodiment, where the input and output are composed of, for example, three features with different sizes.

[0136] (1) The feature space alignment module consists of a residual network module composed of a convolutional module and a downsampling module. Among them, the network complexity of the residual network module becomes more complex as the feature size increases. Specifically, the p3 branch is composed of a 1x1 convolutional module with a stride of 1 (labeled con1×1), and the size of its input feature is (4C, H, W), and the size of the output first feature is (4C, H, W). The p2 branch is a residual network module composed of a 3x3 convolutional module with a stride of 2 (labeled con3×3, stride 2) and a nearest - neighbor downsampling module (downsampling ratio of 2, labeled Downsample 2x), and the size of its input feature is (2C, 2H, 2W), and the size of the output first feature is (2C, H, W). The p1 branch is a residual network module composed of two 3x3 convolutional modules with a stride of 2 (labeled con3×3, stride 2 respectively) and a nearest - neighbor downsampling module (downsampling ratio of 4, labeled Downsample 4x), and the size of its input feature is (C, 4H, 4W), and the size of the output first feature is (C, H, W). The three first features with sizes of (4C, H, W), (2C, H, W), and (C, H, W) are concatenated along the channel dimension to form a second feature with a size of (7C, H, W).

[0137] (2) The first feature enhancement module at the encoding end uses a lightweight attention mechanism, the ECA (Efficient Channel Attention) module, to enhance the features. The ECA module enhances the representation ability of the second feature with a size of (7C, H, W) to obtain the third feature with a size of (7C, H, W).

[0138] (3) The feature channel dimensionality reduction module is a convolutional network with a 1x1 convolutional kernel (labeled con1×1), which reduces the dimensionality of the third feature with a size of (7C, H, W) to obtain the fourth feature with a size of (64, H, W).

[0139] (4) The feature encoding module consists of 10-bit linear quantization (labeled 10-bit linear quantization) and a VTM (a coding tool) encoder (labeled Encode using VTM), which quantizes and encodes the fourth feature to obtain the encoded bitstream.

[0140] (5) The feature decoding module consists of a VTM decoder (labeled Decode using VTM) and 10-bit linear dequantization (labeled 10-bit linear dequantization), which decodes and dequantizes the encoded bitstream to obtain the fifth feature with a size of (64, H, W).

[0141] (6) The feature channel dimensionality increase module corresponds to the encoding end and is a convolutional network with a 1x1 convolutional kernel (labeled con1×1), which increases the dimensionality of the fifth feature with a size of (64, H, w) to obtain the sixth feature with a size of (4C, H, w).

[0142] (7) The feature size restoration module uses a sub-pixel convolution with a 1x1 convolutional kernel to simultaneously restore the sixth feature with a size of (4C, H, W) to the original size of each input feature, obtaining seventh features with sizes of (4C, H, W), (2C, 2H, 2W), and (C, 4H, 4W) respectively.

[0143] Among them, the sizes of the features before and after restoration of the P3 branch are the same, both (4C, H, W), which can be regarded as a critical case of the feature size restoration module. When the smallest feature size is consistent with the output size of the feature channel dimensionality increase module, the feature size restoration module can be composed of an identity function (because theoretically, the identity function also restores the channel and width-height dimensions of the feature simultaneously).

[0144] Among them, the upsampling ratio of the sub-pixel convolution with a 1x1 convolutional kernel in the P2 branch is 2 (labeled Sub-pixel conv1x1 2x Upsample).

[0145] Among them, the upsampling magnification of the sub-pixel convolution with a convolution kernel size of 1x1 in the P1 branch is 4 (marked as Sub-pixel conv1x1 4x Upsample).

[0146] (8) The second feature enhancement module at the decoding end consists of a residual network module composed of max pooling and convolution, without changing the size of the feature layer. That is, the representational capabilities of three seventh features are enhanced to obtain three eighth features, and the sizes of the seventh features and the eighth features of each branch are the same. The network complexity of the residual network module becomes more complex as the feature size increases. Specifically, the p3 branch is a residual network module composed of a convolution module with a convolution kernel of 1x1 and a stride of 1 (marked as con1×1) and a max pooling module with a convolution kernel of 3x3 (marked as Max pool 3x3), and the p2 and p1 branches are residual network modules composed of a convolution module with a convolution kernel of 3x3 and a stride of 1 (marked as con3x3) and a max pooling module with a convolution kernel of 3x3 (marked as Max pool 3x3). This module can effectively improve the encoding and task performance.

[0147] Figure 4 A schematic diagram showing the encoding and decoding scheme of some embodiments of the present disclosure. The following combines Figure 4 to describe the encoding and decoding scheme of this embodiment, where the input and output are composed of four features with different sizes, for example.

[0148] (1) The feature space alignment module consists of a residual network module composed of a convolutional module and a downsampling module. Among them, the network complexity of the residual network module becomes more complex as the feature size increases. Specifically, the p5 branch consists of a 1x1 convolutional module with a stride of 1 (labeled con1×1), the size of its input feature is (8C, H, W), and the size of the output first feature is (8C, H, W). The p4 branch is a residual network module composed of a 3x3 convolutional module with a stride of 2 (labeled con3×3, stride 2) and a nearest neighbor downsampling module (downsampling ratio of 2, labeled Downsample 2x), the size of its input feature is (4C, 2H, 2W), and the size of the output first feature is (4C, H, W). The p3 branch is a residual network module composed of two 3x3 convolutional modules with a stride of 2 (labeled con3×3, stride 2 respectively) and a nearest neighbor downsampling module (downsampling ratio of 4, labeled Downsample 4x), the size of its input feature is (2C, 4H, 4W), and the size of the output first feature is (2C, H, W). The p2 branch is a residual network module composed of three 3x3 convolutional modules with a stride of 2 (labeled con3×3, stride 2 respectively) and a nearest neighbor downsampling module (downsampling ratio of 8, labeled Downsample8x), the size of its input feature is (C, 8H, 8W), and the size of the output first feature is (C, H, W). The four first features with sizes of (8C, H, W), (4C, H, W), (2C, H, W), and (C, H, W) are concatenated along the channel dimension to form a second feature with a size of (15C, H, W).

[0149] (2) The first feature enhancement module at the encoding end uses a lightweight attention mechanism, the ECA module, to enhance the features here. The ECA module enhances the representation ability of the second feature with a size of (15C, H, W) to obtain a third feature with a size of (15C, H, W).

[0150] (3) The feature channel dimensionality reduction module is a convolutional network with a 1x1 convolutional kernel (labeled con1×1), which reduces the dimensionality of the third feature with a size of (15C, H, W) to obtain a fourth feature with a size of (64, H, W).

[0151] (4) The feature encoding module consists of 10-bit linear quantization (labeled 10-bit linear quantization) and a VTM (a coding tool) encoder (labeled Encode using VTM), which quantizes and encodes the fourth feature to obtain an encoded bitstream.

[0152] (5) The feature decoding module consists of a VTM decoder (labeled as Decode using VTM) and 10-bit linear dequantization (labeled as 10-bit linear dequantization), which decodes and dequantizes the encoded bitstream to obtain the fifth feature with a size of (64, H, W).

[0153] (6) The feature channel dimension increasing module corresponds to the encoding end and is a convolutional network with a 1x1 convolutional kernel (labeled as con1×1), which increases the dimension of the fifth feature with a size of (64, H, w) to obtain the sixth feature with a size of (8C, H, w).

[0154] (7) The feature size reduction module uses sub-pixel convolution with a 1x1 convolutional kernel to simultaneously reduce the sixth feature with a size of (8C, H, W) to the original size of each input feature, obtaining seventh features with sizes of (8C, H, W), (4C, 2H, 2W), (2C, 4H, 4W), and (C, 8H, 8W) respectively.

[0155] Among them, the sizes of the features before and after reduction in the P5 branch are the same, both (8C, H, W), which can be regarded as a critical case of the feature size reduction module. When the smallest feature size is the same as the output size of the feature channel dimension increasing module, the feature size reduction module can be composed of an identity function (because theoretically the identity function also reduces the channel and width-height dimensions of the feature simultaneously).

[0156] Among them, the upsampling ratio of the sub-pixel convolution with a 1x1 convolutional kernel in the P4 branch is 2 (labeled as Sub-pixel conv1x1 2x Upsample).

[0157] Among them, the upsampling ratio of the sub-pixel convolution with a 1x1 convolutional kernel in the P3 branch is 4 (labeled as Sub-pixel conv1x1 4x Upsample).

[0158] Among them, the upsampling ratio of the sub-pixel convolution with a 1x1 convolutional kernel in the P2 branch is 8 (labeled as Sub-pixel conv1x1 8x Upsample).

[0159] (8) The second feature enhancement module at the decoding end consists of a residual network module composed of a max-pooling and a convolution, without changing the size of the feature layer. That is, the representational capabilities of the four seventh features are enhanced to obtain four eighth features, and the seventh features of each branch have the same size as the eighth features of that branch. The network complexity of the residual network module becomes more complex as the feature size increases. Specifically, the p5 branch is a residual network module composed of a 1x1 convolutional module with a stride of one (labeled con1×1) and a 3x3 max-pooling module (labeled Max pool 3x3). The p4 and p3 branches are residual network modules composed of a 3x3 convolutional module with a stride of one (labeled con3x3) and a 3x3 max-pooling module (labeled Max pool 3x3). The P2 branch is a residual network module composed of two 3x3 convolutional modules with a stride of one (labeled con3x3) and a 3x3 max-pooling module (labeled Max pool 3x3). This module can effectively improve the encoding and task performance.

[0160] Figure 5 Schematic diagram showing the encoding and decoding schemes of some embodiments of the present disclosure. The following will be combined with Figure 5 Describe the encoding and decoding schemes of this embodiment, where the input and output are composed of, for example, three features with different sizes.

[0161] (1) The feature space alignment module consists of a convolutional module or a max-pooling module. Specifically, the p3 branch is composed of a 1x1 convolutional module (labeled con1×1), with the size of the input feature being (4C, H, W) and the size of the output first feature being (4C, H, W). The p2 branch is composed of a max-pooling module with a stride of two (labeled max-poling, stride2), with the size of the input feature being (2C, 2H, 2W) and the size of the output first feature being (2C, H, W). The p1 branch is composed of a max-pooling module with a stride of 4 (labeled max-poling, stride 4), with the size of the input feature being (C, 4H, 4W) and the size of the output first feature being (C, H, W). The three first features with sizes of (4C, H, W), (2C, H, W), and (C, H, W) are concatenated along the channel dimension to form a second feature with a size of (7C, H, W).

[0162] (2) The first feature enhancement module at the encoding end uses an attention mechanism SE (Squeeze-and-Excitation) network module here to enhance the features. The SE network module enhances the representational capabilities of the second feature with a size of (7C, H, W) to obtain a third feature with a size of (7C, H, W).

[0163] (3) The feature channel dimensionality reduction module is a convolutional network with a 1x1 convolutional kernel (labeled con1×1), which reduces the third feature of size (7C, H, W) to obtain a fourth feature of size (64, H, W).

[0164] (4) The feature encoding module consists of k-means clustering quantization (labeled k-means quantization) and an AI encoder (labeled AI Encoder), which quantizes and encodes the fourth feature to obtain an encoded bitstream.

[0165] (5) The feature decoding module consists of an AI decoder (labeled AIDecoder) and k-means clustering dequantization (labeled k-means dequantization), which decodes and dequantizes the encoded bitstream to obtain a fifth feature of size (64, H, W).

[0166] (6) The feature channel dimensionality increase module corresponds to the encoding end and is a convolutional network with a 3x3 convolutional kernel (labeled con3x3), which increases the dimensionality of the fifth feature of size (64, H, W) to obtain a sixth feature of size (4C, H, w).

[0167] (7) The feature size restoration module is a transposed convolution module and a sub-pixel convolution module respectively, which restores the sixth feature of size (4C, H, W) to the original size of each input feature at the same time, obtaining seventh features of sizes (4C, H, W), (2C, 2H, 2W), and (C, 4H, 4W) respectively. Specifically:

[0168] Among them, the sizes of the features before and after restoration of the P3 branch are the same, both (4C, H, w), which can be regarded as a critical case of the feature size restoration module. When the smallest feature size is consistent with the output size of the feature channel dimensionality increase module, the feature size restoration module can be composed of an identity function (because theoretically the identity function also restores the channel and width-height sizes of the feature at the same time).

[0169] Among them, the upsampling ratio of the transposed convolution with a 1x1 kernel in the P2 branch is 2 (labeled deconv1x1 2xUpsample).

[0170] Among them, the upsampling ratio of the sub-pixel convolution with a 1x1 kernel in the P1 branch is 4 (labeled Sub-pixelconv1x1 4x Upsample).

[0171] (8) The second feature enhancement module at the decoding end is composed of a convolutional module with a kernel size of 3x3 (labeled as con3x3), which does not change the size of the feature layer. That is, the representational capabilities of 3 seventh features are enhanced to obtain 3 eighth features, and the seventh features of each branch have the same size as the eighth features of that branch.

[0172] Figure 6 The structural schematic diagram of an encoding and decoding device according to some embodiments of the present disclosure is shown. The encoding and decoding device is: an encoding device or / and a decoding device.

[0173] As Figure 6 shown, the encoding and decoding device 600 of this embodiment includes: a memory 610 and a processor 620 coupled to the memory 610. The processor 620 is configured to execute the encoding method or / and the decoding method of any of some embodiments based on instructions stored in the memory 610.

[0174] The encoding and decoding device 600 may further include an input / output interface 630, a network interface 640, a storage interface 650, etc. These interfaces 630, 640, 650 and the memory 610 and the processor 620 may be connected through a bus 660, for example.

[0175] Among them, the memory 610 may include a system memory, a fixed non-volatile storage medium, etc., for example. The system memory stores an operating system, application programs, a boot loader, and other programs, for example.

[0176] Among them, the processor 620 may be implemented in the form of a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete hardware components such as discrete gates or transistors.

[0177] Among them, the input / output interface 630 provides connection interfaces for input / output devices such as monitors, mice, keyboards, and touchscreens. The network interface 640 provides connection interfaces for various networking devices. The storage interface 650 provides connection interfaces for external storage devices such as SD cards and USB flash drives. The bus 660 can use any bus structure among a variety of bus structures. For example, the bus structure includes, but is not limited to, the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MCA) bus, and the Peripheral Component Interconnect (PCI) bus.

[0178] Embodiments of the present disclosure provide a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements an encoding method, and / or a decoding method. The storage medium is, for example, a non-transitory storage medium.

[0179] Those skilled in the art should understand that the embodiments of the present disclosure may be provided as a method, a system, or a computer program product. Therefore, the present disclosure may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present disclosure may take the form of a computer program product implemented on one or more non-transitory computer-readable storage media (including, but not limited to, disk memories, CD-ROMs, optical memories, etc.) that contain computer program code.

[0180] The present disclosure is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0181] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing devices to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one or more flows and / or blocks Figure 1The functions specified in one or more boxes.

[0182] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide for implementing the steps of the functions specified in Figure 1 One process or more processes and / or boxes Figure 1 The steps of the functions specified in one or more boxes.

[0183] The above are only the preferred embodiments of the present disclosure and are not intended to limit the present disclosure. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. An encoding device, comprising: A feature space alignment module, configured to perform spatial alignment on N input features with different sizes to obtain N first features with different numbers of channels but the same width and height, and concatenate the N first features along the channel dimension to form a second feature, where N is greater than or equal to 2; A first feature enhancement module, configured to enhance the representation ability of the second feature to obtain a third feature, where the second feature and the third feature have the same size; A feature channel dimensionality reduction module, configured to reduce the dimensionality of the third feature along the channel dimension to obtain a fourth feature; A feature encoding module, configured to quantize and encode the fourth feature to obtain an encoded bitstream.

2. The encoding device according to claim 1, wherein, The feature space alignment module is configured to transform the size of (C i ,2 i-1 ×H,2 i-1 ×W) is used to perform 2 i-1 times downsampling, where C i represents the number of channels of the i-th input feature, i∈[1,N], H and W represent the height and width of the minimum size input feature, respectively.

3. The encoding device according to claim 2, wherein, The feature space alignment module is: a nearest neighbor downsampling module, a bilinear interpolation downsampling module, a pooling downsampling module, a convolutional downsampling module, or a first residual network module, wherein the first residual network module is composed of the addition of a nearest neighbor downsampling module or a bilinear interpolation downsampling module and a pooling downsampling module or a convolutional downsampling module.

4. The encoding device according to any one of claims 1-3, wherein, The first feature enhancement module is: a first attention mechanism module or a second residual network module for enhancing the feature representation ability.

5. The encoding device according to any one of claims 1-4, wherein, The feature channel dimensionality reduction module is composed of a first convolution module with a stride of one and a number of channels being the first number of channels C′, or a third residual network module, or is composed of a channel truncation module, where C i represents the number of channels of the i-th input feature, and C′ represents the number of channels after dimensionality reduction during the encoding process.

6. The encoding device according to any one of claims 1-5, wherein, The feature encoding module is configured to perform quantization using at least one of the following: linear quantization, clustering quantization, principal component analysis (PCA) quantization, discrete cosine transform (DCT) quantization; or / and The feature encoding module is configured to perform encoding using at least one of the following: a non-artificial intelligence encoder or an artificial intelligence encoder.

7. A decoding device, comprising: A feature decoding module, configured to decode and inverse-quantize the encoded bitstream to obtain a fifth feature, wherein the methods of decoding and inverse-quantization are symmetric to the methods of encoding and quantization; A feature channel dimensionality increase module, configured to increase the dimensionality of the fifth feature along the channel dimension to obtain a sixth feature; A feature size restoration module, configured to restore the size of the sixth feature to obtain N seventh features, where the sizes of the N seventh features are the same as the sizes of the N input features; A second feature enhancement module, configured to enhance the representation ability of the N seventh features to obtain N eighth features.

8. The decoding device according to claim 7, Wherein, The feature decoding module is configured to perform inverse-quantization using at least one of the following: linear inverse-quantization, clustering inverse-quantization, principal component analysis (PCA) inverse-quantization, discrete cosine transform (DCT) inverse-quantization; or / and The feature decoding module is configured to perform decoding using at least one of the following: a non-artificial intelligence decoder or an artificial intelligence decoder.

9. The decoding device according to any one of claims 7-8, wherein, The feature channel dimension increasing module is composed of a second convolutional module with a stride of one and a number of channels of a second number of channels C″, or a fourth residual network module, or is composed of a channel padding module, where C i represents the number of channels of the i-th input feature, C′ represents the number of channels after dimension reduction during the encoding process, and C″ represents the number of channels after dimension increase during the decoding process.

10. The decoding device according to any one of claims 7-9, wherein, The feature size reduction module is composed of N parallel feature size reduction sub-modules. The i-th feature size reduction sub-module corresponding to the i-th input feature is implemented by deconvolution or sub-pixel convolution with a channel number of C i , a stride of 2 i-1 , where i ∈ [1, N].

11. The decoding device according to any one of claims 7-10, wherein, The second feature enhancement module is: a second attention mechanism module or a fifth residual network module for enhancing the feature representation ability.

12. An encoding method, comprising: Performing spatial alignment on N input features with different sizes to obtain N first features with different numbers of channels but the same width and height, and concatenating the N first features along the channel dimension to form a second feature, where N is greater than or equal to 2; Enhancing the representation ability of the second feature to obtain a third feature, where the second feature and the third feature have the same size; Reducing the dimensionality of the third feature along the channel dimension to obtain a fourth feature; Quantizing and encoding the fourth feature to obtain an encoded bitstream.

13. The encoding method according to claim 12, wherein, The spatial alignment includes: Downsample the i-th input feature with dimensions (C i , 2 i-1 ×H, 2 i-1 ×W) by a factor of 2 i-1 , where C i represents the number of channels of the i-th input feature, i ∈ [1, N], and H and W represent the height and width of the smallest-sized input feature, respectively.

14. The encoding method according to claim 13, wherein, The downsampling includes: Performing downsampling using a nearest neighbor downsampling module, a bilinear interpolation downsampling module, a pooling downsampling module, a convolutional downsampling module, or a first residual network module. Wherein, the first residual network module is constituted by adding a nearest neighbor downsampling module or a bilinear interpolation downsampling module to a pooling downsampling module or a convolutional downsampling module.

15. The encoding method according to any one of claims 12-14, wherein, The enhancement of the representation ability of the second feature to obtain the third feature includes: Enhancing the representation ability of the second feature to obtain the third feature using a first attention mechanism module or a second residual network module for enhancing the feature representation ability.

16. The encoding method according to any one of claims 12-15, wherein, The reduction of the third feature along the channel dimension to obtain the fourth feature includes: Using a first convolutional module or a third residual network module with a step size of one and a number of channels being the first number of channels C′, the third feature is dimensionally reduced along the channel dimension to obtain a fourth feature, where C i represents the number of channels of the i-th input feature, and C′ represents the number of channels after dimensional reduction during the encoding process; or Using channel truncation to reduce the third feature along the channel dimension to obtain the fourth feature.

17. The encoding method according to any one of claims 12-16, wherein Performing quantization using at least one of the following: linear quantization, clustering quantization, principal component analysis (PCA) quantization, discrete cosine transform (DCT) quantization; or / and Performing encoding using at least one of the following: a non-artificial intelligence encoder or an artificial intelligence encoder.

18. A decoding method, comprising: Decoding and inverse quantizing the encoded bitstream to obtain a fifth feature, wherein the methods of decoding and inverse quantization are symmetric to the methods of encoding and quantization; Upsampling the fifth feature along the channel dimension to obtain a sixth feature; Restoring the size of the sixth feature to obtain N seventh features, and the sizes of the N seventh features are the same as those of the N input features; Enhancing the representation ability of the N seventh features to obtain N eighth features.

19. The decoding method according to claim 18, wherein Performing inverse quantization using at least one of the following: linear inverse quantization, clustering inverse quantization, principal component analysis (PCA) inverse quantization, discrete cosine transform (DCT) inverse quantization; or / and Performing decoding using at least one of the following: a non-artificial intelligence decoder or an artificial intelligence decoder.

20. The decoding method according to any one of claims 18-19, wherein, The upsampling of the fifth feature along the channel dimension to obtain the sixth feature includes: Using a second convolutional module or a fourth residual network module with a step size of one and the number of channels being the second number of channels C″, the fifth feature is dimensionally elevated along the channel dimension to obtain a sixth feature, where C i represents the number of channels of the i-th input feature, C′ represents the number of channels after dimensionality reduction during the encoding process, and C″ represents the number of channels after dimensionality elevation during the decoding process; or Using channel padding to upsample the fifth feature along the channel dimension to obtain the sixth feature.

21. The decoding method according to any one of claims 18-20, wherein, The restoring the size of the sixth feature to obtain N seventh features includes: Using N parallel feature size reduction sub-modules to reduce the size of the sixth feature to obtain N seventh features. The i-th feature size reduction sub-module corresponding to the i-th input feature is implemented by a transposed convolution or sub-pixel convolution with a channel number of C i , a stride of 2 i-1 , where i ∈ [1, N].

22. The decoding method according to any one of claims 18-21, wherein, The enhancing the representation ability of the N seventh features to obtain N eighth features includes: Enhancing the representation ability of the N seventh features to obtain N eighth features using a second attention mechanism module or a fifth residual network module for enhancing the feature representation ability.

23. An encoding device, comprising: A memory; And A processor coupled to the memory, the processor being configured to execute the encoding method according to any one of claims 12-17 based on instructions stored in the memory.

24. A decoding device, comprising: A memory; And A processor coupled to the memory, the processor being configured to execute the decoding method according to any one of claims 18-22 based on instructions stored in the memory.

25. A codec system, comprising: An encoding device configured to execute the encoding method according to any one of claims 12-17; And A decoding device configured to perform the decoding method according to any one of claims 18-22.

26. A computer-readable storage medium storing a computer program which, when executed by a processor, implements the encoding method according to any one of claims 12-17, and / or the decoding method according to any one of claims 18-22.