Image decoding and encoding method, apparatus, device, and storage medium

By extracting and grouping residual data for spatial domain resolution reduction and expansion, the method addresses high time complexity in image coding and decoding, improving efficiency and reducing computational demands.

JP2026511884APending Publication Date: 2026-04-14HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
Filing Date
2024-03-29
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

The high time complexity in the average value prediction process of conventional image coding and decoding is a significant challenge, particularly as the resolution of features increases, leading to inefficient execution.

Method used

The method involves extracting image residual data, performing spatial domain resolution reduction and expansion processes, grouping residual data into multiple groups, and conducting residual reconstruction to reduce time complexity.

Benefits of technology

This approach improves residual reconstruction computation efficiency and reduces time complexity by processing residual data in a grouped manner, enhancing overall image decoding and encoding performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026511884000001_ABST
    Figure 2026511884000001_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of image processing technology and discloses an image decoding and encoding method, apparatus, device, and storage medium. The present invention extracts image residual data or extended residual data from an image bitstream, obtains a plurality of extended residual groups based on the extracted image residual data or extended residual data, performs residual reconstruction for each of the plurality of extended residual groups, obtains image reconstruction features corresponding to each extended residual group, performs spatial domain resolution expansion processing on the image reconstruction features corresponding to each extended residual group, obtains reconstruction feature data, performs image reconstruction based on the reconstruction feature data, and obtains a reconstructed image block. Since the obtained extended residual data is residual data that has undergone spatial domain resolution reduction processing, it can be grouped at a low resolution and residual reconstruction processing can be performed on a group basis, improving the overall residual reconstruction computation efficiency and reducing time complexity.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to the field of image processing technology, and more particularly to image decoding and encoding methods, apparatus, devices, and storage media. [Background technology]

[0002] In deep learning-based image compression, the current mainstream method uses decoded feature points as prior information to predict the mean value for the currently decoded feature points and reduces the spatial redundancy of the image. However, mainstream methods typically employ serial or wavefront encoding and decoding schemes, and as the resolution of the features increases, the degree of seriality increases, and the overall time complexity of execution increases.

[0003] The above information is merely an aid to understanding the technical proposal of the present invention and does not imply that the above information constitutes prior art. [Overview of the project]

[0004] The main objective of the present invention is to provide an image decoding and encoding method, apparatus, device, and storage medium in order to solve the technical problem of high time complexity in the average value prediction process in the conventional image coding and decoding process.

[0005] To achieve the above objective, the present invention The steps include: extracting image residual data or extended residual data from an image bitstream, and obtaining a plurality of extended residual groups based on the extracted image residual data or extended residual data; The steps include performing residual reconstruction for each of the aforementioned multiple extended residual groups and obtaining image reconstruction features corresponding to each extended residual group, The steps include performing spatial domain resolution expansion processing on the image reconstruction features corresponding to each expanded residual group and obtaining the reconstruction feature data, The present invention provides an image decoding method that includes the steps of: performing image reconstruction based on the aforementioned reconstruction feature data and obtaining a reconstructed image block.

[0006] In one possible embodiment of the present invention, the steps of extracting image residual data or extended residual data from the image bitstream and obtaining a plurality of extended residual groups based on the extracted image residual data or extended residual data are: The steps include extracting the image residual data from the image bitstream, A step of performing a spatial domain resolution reduction process on the aforementioned image residual data and obtaining the aforementioned expanded residual data, wherein the spatial domain resolution expansion process is the reverse process of the spatial domain resolution reduction process, The process includes the step of grouping the extended residual data to obtain a plurality of extended residual groups.

[0007] In one possible embodiment of the present invention, the step of performing a spatial domain resolution reduction process on the image residual data and obtaining the expanded residual data is: The method includes the step of obtaining the expanded residual data by reducing the spatial size corresponding to the image residual data and / or increasing the number of feature channels corresponding to the image residual data.

[0008] In one possible embodiment of the present invention, the step of obtaining the expanded residual data by reducing the spatial size corresponding to the image residual data and / or increasing the number of feature channels corresponding to the image residual data is: The method includes the step of obtaining the expanded residual data by reducing the spatial size corresponding to the image residual data and / or increasing the number of feature channels corresponding to the image residual data, based on spatial region information corresponding to the image residual data.

[0009] In one possible embodiment of the present invention, the step of performing residual reconstruction for each of the plurality of extended residual groups and obtaining image reconstruction features corresponding to each extended residual group is: The steps include constructing a residual recovery sequence based on the aforementioned multiple extended residual groups, The process includes the steps of performing residual reconstruction for each of the plurality of extended residual groups based on the residual reconstruction sequence, and obtaining image reconstruction features corresponding to each extended residual group.

[0010] In one possible embodiment of the present invention, the step of performing residual reconstruction for each of the plurality of extended residual groups based on the residual reconstruction sequence and obtaining image reconstruction features corresponding to each extended residual group is: The steps include traversing the residual restoration sequence and obtaining the current expanded residual group, The steps include obtaining auxiliary information output from the auxiliary coding network, The steps include constructing preliminary information based on the aforementioned auxiliary information, The steps include: performing residual reconstruction on the current extended residual group based on the prior information and obtaining image reconstruction features corresponding to the current extended residual group; The process includes the step of obtaining image reconstruction features corresponding to each extended residual group once the traverse is complete.

[0011] In one possible embodiment of the present invention, the step of constructing prior information based on the auxiliary information is: Steps to obtain extended auxiliary information, The steps include detecting whether the current expanded residual group is the first element in the residual restoration sequence, If it is the first element, the step is to construct preliminary information based on the aforementioned extended auxiliary information, If it is not the first element, the process includes the steps of combining the extended auxiliary information and the convolutional processing result corresponding to the image reconstruction features of the restored extended residual group to obtain combined auxiliary information, and constructing prior information based on the combined auxiliary information.

[0012] Furthermore, in order to achieve the above objectives, the present invention is Performing a spatial region resolution reduction process on the image features corresponding to the image to be encoded to obtain enhanced image features; Grouping the enhanced image features to obtain a plurality of enhanced feature groups; Performing a residual calculation on each of the plurality of enhanced feature groups to obtain image residual data corresponding to each enhanced feature group; Generating an image bitstream based on the image residual data and transmitting the image bitstream to the image decoding side, further providing an image encoding method.

[0013] In one possible embodiment of the present invention, the step of performing a spatial region resolution reduction process on the image features corresponding to the image to be encoded to obtain enhanced image features includes: Obtaining the image features corresponding to the image to be encoded; Performing a spatial region resolution reduction process on the data in the image features to obtain enhanced image features.

[0014] In one possible embodiment of the present invention, the step of performing a spatial region resolution reduction process on the image features corresponding to the image to be encoded to obtain enhanced image features includes: Reducing the spatial size corresponding to the image features of the image to be encoded and / or increasing the number of feature channels corresponding to the image features to obtain enhanced image features.

[0015] In one possible embodiment of the present invention, the step of reducing the spatial size corresponding to the image features of the image to be encoded and / or increasing the number of feature channels corresponding to the image features to obtain enhanced image features includes: Based on the spatial region information corresponding to the image features of the image to be encoded, reducing the spatial size corresponding to the image features and / or increasing the number of feature channels corresponding to the image features to obtain enhanced Image features features.

[0016] In one possible embodiment of the present invention, the step of performing residual calculation for each of the plurality of extended feature groups and obtaining image residual data corresponding to each extended feature group is as follows: Constructing a residual calculation sequence based on the plurality of extended feature groups; Performing residual calculation for each of the plurality of extended feature groups based on the residual calculation sequence, and obtaining image residual data corresponding to each extended feature group.

[0017] In one possible embodiment of the present invention, the step of performing residual calculation for each of the plurality of extended feature groups based on the residual calculation sequence and obtaining image residual data corresponding to each extended feature group is as follows: Traversing the residual calculation sequence to obtain the current extended feature group; Obtaining auxiliary information output from the auxiliary encoding network; Constructing prior information based on the auxiliary information; Performing residual calculation on the current extended feature group based on the prior information, and obtaining image residual data corresponding to the current extended feature group; After the traversal is completed, obtaining image residual data corresponding to each extended feature group.

[0018] In one possible embodiment of the present invention, the step of grouping the extended image features and obtaining a plurality of extended feature groups is as follows: Based on the feature channels corresponding to the extension Image features Grouping the extended image features and obtaining a plurality of extended feature groups.

[0019] In one possible embodiment of the present invention, the step of generating an image bitstream based on the image residual data and transmitting the image bitstream to the image decoding side is as follows: For the image corresponding to each extended feature group Residual dataA step of performing a spatial domain resolution expansion process on the image to be encoded and obtaining image residual data corresponding to the image to be encoded, wherein the spatial domain resolution expansion process is the reverse process of the spatial domain resolution reduction process, The process includes the steps of generating an image bitstream based on image residual data corresponding to the image to be encoded, and transmitting the image bitstream to the image decoding side.

[0020] Furthermore, in order to achieve the above objectives, the present invention is A bitstream decoding module for extracting image residual data or extended residual data from an image bitstream, and for obtaining multiple extended residual groups based on the extracted image residual data or extended residual data, A residual reconstruction module for performing residual reconstruction on each of the aforementioned multiple extended residual groups and obtaining image reconstruction features corresponding to each extended residual group, A data combination module for obtaining reconstructed feature data by performing spatial domain resolution expansion processing on image reconstruction features corresponding to each expanded residual group, The present invention further provides an image decoding device, which includes an image reconstruction module for performing image reconstruction based on the aforementioned reconstruction feature data and obtaining a reconstructed image block.

[0021] Furthermore, in order to achieve the above objectives, the present invention is A feature extraction module for obtaining augmented image features by performing spatial domain resolution reduction processing on image features corresponding to the image to be encoded, A data grouping module for grouping the aforementioned extended image features and obtaining multiple extended feature groups, A residual calculation module for performing residual calculations for each of the aforementioned multiple extended feature groups and obtaining image residual data corresponding to each extended feature group, The present invention further provides an image encoding apparatus that includes a bitstream generation module for generating an image bitstream based on the aforementioned image residual data and transmitting the image bitstream to the image decoding side.

[0022] Furthermore, in order to achieve the above objectives, the present invention is The process involves extracting image residual data from an image bitstream, performing a spatial domain resolution reduction process on the image residual data to obtain augmented residual data, and grouping the augmented residual data to obtain multiple augmented residual groups. The steps include performing residual reconstruction for each of the aforementioned multiple extended residual groups and obtaining image reconstruction features corresponding to each extended residual group, A step of performing a spatial domain resolution expansion process on the image reconstruction features corresponding to each expanded residual group and obtaining reconstruction feature data, wherein the spatial domain resolution expansion process is the reverse process of the spatial domain resolution reduction process. The present invention further provides an image decoding method that includes the step of performing image reconstruction based on the aforementioned reconstructed feature data and obtaining a reconstructed image block.

[0023] In one possible embodiment of the present invention, the step of performing a spatial domain resolution reduction process on the image residual data and obtaining augmented residual data is: The method includes the step of obtaining the expanded residual data by reducing the spatial size corresponding to the image residual data and / or increasing the number of feature channels corresponding to the image residual data.

[0024] In one possible embodiment of the present invention, the step of obtaining the expanded residual data by reducing the spatial size corresponding to the image residual data and / or increasing the number of feature channels corresponding to the image residual data is: Based on the spatial region information corresponding to the image residual data, the expanded residual data is obtained by reducing the spatial size corresponding to the image residual data, or, By increasing the number of feature channels corresponding to the image residual data based on the spatial region information corresponding to the image residual data, the extended residual data is acquired, or The method includes the steps of obtaining the expanded residual data by reducing the spatial size corresponding to the image residual data based on spatial region information corresponding to the image residual data, and increasing the number of feature channels corresponding to the image residual data based on spatial region information corresponding to the image residual data.

[0025] In one possible embodiment of the present invention, the step of grouping the extended residual data to obtain a plurality of extended residual groups is: The process includes the step of uniformly grouping the expanded residual data based on the feature channels corresponding to the expanded residual data to obtain a plurality of expanded residual groups.

[0026] In one possible embodiment of the present invention, when performing spatial domain resolution expansion processing on image reconstruction features corresponding to each expanded residual group, the spatial domain resolution expansion processing is used to restore the spatial size corresponding to the image reconstruction feature to match the spatial size corresponding to the image obtained by feature extraction from the original image, and the spatial domain resolution expansion processing is used to restore the number of channels corresponding to the image reconstruction feature to match the number of channels corresponding to the image obtained by feature extraction from the original image.

[0027] Furthermore, in order to achieve the above objectives, the present invention is The steps include: performing spatial domain resolution reduction on image features corresponding to the image to be encoded to obtain augmented image features; The steps include grouping the aforementioned extended image features and obtaining multiple extended feature groups, A step of performing residual calculations for each of the aforementioned multiple extended feature groups and obtaining image residual data corresponding to each extended feature group, wherein after obtaining the image reconstruction features corresponding to each extended feature group, a spatial domain resolution expansion process is performed on the image reconstruction features corresponding to each extended feature group, and image residual data corresponding to each extended feature group is obtained, and the spatial domain resolution expansion process is the reverse process of the spatial domain resolution reduction process. The present invention further provides an image encoding method that includes the steps of generating an image bitstream based on the image residual data and transmitting the image bitstream to the image decoding side.

[0028] Furthermore, in order to achieve the above objectives, the present invention is A bitstream decoding module for extracting image residual data from an image bitstream, performing spatial domain resolution reduction processing on the image residual data to obtain augmented residual data, and grouping the augmented residual data to obtain multiple augmented residual groups, A residual reconstruction module for performing residual reconstruction on each of the aforementioned multiple extended residual groups and obtaining image reconstruction features corresponding to each extended residual group, A data combination module for obtaining reconstructed feature data by performing a spatial domain resolution expansion process on image reconstruction features corresponding to each expanded residual group, wherein the spatial domain resolution expansion process is the reverse process of the spatial domain resolution reduction process, and The present invention further provides an image decoding device, which includes an image reconstruction module for performing image reconstruction based on the aforementioned reconstruction feature data and obtaining a reconstructed image block.

[0029] Furthermore, in order to achieve the above objectives, the present invention is A feature extraction module for obtaining augmented image features by performing spatial domain resolution reduction processing on image features corresponding to the image to be encoded, A data grouping module for grouping the aforementioned extended image features and obtaining multiple extended feature groups, A residual calculation module for performing residual calculations for each of the aforementioned multiple extended feature groups and obtaining image residual data corresponding to each extended feature group, wherein after obtaining the image reconstruction features corresponding to each extended feature group, the residual calculation module performs a spatial domain resolution expansion process on the image reconstruction features corresponding to each extended feature group and obtains image residual data corresponding to each extended feature group, the spatial domain resolution expansion process being the reverse process of the spatial domain resolution reduction process, and The present invention further provides an image encoding apparatus that includes a bitstream generation module for generating an image bitstream based on the aforementioned image residual data and transmitting the image bitstream to the image decoding side.

[0030] Furthermore, in order to achieve the above objectives, the present invention is Steps include grouping augmented residual data to obtain multiple augmented residual groups, The steps include constructing a residual reconstruction sequence based on the plurality of extended residual groups, performing residual reconstruction for each of the plurality of extended residual groups based on the residual reconstruction sequence, and obtaining image reconstruction features corresponding to each extended residual group, The steps include performing spatial domain resolution expansion processing on the image reconstruction features corresponding to each expanded residual group and obtaining the reconstruction feature data, The present invention further provides an image decoding method that includes the step of performing image reconstruction based on the aforementioned reconstructed feature data and obtaining a reconstructed image block.

[0031] In one possible embodiment of the present invention, the step of performing residual reconstruction for each of the plurality of extended residual groups based on the residual reconstruction sequence and obtaining image reconstruction features corresponding to each extended residual group is: The steps include traversing the residual restoration sequence and obtaining the current expanded residual group, Steps to obtain prior information, The steps include: performing residual reconstruction on the current extended residual group based on the prior information and obtaining image reconstruction features corresponding to the current extended residual group; The process includes the step of obtaining image reconstruction features corresponding to each extended residual group once the traverse is complete.

[0032] In one possible embodiment of the present invention, the step of obtaining the prior information is: Steps include obtaining augmented auxiliary information based on information output from the auxiliary coding network, The steps include detecting whether the current expanded residual group is the first element in the residual restoration sequence, If the current extended residual group is the first element, the steps include constructing prior information based on the extended auxiliary information, If the current extended residual group is not the first element, the process includes the steps of combining the extended auxiliary information with the convolution results corresponding to the image reconstruction features of the restored extended residual group to obtain combined auxiliary information, and constructing prior information based on the combined auxiliary information.

[0033] In one possible embodiment of the present invention, the step of performing residual reconstruction on the current extended residual group based on the prior information and obtaining image reconstruction features corresponding to the current extended residual group is: The steps include processing the prior information using a convolutional neural network to obtain a predicted average value, The method includes the step of adding the predicted mean value and the residuals in the current extended residual group to obtain an image reconstruction feature corresponding to the current extended residual group.

[0034] In one possible embodiment of the present invention, when performing spatial domain resolution expansion processing on image reconstruction features corresponding to each expanded residual group, the spatial domain resolution expansion processing is used to restore the spatial size corresponding to the image reconstruction feature to match the spatial size corresponding to the image obtained by feature extraction from the original image, and the spatial domain resolution expansion processing is used to restore the number of channels corresponding to the image reconstruction feature to match the number of channels corresponding to the image obtained by feature extraction from the original image.

[0035] Furthermore, in order to achieve the above objectives, the present invention is The steps include: performing spatial domain resolution reduction on image features corresponding to the image to be encoded to obtain augmented image features; The steps include grouping the aforementioned extended image features and obtaining multiple extended feature groups, A step comprising: constructing a residual calculation sequence based on the plurality of extended feature groups; performing residual calculations for each of the plurality of extended feature groups based on the residual calculation sequence; and obtaining image residual data corresponding to each extended feature group, wherein after obtaining the image reconstruction features corresponding to each extended feature group, a spatial domain resolution expansion process is performed on the image reconstruction features corresponding to each extended feature group; and obtaining image residual data corresponding to each extended feature group, wherein the spatial domain resolution expansion process is the reverse process of the spatial domain resolution reduction process. The present invention further provides an image encoding method that includes the steps of generating an image bitstream based on the image residual data and transmitting the image bitstream to the image decoding side.

[0036] Furthermore, in order to achieve the above objectives, the present invention is A bitstream decoding module for grouping extended residual data to obtain multiple extended residual groups, A residual reconstruction module for constructing a residual reconstruction sequence based on the plurality of extended residual groups, performing residual reconstruction for each of the plurality of extended residual groups based on the residual reconstruction sequence, and obtaining image reconstruction features corresponding to each extended residual group, A data combination module for obtaining reconstructed feature data by performing spatial domain resolution expansion processing on image reconstruction features corresponding to each expanded residual group, The present invention further provides an image decoding device, which includes an image reconstruction module for performing image reconstruction based on the aforementioned reconstruction feature data and obtaining a reconstructed image block.

[0037] Furthermore, in order to achieve the above objectives, the present invention is A feature extraction module for obtaining augmented image features by performing spatial domain resolution reduction processing on image features corresponding to the image to be encoded, A data grouping module for grouping the aforementioned extended image features and obtaining multiple extended feature groups, A residual calculation module for constructing a residual calculation sequence based on the aforementioned multiple extended feature groups, performing residual calculations for each of the aforementioned multiple extended feature groups based on the residual calculation sequence, and obtaining image residual data corresponding to each extended feature group, wherein after obtaining image reconstruction features corresponding to each extended feature group, a spatial domain resolution expansion process is performed on the image reconstruction features corresponding to each extended feature group, and image residual data corresponding to each extended feature group is obtained, and the spatial domain resolution expansion process is the reverse process of the spatial domain resolution reduction process, The present invention further provides an image encoding apparatus that includes a bitstream generation module for generating an image bitstream based on the aforementioned image residual data and transmitting the image bitstream to the image decoding side.

[0038] Furthermore, in order to achieve the above objectives, the present invention is Steps include obtaining multiple extended residual groups, The steps include performing residual reconstruction for each of the aforementioned multiple extended residual groups and obtaining image reconstruction features corresponding to each extended residual group, The steps include performing spatial domain resolution expansion processing on the image reconstruction features corresponding to each expanded residual group and obtaining the reconstruction feature data, The present invention provides an image decoding method comprising the steps of: obtaining a reconstructed image block by performing a synthetic transformation process on the reconstructed feature data using a pre-constructed synthetic transformation network, wherein the synthetic transformation network is a network constructed based on deep learning or a neural network.

[0039] In one possible embodiment of the present invention, when performing spatial domain resolution expansion processing on image reconstruction features corresponding to each expanded residual group, the spatial domain resolution expansion processing is used to restore the spatial size corresponding to the image reconstruction feature to match the spatial size corresponding to the image obtained by feature extraction from the original image, and the spatial domain resolution expansion processing is used to restore the number of channels corresponding to the image reconstruction feature to match the number of channels corresponding to the image obtained by feature extraction from the original image.

[0040] In one possible embodiment of the present invention, the step of performing residual reconstruction for each of the plurality of extended residual groups and obtaining image reconstruction features corresponding to each extended residual group is: The steps include constructing a residual recovery sequence based on the aforementioned multiple extended residual groups, The process includes the steps of performing residual reconstruction for each of the plurality of extended residual groups based on the residual reconstruction sequence, and obtaining image reconstruction features corresponding to each extended residual group.

[0041] In one possible embodiment of the present invention, the step of obtaining the plurality of extended residual groups is: The process includes extracting image residual data from an image bitstream, performing a spatial domain resolution reduction process on the image residual data to obtain augmented residual data, and grouping the augmented residual data to obtain multiple augmented residual groups, wherein the spatial domain resolution expansion process is the reverse process of the spatial domain resolution reduction process.

[0042] In one possible embodiment of the present invention, the step of obtaining the plurality of extended residual groups is: The process includes extracting extended residual data from an image bitstream and grouping the extended residual data to obtain a plurality of extended residual groups.

[0043] Furthermore, in order to achieve the above objectives, the present invention is The steps include: performing spatial domain resolution reduction on image features corresponding to the image to be encoded to obtain augmented image features; The steps include grouping the aforementioned extended image features and obtaining multiple extended feature groups, A step of performing residual calculations for each of the plurality of extended feature groups and obtaining image residual data corresponding to each extended feature group, wherein after obtaining image reconstruction features corresponding to each extended feature group, a spatial domain resolution expansion process is performed on the image reconstruction features corresponding to each extended feature group, and image residual data corresponding to each extended feature group is obtained, the spatial domain resolution expansion process is the reverse process of the spatial domain resolution reduction process, and a composite transformation process is performed on the image reconstruction features by a pre-constructed composite transformation network to obtain a reconstructed image block, wherein the composite transformation network is a network constructed based on deep learning or a neural network. The present invention further provides an image encoding method that includes the steps of generating an image bitstream based on the image residual data and transmitting the image bitstream to the image decoding side.

[0044] A bitstream decoding module for obtaining multiple extended residual groups, A residual reconstruction module for performing residual reconstruction on each of the aforementioned multiple extended residual groups and obtaining image reconstruction features corresponding to each extended residual group, A data combination module for obtaining reconstructed feature data by performing spatial domain resolution expansion processing on image reconstruction features corresponding to each expanded residual group, The present invention further provides an image decoding device, which includes an image reconstruction module for obtaining a reconstructed image block obtained by performing a synthetic transformation process on the reconstructed feature data using a pre-constructed synthetic transformation network, wherein the synthetic transformation network is a network constructed based on deep learning or a neural network.

[0045] Furthermore, in order to achieve the above objectives, the present invention is A feature extraction module for obtaining augmented image features by performing spatial domain resolution reduction processing on image features corresponding to the image to be encoded, A data grouping module for grouping the aforementioned extended image features and obtaining multiple extended feature groups, A residual calculation module for performing residual calculations for each of the aforementioned multiple extended feature groups and obtaining image residual data corresponding to each extended feature group, wherein after obtaining image reconstruction features corresponding to each extended feature group, a spatial domain resolution expansion process is performed on the image reconstruction features corresponding to each extended feature group, and image residual data corresponding to each extended feature group is obtained, the spatial domain resolution expansion process is the reverse process of the spatial domain resolution reduction process, and a composite transformation process is performed on the image reconstruction features by a pre-constructed composite transformation network to obtain a reconstructed image block, wherein the composite transformation network is a network constructed based on deep learning or a neural network, and the residual calculation module, The present invention further provides an image encoding apparatus that includes a bitstream generation module for generating an image bitstream based on the aforementioned image residual data and transmitting the image bitstream to the image decoding side.

[0046] Furthermore, in order to achieve the above objective, the present invention further provides a decoding device comprising a processor, a memory, and an image decoding program stored in the memory and executable on the processor, wherein when the image decoding program is executed by the processor, the above image decoding method is performed.

[0047] Furthermore, in order to achieve the above objective, the present invention further provides an encoding device comprising a processor, a memory, and an image decoding program and / or image encoding program stored in the memory and executable on the processor, wherein when the image decoding program is executed by the processor, the above image decoding method is performed, and when the image encoding program is executed by the processor, the above image encoding method is performed.

[0048] Furthermore, in order to achieve the above objective, the present invention provides a storage medium in which an image decoding program and / or an image encoding program are stored, wherein the image decoding program is By the processor When executed, the above image decoding method is performed, and the image encoding program By the processor The present invention further provides a storage medium on which the above image encoding method is performed when executed.

[0049] Furthermore, in order to achieve the above objective, the present invention further provides a computer program configured such that, when executed by a processor having memory, the above-described image decoding method or the above-described image encoding method is performed.

[0050] Furthermore, in order to achieve the above objective, the present invention further provides a computer program product including computer program instructions, wherein the computer program instructions are configured such that when executed by a processor having memory, the above-described image decoding method or the above-described image encoding method is performed.

[0051] The present invention extracts image residual data or extended residual data from an image bitstream, obtains a plurality of extended residual groups based on the extracted image residual data or extended residual data, performs residual reconstruction for each of the plurality of extended residual groups, obtains image reconstruction features corresponding to each extended residual group, performs spatial domain resolution expansion processing on the image reconstruction features corresponding to each extended residual group, obtains reconstruction feature data, performs image reconstruction based on the reconstruction feature data, and obtains a reconstructed image block. Since the obtained extended residual data is residual data that has undergone spatial domain resolution reduction processing, it can be grouped at a low resolution and residual reconstruction processing can be performed on a group basis, improving the overall residual reconstruction computation efficiency and reducing time complexity. [Brief explanation of the drawing]

[0052] [Figure 1] This is a schematic diagram of the structure of an electronic device in a hardware execution environment according to an embodiment of the present invention. [Figure 2] This is a schematic flowchart of the first embodiment of the image decoding method of the present invention. [Figure 3] This is an overall block diagram of image compression according to one embodiment of the present invention. [Figure 4] This is a schematic flowchart of a second embodiment of the image decoding method of the present invention. [Figure 5] This is a schematic flowchart of the spatial domain resolution processing according to one embodiment of the present invention. [Figure 6] This is a schematic flowchart of a third embodiment of the image decoding method of the present invention. [Figure 7] This is a schematic flowchart of the image decoding and grouping execution process according to one embodiment of the present invention. [Figure 8] This is a schematic flowchart of a two-step grouping execution in one embodiment of the present invention. [Figure 9] This is a schematic flowchart of the feature enhancement grouping execution process in one embodiment of the present invention. [Figure 10] This is a schematic flowchart of a first embodiment of the image encoding method of the present invention. [Figure 11]This is a schematic flowchart of a second embodiment of the image encoding method of the present invention. [Figure 12] This is a structural block diagram of the first embodiment of the image decoding device of the present invention. [Figure 13] This is a structural block diagram of a first embodiment of the image encoding device of the present invention. The realization of the objectives, functional features, and advantages of the present invention will be further described in conjunction with the examples and with reference to the drawings. [Modes for carrying out the invention]

[0053] The specific examples described herein are used solely for the purpose of illustrating the present invention and should not be considered limiting.

[0054] Referring to Figure 1, Figure 1 is a schematic diagram of the structure of a decoding device or encoding device in a hardware execution environment according to an embodiment of the present invention.

[0055] As shown in Figure 1, the electronic device may include a processor 1001 such as a Central Processing Unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. The communication bus 1002 is used to enable connection communication between these components. The user interface 1003 may include input units such as a display and a keyboard, and optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (e.g., a Wireless-Fidelity (Wi-Fi) interface). The memory 1005 may be high-speed random access memory (RAM) or stable non-volatile memory (NVM) such as magnetic disk memory. Optionally, the memory 1005 may be a storage device independent of the processor 1001.

[0056] As those skilled in the art will understand, the structure shown in Figure 1 is not limited to electronic devices and may include more or fewer components, combinations of components, or different arrangements of components than those shown.

[0057] As shown in Figure 1, the memory 1005 as a storage medium may include an operating system, a network communication module, a user interface module, an image decoding program, and / or an image encoding program.

[0058] In the electronic device shown in Figure 1, the network interface 1004 is mainly used for data communication with a network server, and the user interface 1003 is mainly used for data interaction with a user. The processor 1001 and memory 1005 in the electronic device of the present invention may be provided in a decoding device or an encoding device. The electronic device calls an image decoding program or an image encoding program stored in the memory 1005 via the processor 1001 and executes an image decoding method or an image encoding method according to an embodiment of the present invention.

[0059] The embodiments of the present invention provide an image decoding method, and Figure 2 is a schematic flowchart of the first embodiment of the image decoding method of the present invention.

[0060] In this embodiment, the image decoding method includes the following steps.

[0061] Step S10: Extract image residual data or extended residual data from the image bitstream, and obtain multiple extended residual groups based on the extracted image residual data or extended residual data.

[0062] In this embodiment, the implementing entity may be a decoding device used when performing decoding processing on image data. This decoding device may be an electronic device such as a personal computer or a server, or of course, any other device capable of performing the same or similar functions. This embodiment is not limited to this, and in this embodiment and subsequent embodiments, the image decoding method of the present invention will be explained using the decoding device as an example.

[0063] In the image coding process, the coding device typically decodes the coded image bitstream after coding is complete and determines whether or not it is necessary to adjust the parameters used during coding based on the image quality obtained by coding. Therefore, the implementing entity in this embodiment may be the coding device.

[0064] The image bitstream may be a bitstream generated after an encoding device has performed encoding on image data that needs to be compressed and encoded. When generating the image bitstream, the encoding device reduces the spatial domain resolution of image features to reduce the time complexity of mean prediction in the encoding process. Finally, the encoding device directly encodes the generated image residual data or extended residual data into the image bitstream. In this case, the decoding device can directly extract the image residual data or extended residual data from the image bitstream. In this case, the decoding device can process the extracted image residual data or extended residual data to obtain multiple extended residual groups, perform residual reconstruction for each group, and reduce the time complexity of mean prediction in the image decoding process.

[0065] Technical terms related to image encoding or decoding include JPEG (Joint Photographic Experts Group), JPEG-AI (Joint Photographic Experts Group Artificial Intelligence), entropy encoding, neural network (NN), convolutional neural network (CNN), feature, and rate-distortion optimized, which will be explained here.

[0066] JPEG (Joint Photographic Experts Group) is a standard for compressing continuous-tone still images. Its file extensions are .jpg or .jpeg, and it is the most common image file format. It primarily employs a joint coding scheme using predictive coding (e.g., Differential Pulse Code Modulation, DPCM), discrete cosine transform (DCT), and entropy coding to remove redundant image and color data. It is a lossy compression format, capable of compressing images into a small memory space, but it does cause some damage to the image data. In particular, if the compression ratio is too high, the quality of the decompressed image will deteriorate, so it is not advisable to use excessively high compression ratios when pursuing high-quality images.

[0067] The scope of JPEG AI is to create a machine learning-based image coding standard that provides a single-stream, compact, compressed region representation that is visible to humans and offers significantly improved compression efficiency compared to commonly used image coding standards at the same subjective quality, effectively enhancing performance in image processing and computer vision tasks. JPEG AI is intended for a wide range of applications, including cloud storage, vision surveillance, autonomous driving of automobiles and equipment, image acquisition, storage and management, real-time monitoring of vision data, and media distribution. Its goal is to design a coding solution that significantly improves the compression efficiency of commonly used coding standards at the same subjective quality and provides effective compressed region processing for machine learning-based image processing and computer vision tasks. Other important requirements include hardware / software-friendly coding and decoding, support for 8-bit and 10-bit depths, and efficient coding and progressive decoding of images with text and graphics.

[0068] Entropy coding is a method of coding in which no information is lost during the coding process, according to the principle of entropy. Information entropy is the average amount of information (degree of uncertainty) in the source. Common entropy coding methods include Shannon coding, Huffman coding, and arithmetic coding.

[0069] The neural network of this invention refers to an artificial neural network, not a biological neural network. A neural network is a computational model composed of a large number of nodes (called neurons) connected to one another. In an artificial neural network, neuron processing units can represent different objects, such as features, alphabets, concepts, or several meaningful abstract modes. There are three types of processing units in the network: input units, output units, and hidden units. Input units receive external signals and data, output units realize the output of the system processing results, and hidden units are units that are between the input and output units and cannot be observed from outside the system. The connection weights between neurons reflect the strength of the connections between units, and the representation and processing of information are reflected in the connection relationships of the network processing units. An artificial neural network is an unprogrammed, brain-like information processing method, and its essence is to acquire parallel and distributed information processing capabilities through network transformation and dynamic activity, mimicking the information processing capabilities of the human brain and nervous system to different degrees and levels. Currently, commonly used neural networks in the field of video processing include convolutional neural networks (CNNs), recurrent neural networks (RNNs), and fully connected networks.

[0070] Convolutional neural networks (CNNs) are feedforward neural networks and one of the representative network structures in deep learning techniques. Their artificial neurons can respond to peripheral units within a certain coverage area, demonstrating excellent performance in large-scale image processing. Generally, the basic structure of a CNN includes two layers: one is a feature extraction layer (also called a convolutional layer), where the input of each neuron is connected to the local receptive field of the previous layer, extracting local features. Once local features are extracted, their positional relationship with other features is also determined. The other layer is a feature mapping layer (also called an activation layer), where each computational layer of the network consists of multiple feature mappings. Each feature mapping is a plane, and the weights of all neurons in that plane are equal. The feature mapping structure may use functions such as the Sigmoid function, ReLU function, Leaky-ReLU function, PReLU function, or GDN (Generalized Difference Network) function as the activation function of the convolutional network. Furthermore, because neurons in a single mapping plane share weights, the number of free parameters in the network is reduced. One advantage of CNNs compared to conventional image processing algorithms is that they can avoid complex pre-processing processes for images (such as extracting artificial features), directly inputting original images and performing end-to-end learning. Another advantage of CNNs compared to conventional neural networks is that while conventional neural networks employ a fully connected architecture, meaning all neurons from the input layer to the hidden layer are connected, resulting in a huge number of parameters and making network training time-consuming and difficult, CNNs avoid this difficulty by using methods such as local connections and weight sharing.

[0071] A key feature of the present invention is a three-dimensional feature matrix of C × W × H (as shown in Figure 3, which is a schematic diagram of the matrix structure of this embodiment). C represents the number of channels, H represents the feature height, and W represents the feature width. The feature matrix may be the input to a neural network or the output of a neural network.

[0072] Various metrics can be used to evaluate coding efficiency, including bitrate, PSNR, MS-SSIM, VMAF FSIM, and PSNRHVS. Of course, more metrics may be included, and this is not limited to those mentioned here. A smaller bitstream results in a higher compression ratio, and a higher PSNR indicates better image coding efficiency. When selecting a mode, the discriminant is essentially a comprehensive evaluation of both. The cost corresponding to a mode is given by J(mode) = D + λ*R. Here, D represents distortion, usually evaluated using the SSE (Sum of Squared Errors) metric, where SSE refers to the mean square sum of the differences between the reconstructed block and the source image; λ is the Lagrangian multiplier; and R is the actual number of bits required to encode the image block in that mode, including the total number of bits needed to encode mode information, residuals, etc.

[0073] In one possible embodiment of this embodiment, the grouping of acquired extended residual data can be performed based on the feature channel corresponding to each extended residual data, in which case step S10 of this embodiment is Steps include extracting extended residual data from the image bitstream, The process may also include the step of grouping the augmented residual data based on feature channels corresponding to the augmented residual data to obtain a plurality of augmented residual groups.

[0074] Grouping extended residual data based on the feature channels corresponding to the extended residual data and obtaining multiple extended residual groups may also be done by uniformly dividing the extended residual data into multiple groups based on the corresponding feature channels. For example, if the total number of feature channels corresponding to the extended residual data is 20, then the extended residual data with corresponding feature channels 1 to 10 may be divided into one group, and the extended residual data with corresponding feature channels 11 to 20 may be divided into another group. The number of groups to be uniformly divided may be predetermined by the administrator of the decoding device and is not limited in this embodiment.

[0075] Of course, when grouping specifically, the data may be divided unevenly. In this case, grouping the extended residual data based on the feature channels corresponding to the extended residual data and obtaining multiple extended residual groups may also be done by dividing the extended residual data into multiple groups based on the corresponding feature channels according to a pre-configured grouping rule. Here, the pre-configured grouping rule may be pre-configured by the administrator of the decoding device according to the actual needs. For example, the pre-configured grouping rule may be set to divide the first m / n (where n is the total number of feature channels, m is a pre-configured value, and the range is [1, n]) of extended residual data into one group and the remaining extended residual data into another group.

[0076] In actual use, when grouping augmented residual data based on the feature channels corresponding to the augmented residual data and obtaining multiple augmented residual groups, the augmented residual data corresponding to one feature channel may be divided into one group. For example, if the total number of feature channels corresponding to the augmented residual data is 20, then the augmented residual data may be divided into 20 groups based on the differences in the feature channels.

[0077] Step S20: Residual reconstruction is performed for each of the multiple extended residual groups, and image reconstruction features corresponding to each extended residual group are obtained.

[0078] Performing residual reconstruction on the extended residual group and obtaining image reconstruction features corresponding to the extended residual group may also be done by performing mean prediction on the extended residual group and adding the residual data in the extended residual group with the mean value obtained by the prediction to obtain image reconstruction features corresponding to the extended residual group.

[0079] Step S30: Spatial domain resolution expansion processing is performed on the image reconstruction features corresponding to each expanded residual group, and reconstruction feature data is obtained.

[0080] The extended residual data has undergone a spatial domain resolution reduction process, and the spatial size and number of channels corresponding to each data point differ from those of the image features obtained by the encoding device when it first extracted features from the original image. In this case, to ensure the smooth execution of image reconstruction, a spatial domain resolution expansion process can be performed on the image reconstruction features corresponding to each extended residual group to restore the spatial size and number of channels corresponding to the image reconstruction features to match those of the image features obtained by the feature extraction process from the original image.

[0081] The spatial domain resolution expansion process may be the reverse process of the spatial domain resolution reduction process performed in the encoding device.

[0082] Step S40: Image reconstruction is performed based on the reconstruction feature data, and a reconstructed image block is obtained.

[0083] After obtaining reconstruction feature data in which the spatial size and number of channels match the image features corresponding to the original image, image reconstruction can be performed based on the reconstruction feature data, and a reconstructed image block can be obtained.

[0084] Here, performing image reconstruction based on the reconstructed feature data and obtaining a reconstructed image block may also be performed by applying a synthetic transformation process to the reconstructed feature data using a pre-constructed synthetic transformation network to achieve image reconstruction and obtain a reconstructed image block. Here, the synthetic transformation network may be a network constructed based on deep learning or a neural network.

[0085] For ease of understanding, we will refer to Figure 3, but this does not limit the proposed technology. Figure 3 is an overall block diagram of image compression. As shown in Figure 3, on the encoding device side, x is the input image, the main encoder (Analysis Transform Net) generates a latent representation y, and a large bitrate is consumed to directly encode y. Therefore, a context (Context Model Net) and a hyperparameter coding network (Hyper Encoder Net) are introduced, prediction is performed by a hyperparameter decoding network (Hyper Decoder Net), the prediction result μ is obtained, and the residuals

number

number

[0086] On the decoding device side, entropy decoding is performed on the image bitstream.

number

number

number

[0087] This embodiment extracts image residual data or extended residual data from an image bitstream, obtains multiple extended residual groups based on the extracted image residual data or extended residual data, performs residual reconstruction on each of the multiple extended residual groups, obtains image reconstruction features corresponding to each extended residual group, performs spatial domain resolution expansion processing on the image reconstruction features corresponding to each extended residual group, obtains reconstruction feature data, performs image reconstruction based on the reconstruction feature data, and obtains a reconstructed image block. Since the obtained extended residual data is residual data that has undergone spatial domain resolution reduction processing, it can be grouped at a low resolution and residual reconstruction processing can be performed on a group basis, improving the overall residual reconstruction computation efficiency and reducing time complexity.

[0088] Figure 4 is a schematic flowchart of a second embodiment of the image decoding method of the present invention.

[0089] Based on the first embodiment described above, step S10 of the image decoding method of this embodiment includes the following steps.

[0090] Step S101: Extract image residual data from the image bitstream.

[0091] The encoding device may perform spatial domain resolution augmentation on the generated augmented residual data to restore image residual data whose spatial size and number of channels match the image features corresponding to the original image, and then encode the image residual data into an image bitstream. In this case, when the decoding device decodes the image bitstream, only the image residual data can be extracted from the image bitstream.

[0092] Step S102: The image residual data is subjected to spatial domain resolution reduction processing to obtain expanded residual data.

[0093] To facilitate subsequent grouping and reduce temporal complexity, after acquiring image residual data, spatial domain resolution reduction processing may be performed on the image residual data to obtain expanded residual data. To ensure decoding accuracy, the methods employed by the encoding device and decoding device when performing spatial domain resolution reduction processing must be the same, and the spatial domain resolution expansion processing must be the reverse process of the spatial domain resolution reduction processing.

[0094] After applying spatial domain resolution reduction processing to image residual data, the spatial size becomes smaller. Consequently, when processing the image residual data, a smaller convolutional kernel can be used. For example, if a convolutional kernel with a size of 5x5 is used in the original image residual data, applying spatial domain resolution reduction processing to the receptive field of the convolutional kernel is equivalent to using a convolutional kernel with a size of 3x3.

[0095] In actual use, the step of performing spatial domain resolution reduction processing on the image residual data of this embodiment and obtaining expanded residual data is: The method may also include the step of obtaining the expanded residual data by reducing the spatial size corresponding to the image residual data and / or increasing the number of feature channels corresponding to the image residual data.

[0096] The reduction amount for reducing the spatial size corresponding to the image residual data, and the increase amount for increasing the number of feature channels corresponding to the image residual data, may be pre-set by the administrator of the decoding device, and are not limited to this embodiment.

[0097] When actually executing the process, you may set only the amount of spatial size reduction, or only the amount of feature channel increase, and the decoding device will adaptively adjust the spatial size or the number of feature channels. Of course, you may also set both the amount of spatial size reduction and the amount of feature channel increase and have the device execute the process without using the device's adaptive function.

[0098] For example, if the image features corresponding to the image residual data are y∈R {H,W,C} If H is the height of the image feature, W is the width of the image feature, and C is the number of feature channels corresponding to the image feature, then in this case, H and W can be reduced to half of the original, and in this case, in order to keep the amount of data the same, the number of feature channels becomes four times the original, and in this case, the image feature corresponding to the obtained expanded residual data is y∈R {H / 2,W / 2,4C} It can be expressed as follows.

[0099] In concrete implementation, spatial domain resolution reduction processing may be performed based on spatial domain information or frequency domain information corresponding to the image residual data, or spatial domain resolution reduction processing may be performed by a pre-set convolutional layer. In this case, the step of acquiring the expanded residual data by reducing the spatial size corresponding to the image residual data in this embodiment and / or increasing the number of feature channels corresponding to the image residual data is: A step of acquiring expanded residual data by reducing the spatial size corresponding to the image residual data and / or increasing the number of feature channels corresponding to the image residual data, based on spatial region information corresponding to the image residual data. Or, A step of acquiring expanded residual data by reducing the spatial size corresponding to the image residual data and / or increasing the number of feature channels corresponding to the image residual data, based on frequency domain information corresponding to the image residual data. Or, The process may include the step of acquiring expanded residual data by reducing the spatial size corresponding to the image residual data and / or increasing the number of feature channels corresponding to the image residual data, based on a pre-configured convolutional layer.

[0100] To facilitate understanding, we will refer to Figure 5, but this does not limit the proposed technology. Figure 5 is a schematic flow diagram of the spatial domain resolution processing in this embodiment. As shown in method (a) in Figure 5, SpaceShuffle is one method of spatial domain resolution reduction processing. If the input to the processing process is a1 ∈ R{H, W, C} and its output is a2 ∈ R{H / 2, W / 2, 4* C}, then its mathematical expression is as follows. a2[h,w,c]=a1[2*h,2*w,c] a2[h,w,C+c]=a1[2*h+1,2*w+1,c] a2[h,w,2*C+c]=a1[2*h,2*w+1,c] a2[h,w,3*C+c]=a1[2*h+1,2*w,c] Here, h represents the index value in the height dimension, with a range of 0 to H-1; w represents the index value in the width dimension, with a range of 0 to W-1; and c represents the index value in the channel dimension, with a range of 0 to C-1. Furthermore, a1[h,w,c] represents the values ​​of the corresponding h, w, and c index positions in a1. In method (a) of Figure 5, unSpaceShuffle is the inverse process of SpaceShuffle, and unSpaceShuffle is one method of spatial domain resolution expansion processing. In this process, the output result is the same as the original input after the input has undergone SpaceShuffle and unSpaceShuffle processing, so the process is reversible. The input and output of the above processing process may be other data, and are not limited in this embodiment.

[0101] As shown in method (b), method (b) is another method that performs spatial domain resolution processing based on spatial domain information corresponding to image residual data. The specific sampling method of Pixshuffle is similar to method (a), but the difference is that the data is arranged alternately in the channel dimension. The data from one channel before processing is divided into four channels according to the spatial domain and arranged sequentially.

[0102] Method (c) in Figure 5 is a method that performs spatial domain resolution processing based on frequency domain information corresponding to image residual data. As shown in method (c), the Wavelet transform is a two-dimensional wavelet transform that can perform frequency domain partitioning on the data and outputs four frequency domain subbands with dimensions [H / 2, W / 2, C], each representing different frequency domain characteristics. The Inv Wavelet transform is the inverse process, which synthesizes the frequency domain subbands with the original data, and this process is reversible.

[0103] Method (d) in Figure 5 is a method that performs spatial domain resolution processing based on a pre-configured convolutional layer. As shown in method (d), Convolution is a convolutional transformation, which is directly input to the convolutional layer for processing, obtaining the output [H / 2, W / 2, 4*C] and achieving the objective of reducing the spatial domain resolution. The corresponding inverse process is similarly processed in the convolutional layer and restored to the original image size, but this method is irreversible, and the data after processing in both ways is different from the original data.

[0104] In one possible embodiment of this embodiment, the image residual data may be grouped and the spatial domain resolution reduction process performed on each group before performing the spatial domain resolution reduction process on the image residual data. In this case, step S102 of this embodiment is: The steps include: performing data grouping based on feature channels corresponding to the aforementioned image residual data to obtain at least one image residual group; The process includes the step of performing a spatial domain resolution reduction process on the data in at least one image residual group to obtain augmented residual data.

[0105] When grouping data based on feature channels corresponding to image residual data, the same or similar method as when grouping augmented residual data may be used. Performing spatial domain resolution reduction on data in at least one image residual group to obtain augmented residual data may also be performed by performing spatial domain resolution reduction on data in at least one image residual group, and then aggregating the data after spatial domain resolution reduction to obtain augmented residual data.

[0106] For example, if the image features corresponding to the image residual data are y∈R {H,W,C} If so, in this case, we can first divide it uniformly into two groups based on the feature channels, and in this case, the image features corresponding to the data in the image residual group of each group are y∈R {H,W,C / 2} It can be expressed as follows.

[0107] Step S103: The extended residual data is grouped to obtain multiple extended residual groups.

[0108] The aforementioned grouping of extended residual data to obtain multiple extended residual groups may also involve uniformly dividing the extended residual data into multiple groups based on the corresponding feature channels. For example, if the total number of feature channels corresponding to the extended residual data is 20, then the extended residual data with corresponding feature channels 1 to 10 may be divided into one group, and the extended residual data with corresponding feature channels 11 to 20 may be divided into another group. The number of groups to which the data is uniformly divided may be predetermined by the administrator of the decoding device, and is not limited in this embodiment.

[0109] Of course, when grouping specifically, the data may be divided unevenly, and grouping the extended residual data based on the feature channels corresponding to the extended residual data and obtaining multiple extended residual groups may also be done by dividing the extended residual data into multiple groups based on the corresponding feature channels according to a pre-configured grouping rule, where the pre-configured grouping rule may be pre-configured by the administrator of the decoding device according to the actual needs, for example, the pre-configured grouping rule may be set to divide the first m / n (where n is the total number of feature channels, m is a pre-configured value, and the range is [1, n)) of extended residual data into one group and divide the remaining extended residual data into another group.

[0110] In actual use, when grouping augmented residual data based on the feature channels corresponding to the augmented residual data and obtaining multiple augmented residual groups, the augmented residual data corresponding to one feature channel may be divided into one group. For example, if the total number of feature channels corresponding to the augmented residual data is 20, then the augmented residual data may be divided into 20 groups based on the differences in the feature channels.

[0111] In this embodiment, before grouping, it is detected whether the extracted data is image residual data or extended residual data, and if it is image residual data, spatial domain resolution reduction processing is performed. This ensures that even if the image bitstream transmitted from the encoding side contains image residual data, it can be properly grouped and processed in groups after processing, thereby improving the versatility of the image decoding method in this embodiment.

[0112] Figure 6 is a schematic flowchart of a third embodiment of the image decoding method of the present invention.

[0113] Based on the first embodiment described above, step S20 of the image decoding method of this embodiment includes the following steps.

[0114] Step S201: Construct a residual recovery sequence based on the multiple extended residual groups.

[0115] By separating the extended residuals into multiple groups and then performing residual reconstruction for each group, the time complexity of mean prediction in the image decoding process can be reduced. In this case, a residual reconstruction sequence may be constructed based on multiple extended residual groups to determine the residual reconstruction order for each extended residual group.

[0116] Step S202: Based on the residual restoration sequence, residual restoration is performed for each of the multiple extended residual groups, and image reconstruction features corresponding to each extended residual group are obtained.

[0117] Performing residual reconstruction for each of multiple extended residual groups based on the residual reconstruction sequence may also be performed sequentially for each of the multiple extended residual groups based on the sequence order in the residual reconstruction sequence.

[0118] In actual use, residual reconstruction may be performed sequentially using a sequence traverse method, in which case step S202 of this embodiment is, The steps include traversing the residual restoration sequence and obtaining the current expanded residual group, The steps include obtaining auxiliary information output from the auxiliary coding network, The steps include constructing preliminary information based on the aforementioned auxiliary information, The steps include: performing residual reconstruction on the current extended residual group based on the prior information and obtaining image reconstruction features corresponding to the current extended residual group; The process may include the step of obtaining image reconstruction features corresponding to each extended residual group once the traverse is complete.

[0119] Traversing the residual reconstruction sequence to obtain the current extended residual group may also be interpreted as traversing the residual reconstruction sequence and using the traversed extended residual group as the current extended residual group. The auxiliary coding network may be the auxiliary network shown in Figure 3 above (Hyper Encoder Net or Hyper Decoder Net).

[0120] In actual use, residual reconstruction for the current extended residual group based on prior information and obtaining image reconstruction features corresponding to the current extended residual group may be performed by processing the prior information using a prediction parameter fusion network to obtain a predicted mean value, adding the predicted mean value to the residuals in the current extended residual group to achieve residual reconstruction, and obtaining image reconstruction features corresponding to the current extended residual group.

[0121] In actual use, the step of constructing prior information based on the auxiliary information of this embodiment is, Steps to obtain extended auxiliary information, The steps include detecting whether the current expanded residual group is the first element in the residual restoration sequence, If it is the first element, the step is to construct preliminary information based on the aforementioned extended auxiliary information, If it is not the first element, the process may include the steps of: obtaining combined auxiliary information by combining the extended auxiliary information with the convolutional processing results corresponding to the image reconstruction features of the restored extended residual group; and constructing prior information based on the combined auxiliary information.

[0122] The spatial size and number of feature channels corresponding to the auxiliary information output from the auxiliary coding network actually match those of the original image. However, since the extended residual data has actually undergone spatial domain resolution reduction processing, in order to ensure smooth channel joining, it is necessary to first process the auxiliary information to obtain extended auxiliary information and then construct prior information based on the extended auxiliary information.

[0123] When constructing prior information based on augmented auxiliary information, in order to improve the accuracy of mean prediction, prior information may be constructed by combining image features corresponding to already reconstructed augmented residual groups. If the current augmented residual group is the first element in the residual reconstruction sequence, it is indicated that this augmented residual group is the first augmented residual group to be reconstructed. In this case, since there are no already reconstructed augmented residual groups, prior information can be directly constructed based on augmented auxiliary information.

[0124] If the current extended residual group is not the first element in the residual reconstruction sequence, then, since there is already a reconstructed extended residual group, a convolutional operation can be performed on the image reconstruction features of the reconstructed extended residual group using a convolutional layer. The extended auxiliary information and the convolutional results corresponding to the image reconstruction features of the reconstructed extended residual group are channel-joined, and prior information can be constructed based on the combined auxiliary information obtained by the joining process. Here, when constructing the prior information and selecting the reconstructed extended residual groups, all reconstructed extended residual groups may be selected, or only some of the reconstructed extended residual groups may be selected.

[0125] In one possible embodiment of this example, when combining the results, feature enhancement may be first performed on the image reconstruction features of the restored extended residual group to further improve the prediction effect, and the step of combining the extended auxiliary information and the convolution processing results corresponding to the image reconstruction features of the restored extended residual group to obtain combined auxiliary information is: A step to obtain image reconstruction features corresponding to the restored extended residual group, The steps include: performing feature enhancement on the aforementioned image reconstruction features to obtain enhanced reconstruction features; The process may also include the step of obtaining combined auxiliary information by combining the auxiliary information and the convolutional processing result corresponding to the enhanced reconstruction feature.

[0126] In actual use, performing feature enhancement on image reconstruction features and obtaining enhanced reconstruction features may involve performing operations such as missing value handling and outlier handling on the image reconstruction features.

[0127] Before combining auxiliary information with image reconstruction features, feature enhancement is first performed on the image reconstruction features to obtain enhanced reconstruction features. This improves the reliability of the enhanced reconstruction features, thereby improving the reliability of the prior information and increasing the accuracy when making mean predictions based on the prior information.

[0128] In actual use, when performing feature enhancement on image reconstruction features corresponding to a restored extended residual group, the predicted mean, auxiliary information, image residual data, and / or residual data variance corresponding to the restored extended residual group can be employed. In this case, the step of performing feature enhancement on the image reconstruction features and obtaining the enhanced reconstruction features in this embodiment is: The steps include obtaining the predicted mean, auxiliary information, image residual data and / or residual data variance corresponding to the restored extended residual group, The process may also include the step of performing feature enhancement on the image reconstruction features based on the predicted mean values, auxiliary information, image residual data and / or residual data variances corresponding to the restored extended residual group, and obtaining enhanced reconstruction features.

[0129] The predicted mean value corresponding to the restored extended residual group may be the value obtained when predicting the mean value during residual restoration for the restored extended residual group. The residual data variance may be the variance value of the image residual data corresponding to the restored extended residual group.

[0130] In one possible embodiment of this embodiment, in order to improve the reconstruction effect of image reconstruction, step S30 of this embodiment is performed The steps include: performing feature enhancement on the image reconstruction features corresponding to each extended residual group to obtain enhanced reconstruction features corresponding to each extended residual group; The step may also include performing a spatial domain resolution expansion process on the enhanced reconstruction features corresponding to each expanded residual group to obtain the reconstructed feature data.

[0131] Performing feature enhancement on image reconstruction features corresponding to the extended residual group and obtaining enhanced reconstruction features corresponding to the extended residual group may involve performing feature enhancement on image reconstruction features corresponding to the extended residual group by employing the predicted mean, auxiliary information, image residual data, and / or residual data variance corresponding to the extended residual group.

[0132] By performing spatial domain resolution augmentation on the image reconstruction features corresponding to each extended residual group, and then performing feature enhancement on the image reconstruction features corresponding to each extended residual group before acquiring the reconstruction feature data, and then performing spatial domain resolution augmentation on the enhanced reconstruction features corresponding to each extended residual group before acquiring the reconstruction feature data, it is possible to ensure that the reliability of the final constructed reconstruction feature data is higher, and the quality of the reconstructed image blocks acquired after image reconstruction is higher.

[0133] For ease of understanding, the explanation will refer to Figures 7, 8, and 9, but will not limit the scope of this technical proposal. Figure 7 is a schematic flowchart of the image decoding grouping execution in this embodiment, Figure 8 is a schematic flowchart of the double grouping execution in this embodiment, and Figure 9 is a schematic flowchart of the feature enhancement grouping execution in this embodiment.

[0134] As shown in Figure 7, the image features corresponding to the image residual data are y∈R {H,W,C} First, we perform a spatial domain resolution reduction process on it, and the resulting value is y∈R {H / 2,W / 2,4C} At this point, the data is divided into two groups (group1 and group2), a residual reconstruction sequence (group1-group2) is constructed, the mean value mu of Group1 points is obtained via the network using auxiliary (Psi) information to obtain the image reconstruction features of Group1 points, features are extracted using the image reconstruction features of Group1 points, channel coupling is performed with Psi, the mean value mu of Group2 is obtained via the network, and decoding is performed to obtain Group2. The Group1 points and Group2 points are combined and channel coupling is performed, and further special spatial domain resolution expansion processing (i.e., the reverse process of spatial domain resolution reduction processing) is performed to obtain reconstruction feature data, and image reconstruction is performed based on the reconstruction feature data.

[0135] The execution process when grouping is performed once before spatial domain resolution reduction is shown in Figure 8. The image features corresponding to the image residual data are y∈R {H,W,C} And if we uniformly divide it into y1 and y2 based on the corresponding feature channel, then y1∈R {H,W,C / 2} Then, after performing spatial domain resolution reduction processing on each and dividing y1 into part1, part2, part3, and part4, we get the following. Part 1 represents the even-numbered rows and columns in y1 whose spatial domain position is at the y1 spatial domain index, and contains data information for all channels. Part 2 represents the odd-numbered rows and columns in y1 whose spatial domain position is at the y1 spatial domain index, and includes data information for all channels. Part 3 represents the even-numbered rows and odd-numbered columns in y1 whose spatial domain position is at the y1 spatial domain index, and includes data information for all channels. Part 4 represents the odd-numbered rows and even-numbered columns in y1 whose spatial domain position is at the y1 spatial domain index, and includes data information for all channels.

[0136] Similarly, y2 may be divided into four similar parts: part 5, part 6, part 7, and part 8. In this case, the constructed residual reconstruction sequence is "part 1-part 2-part 3-part 4-part 5-part 6-part 7-part 8". In this case, when performing residual reconstruction on part 1, prior information can be directly generated based on auxiliary information. When performing residual reconstruction on part 2, prior information can be generated based on the image reconstruction features and auxiliary information of part 1. When performing residual reconstruction on part 3, prior information can be generated based on the image reconstruction features and auxiliary information of part 1 and part 2. ...Image reconstruction features corresponding to all parts are obtained.

[0137] Of course, two residual reconstruction sequences may be constructed based on y1 and y2, in which case the two residual reconstruction sequences would be "part1-part2-part3-part4" and "part5-part6-part7-part8," respectively. In this case, first, an image reconstruction is performed on y1 based on the residual reconstruction sequence consisting of parts 1-4 using a process similar to the above, and image reconstruction features corresponding to each part of y1 are obtained. Next, prior information is generated based on the image reconstruction features and auxiliary information corresponding to y1, and an image reconstruction is performed on y2 based on the residual reconstruction sequence consisting of parts 5-8 and the generated prior information, and image reconstruction features corresponding to each part of y2 are obtained.

[0138] After dividing y into y1 and y2, in order to perform the spatial region resolution reduction process for each of them, in this case, after obtaining the image reconstruction features corresponding to each part in y1 and y2, by performing the spatial region resolution enlargement process on each of the image reconstruction features corresponding to each part in y1 and y2 and then aggregating them, complete reconstruction feature data can be obtained.

[0139] The specific execution process when feature enhancement is performed after grouping is shown in Fig. 9(a). The image features corresponding to the image residual data are y ∈ R {H,W,C} and first, a spatial region resolution reduction process is performed on it, and the obtained result is y ∈ R {H / 2,W / 2,4C} At this time, it is divided into two groups (group1 and group2), the average value mu of Group1 points is obtained through the network by means of auxiliary (Psi) information to obtain the image reconstruction features of Group1, feature enhancement is performed on the image reconstruction features of Group1 using the predicted average value corresponding to Group1, and the enhanced image feature Group1_E is obtained. The features of the enhanced image feature are extracted, channel combination is performed with Psi, the average value mu of Group2 is obtained through the network, decoding is performed to obtain the image reconstruction features of Group2, and feature enhancement is performed on the image reconstruction features of Group2 based on the average value mu of Group2 to obtain the enhanced image feature Group2_E of Group2. Channel combination is performed on Group1_E and Group2_E, and further a special spatial region resolution enlargement process (i.e., the reverse process of the spatial region resolution reduction process) is performed to obtain reconstruction feature data, and image reconstruction is performed based on the reconstruction feature data. Here, when feature enhancement is performed in Fig. 9(a), the specific structure of the network (Enhance_Net) used for the feature enhancement is shown in Fig. 9(b).

[0140] This embodiment constructs a residual restoration sequence based on the multiple extended residual groups, performs residual restoration for each of the multiple extended residual groups based on the residual restoration sequence, and obtains image reconstruction features corresponding to each extended residual group. Since a residual restoration sequence is constructed based on multiple extended residual groups and the order of residual restoration can be determined by the residual restoration sequence, it is possible to quickly determine whether or not a restored extended residual group exists, and if a restored extended residual group exists, more accurate prior information can be constructed based on the image feature data corresponding to the restored extended residual group.

[0141] The embodiments of the present invention provide an image coding method, and Figure 10 is a schematic flowchart of the first embodiment of the image coding method of the present invention.

[0142] In this embodiment, the image encoding method includes the following steps.

[0143] Step S100: Spatial domain resolution reduction processing is performed on the image features corresponding to the image to be encoded, and augmented image features are obtained.

[0144] To facilitate subsequent grouping and reduce temporal complexity, after obtaining image features corresponding to the image to be encoded, the image Features A spatial domain resolution reduction process may be performed on the image to obtain augmented image features. The image to be encoded is the original image described in the embodiment of the image decoding method described above.

[0145] After performing spatial domain resolution reduction on image features, the spatial size becomes smaller, allowing for the use of a smaller convolution kernel when processing the image features. For example, the original image Features In this case, if a convolutional kernel with a size of 5x5 is used, it is equivalent to using a convolutional kernel with a size of 3x3 after performing a spatial domain resolution reduction process on the receptive field of the convolutional kernel.

[0146] In one possible embodiment of this embodiment, step S100 of this embodiment is The process may also include the step of obtaining augmented image features by reducing the spatial size corresponding to the image features of the image to be encoded and / or increasing the number of feature channels corresponding to the image features.

[0147] The reduction amount for reducing the spatial size corresponding to image features, and the increase amount for increasing the number of feature channels corresponding to image features, may be pre-set by the administrator of the encoding device, and are not limited to this embodiment.

[0148] When actually executing the process, you may set only the amount of spatial size reduction, or only the amount of feature channel increase, and the decoding device will adaptively adjust the spatial size or the number of feature channels. Of course, you may also set both the amount of spatial size reduction and the amount of feature channel increase and have the device execute the process without using the device's adaptive function.

[0149] In concrete implementation, spatial domain resolution reduction processing may be performed based on spatial domain information or frequency domain information corresponding to image features, or spatial domain resolution reduction processing may be performed by a pre-configured convolutional layer. In this case, the step of obtaining the extended image features by reducing the spatial size corresponding to the image features of the image to be encoded in this embodiment and / or increasing the number of feature channels corresponding to the image features is: A step of obtaining an augmented image feature by reducing the spatial size corresponding to the image feature and / or increasing the number of feature channels corresponding to the image feature, based on spatial region information corresponding to the image feature of the image to be encoded. Or, A step of obtaining an augmented image feature by reducing the spatial size corresponding to the image feature and / or increasing the number of feature channels corresponding to the image feature, based on frequency domain information corresponding to the image feature of the image to be encoded. Or, The process may include a step of obtaining augmented image features by reducing the spatial size corresponding to the image features and / or increasing the number of feature channels corresponding to the image features, based on a pre-configured convolutional layer.

[0150] For specific embodiments, please refer to the explanation of Figure 5 in the above-described embodiment of the image decoding method; the explanation will be omitted here.

[0151] Step S200: The extended image features are grouped to obtain multiple extended feature groups.

[0152] In concrete implementation, when grouping augmented image features, grouping may be done by referring to the feature channel corresponding to each augmented image feature. In this case, step S200 of this embodiment is: expansion Image features The process includes grouping the enhanced image features based on the corresponding feature channels to obtain a plurality of enhanced feature groups.

[0153] expansion Image features Grouping augmented image features based on their corresponding feature channels and obtaining multiple augmented feature groups may also be done by uniformly dividing the augmented image features into multiple groups based on their corresponding feature channels. For example, if the total number of feature channels corresponding to an augmented image feature is 20, then augmented image features with corresponding feature channels 1 to 10 may be divided into one group, and augmented image features with corresponding feature channels 11 to 20 may be divided into another group. The number of groups to be uniformly divided may be predetermined by the administrator of the encoding device and is not limited in this embodiment.

[0154] Of course, when grouping specifically, the division may be uneven, and grouping augmented image features based on the feature channels corresponding to the augmented image features and obtaining multiple augmented feature groups may also be done by dividing the augmented image features into multiple groups based on the corresponding feature channels according to a pre-configured grouping rule, where the pre-configured grouping rule may be pre-configured by the administrator of the encoding device according to the actual needs, for example, the pre-configured grouping rule may be set to divide the first m / n (where n is the total number of feature channels, m is a pre-configured value, and the range is [1, n)) augmented image features into one group and divide the remaining augmented image features into another group.

[0155] In actual use, when grouping augmented image features based on the feature channels corresponding to the augmented image features and obtaining multiple augmented feature groups, it is also possible to divide an augmented image feature corresponding to one feature channel into one group. For example, if the total number of feature channels corresponding to the augmented image features is 20, then in this case, the augmented image features may be divided into 20 groups based on the differences in the feature channels.

[0156] Step S300: Perform residual calculations for each of the multiple extended feature groups and obtain image residual data corresponding to each extended feature group.

[0157] Performing residual calculations for each extended feature group and obtaining image residual data corresponding to the extended feature group may also be done by performing mean prediction for the extended feature group and subtracting the image features in the extended feature group from the predicted mean to obtain image residual data corresponding to the extended feature group.

[0158] Step S400: An image bitstream is generated based on the image residual data, and the image bitstream is transmitted to the image decoding side.

[0159] Generating an image bitstream based on image residual data may also involve writing the image residual data to the image bitstream using entropy coding.

[0160] In one possible embodiment of this embodiment, before performing spatial domain resolution reduction on the image features corresponding to the image to be encoded, the image features may be grouped and the spatial domain resolution reduction may be performed on each of them, in which case step S100 of this embodiment is, The steps include obtaining image features corresponding to the image to be encoded, The steps include: grouping the image features based on the feature channels corresponding to the image features and obtaining at least one group of image features; The process may also include the step of performing a spatial domain resolution reduction process on the data in at least one image feature group to obtain augmented image features.

[0161] In one possible embodiment of the present invention, performing spatial domain resolution reduction processing on image features corresponding to the image to be encoded and obtaining augmented image features may also be performed by first obtaining image features corresponding to the image to be encoded, and then performing spatial domain resolution reduction processing on the data in the image features and obtaining augmented image features. Obtaining image features corresponding to the image to be encoded may also be done by extracting image features corresponding to the image to be encoded via a pre-configured feature extraction network. When grouping data based on feature channels corresponding to image features, the same or similar method as when grouping augmented image features may be employed.

[0162] In one possible embodiment of this embodiment, step S400 of this embodiment is Image residuals corresponding to each extended feature group data The steps include performing spatial domain resolution expansion processing on the image to be encoded and obtaining image residual data corresponding to the image to be encoded, The process may also include the steps of generating an image bitstream based on image residual data corresponding to the image to be encoded, and transmitting the image bitstream to the image decoding side.

[0163] The spatial domain resolution expansion process may be the reverse process of the spatial domain resolution reduction process described above, and the image residuals corresponding to each expanded feature group. data After obtaining the image residuals corresponding to each extended feature group, data The spatial domain resolution can be expanded, and the corresponding spatial size and number of channels can be restored to match the image features corresponding to the image to be encoded, thereby obtaining image residual data corresponding to the image to be encoded. Subsequently, entropy encoding is performed on the image residual data corresponding to the image to be encoded to generate an image bitstream, and the generated image bitstream is sent to the image decoding side.

[0164] This embodiment performs spatial resolution reduction processing on image features corresponding to the image to be encoded, obtains extended image features, groups the extended image features to obtain multiple extended feature groups, performs residual calculation for each of the multiple extended feature groups, obtains image residual data corresponding to each extended feature group, generates an image bitstream based on the image residual data, and transmits the image bitstream to the image decoding side. By obtaining image features of the image to be encoded, performing spatial resolution reduction processing on them, and further grouping them into multiple extended feature groups, residual calculation processing can be performed on a group-by-group basis by grouping at a low resolution, improving the overall residual reconstruction calculation efficiency and reducing time complexity.

[0165] Figure 11 is a schematic flowchart of a second embodiment of the image coding method of the present invention.

[0166] Based on the first embodiment described above, step S300 of the image encoding method of this embodiment includes the following steps.

[0167] Step S3001: Construct a residual calculation sequence based on the multiple extended feature groups.

[0168] After multiple extended feature groups have been separated, residual calculations can be performed on the entire group. In this case, a residual calculation sequence may be constructed based on the multiple extended feature groups to determine the order in which residual calculations are performed for each extended feature group.

[0169] Step S3002: Based on the residual calculation sequence, residual calculation is performed for each of the multiple extended feature groups, and image residual data corresponding to each extended feature group is obtained.

[0170] Performing residual calculations for each of multiple extended feature groups based on a residual calculation sequence may also be performed sequentially for each of multiple extended feature groups based on the sequence order in the residual calculation sequence.

[0171] In actual use, residual calculations may be performed sequentially using a sequence traverse method, in which case step S3002 of this embodiment is, The steps include traversing the residual calculation sequence and obtaining the current extended feature group, The steps include obtaining auxiliary information output from the auxiliary coding network, The steps include constructing preliminary information based on the aforementioned auxiliary information, The steps include: performing residual calculations for the current extended feature group based on the prior information and obtaining image residual data corresponding to the current extended feature group; The process may include the step of obtaining image residual data corresponding to each extended feature group once the traverse is complete.

[0172] Traversing the residual calculation sequence to obtain the current extended feature group may also be interpreted as traversing the residual calculation sequence and making the traversed extended feature group the current extended feature group. The auxiliary coding network may be the auxiliary network shown in Figure 3 above (Hyper Encoder Net or Hyper Decoder Net).

[0173] In actual use, performing residual calculations for the current extended feature group based on prior information and obtaining image residual data corresponding to the current extended feature group may be done by processing the prior information using Prediction Fusion Net to obtain predicted mean values, subtracting the features in the current extended feature group from the predicted mean values ​​to perform residual calculations and obtain image residual data corresponding to the current extended feature group.

[0174] When constructing prior information based on auxiliary information, in order to improve the accuracy of mean prediction, prior information may be constructed by combining image features corresponding to extended feature groups for which residuals have already been calculated. The specific method is the same as the method applied in the image decoding process, and the specific implementation steps can refer to the method for constructing prior information based on auxiliary information relating to any of the above-described examples of image decoding methods.

[0175] This embodiment constructs a residual calculation sequence based on multiple extended feature groups, performs residual calculations for each of the multiple extended feature groups based on the residual calculation sequence, and obtains image residual data corresponding to each extended feature group. Since the residual calculation sequence is constructed based on multiple extended feature groups and the order of residual calculations can be determined by the residual calculation sequence, it is possible to quickly determine whether or not a pre-calculated extended feature group exists, and if a pre-calculated extended feature group exists, calculation Completed extensions Features Group image Residual data Based on the convolution results corresponding to this, more accurate prior information can be constructed.

[0176] Furthermore, embodiments of the present invention provide a storage medium in which an image decoding and / or image encoding program is stored, wherein when the image decoding program is executed by a processor, the above-described image decoding method is performed, and when the image encoding program is executed by a processor, the above-described image encoding method is performed.

[0177] Figure 12 is a structural block diagram of a first embodiment of the image decoding device of the present invention.

[0178] As shown in Figure 12, the image decoding device according to an embodiment of the present invention is A bitstream decoding module 10 for extracting image residual data or extended residual data from an image bitstream, and for obtaining multiple extended residual groups based on the extracted image residual data or extended residual data, A residual reconstruction module 20 performs residual reconstruction on each of the aforementioned multiple extended residual groups and obtains image reconstruction features corresponding to each extended residual group. A data combination module 30 for performing spatial domain resolution expansion processing on image reconstruction features corresponding to each expanded residual group and obtaining reconstruction feature data, The system includes an image reconstruction module 40 for performing image reconstruction based on the aforementioned reconstruction feature data and obtaining a reconstructed image block.

[0179] This embodiment extracts image residual data or extended residual data from an image bitstream, obtains multiple extended residual groups based on the extracted image residual data or extended residual data, performs residual reconstruction on each of the multiple extended residual groups, obtains image reconstruction features corresponding to each extended residual group, performs spatial domain resolution expansion processing on the image reconstruction features corresponding to each extended residual group, obtains reconstruction feature data, performs image reconstruction based on the reconstruction feature data, and obtains a reconstructed image block. Since the obtained extended residual data is residual data that has undergone spatial domain resolution reduction processing, it can be grouped at a low resolution and residual reconstruction processing can be performed on a group basis, improving the overall residual reconstruction computation efficiency and reducing time complexity.

[0180] In one possible embodiment of this embodiment, the bitstream decoding module 10 is further used to extract augmented residual data from the image bitstream and to group the augmented residual data based on the feature channels corresponding to the augmented residual data to obtain a plurality of augmented residual groups.

[0181] In one possible embodiment of this embodiment, the bitstream decoding module 10 further extracts image residual data from the image bitstream, performs a spatial domain resolution reduction process on the image residual data to obtain augmented residual data, the spatial domain resolution expansion process is the reverse process of the spatial domain resolution reduction process and is used to group the augmented residual data to obtain a plurality of augmented residual groups.

[0182] In one possible embodiment of this embodiment, the bitstream decoding module 10 is further used to perform data grouping based on feature channels corresponding to the image residual data to obtain at least one image residual group, and to perform spatial domain resolution reduction processing on the data in the image residual group to obtain augmented residual data.

[0183] In one possible embodiment of this example, the bitstream decoding module 10 is further used to obtain extended residual data by reducing the spatial size corresponding to the image residual data and / or increasing the number of feature channels corresponding to the image residual data.

[0184] In one possible embodiment of this example, the bitstream decoding module 10 is further used to obtain extended residual data by reducing the spatial size corresponding to the image residual data and / or increasing the number of feature channels corresponding to the image residual data based on the spatial region information corresponding to the image residual data, or by reducing the spatial size corresponding to the image residual data and / or increasing the number of feature channels corresponding to the image residual data based on the frequency region information corresponding to the image residual data, or by reducing the spatial size corresponding to the image residual data and / or increasing the number of feature channels corresponding to the image residual data based on a preset convolutional layer.

[0185] In one possible embodiment of this example, the residual restoration module 20 is further used to construct a residual restoration sequence based on the plurality of extended residual groups, and perform residual restoration on each of the plurality of extended residual groups based on the residual restoration sequence to obtain image reconstruction features corresponding to each extended residual group.

[0186] In one possible embodiment of this example, the residual restoration module 20 is further used to traverse the residual restoration sequence, obtain the current extended residual group, obtain the auxiliary information output from the auxiliary encoding network, construct prior information based on the auxiliary information, perform residual restoration on the current extended residual group based on the prior information, obtain image reconstruction features corresponding to the current extended residual group, and when the traversal is completed, is used to obtain image reconstruction features corresponding to each extended residual group.

[0187] In one possible embodiment of this embodiment, the residual restoration module 20 further obtains extended auxiliary information, detects whether the current extended residual group is the first element in the residual restoration sequence, and if it is the first element, constructs prior information based on the extended auxiliary information; if it is not the first element, combines the extended auxiliary information with the convolution processing result corresponding to the image reconstruction feature of the restored extended residual group to obtain combined auxiliary information, and is used to construct prior information based on the combined auxiliary information.

[0188] In one possible embodiment of this embodiment, the residual restoration module 20 further obtains the image reconstruction feature corresponding to the restored extended residual group, performs feature enhancement on the image reconstruction feature to obtain an enhanced reconstruction feature, and combines the auxiliary information with the convolution processing result corresponding to the enhanced reconstruction feature to obtain combined auxiliary information.

[0189] In one possible embodiment of this embodiment, the residual restoration module 20 further obtains the predicted average value, auxiliary information, image residual data, and / or residual data variance corresponding to the restored extended residual group, and performs feature enhancement on the image reconstruction feature based on the predicted average value, auxiliary information, image residual data, and / or residual data variance corresponding to the restored extended residual group to obtain an enhanced reconstruction feature.

[0190] In one possible embodiment of this embodiment, the data combination module 30 further performs feature enhancement on the image reconstruction feature corresponding to each extended residual group to obtain the enhanced reconstruction feature corresponding to each extended residual group, and performs spatial region resolution enlargement processing on the enhanced reconstruction feature corresponding to each extended residual group to obtain reconstruction feature data.

[0191] In one possible embodiment of this embodiment, the image reconstruction module 40 is further used to perform a synthetic transformation process on the reconstruction feature data using a pre-constructed synthetic transformation network to achieve the image reconstruction and to obtain the reconstructed image block, wherein the synthetic transformation network is a network constructed based on deep learning or a neural network.

[0192] In one possible embodiment of this embodiment, the bitstream decoding module 10 is further used to uniformly group the augmented residual data based on feature channels corresponding to the augmented residual data.

[0193] Figure 13 is a structural block diagram of the first embodiment of the image encoding device of the present invention.

[0194] As shown in Figure 13, the image encoding device according to an embodiment of the present invention is A feature extraction module 100 performs spatial domain resolution reduction processing on image features corresponding to the image to be encoded, and obtains augmented image features, A data grouping module 200 for grouping the aforementioned extended image features and obtaining multiple extended feature groups, A residual calculation module 300 for performing residual calculations for each of the aforementioned multiple extended feature groups and obtaining image residual data corresponding to each extended feature group, The system includes a bitstream generation module 400 for generating an image bitstream based on the image residual data and transmitting the image bitstream to the image decoding side.

[0195] This embodiment performs spatial resolution reduction processing on image features corresponding to the image to be encoded, obtains extended image features, groups the extended image features to obtain multiple extended feature groups, performs residual calculation for each of the multiple extended feature groups, obtains image residual data corresponding to each extended feature group, generates an image bitstream based on the image residual data, and transmits the image bitstream to the image decoding side. By obtaining image features of the image to be encoded, performing spatial resolution reduction processing on them, and further grouping them into multiple extended feature groups, residual calculation processing can be performed on a group-by-group basis by grouping at a low resolution, improving the overall residual reconstruction calculation efficiency and reducing time complexity.

[0196] In one possible embodiment of this embodiment, the feature extraction module 100 is further used to acquire image features corresponding to the image to be encoded, perform spatial domain resolution reduction processing on the data in the image features, and acquire augmented image features.

[0197] In one possible embodiment of this embodiment, the feature extraction module 100 is further used to obtain augmented image features by reducing the spatial size corresponding to the image features of the image to be encoded and / or increasing the number of feature channels corresponding to the image features.

[0198] In one possible embodiment of this embodiment, the feature extraction module 100 is further used to obtain augmented image features by reducing the spatial size corresponding to the image features and / or increasing the number of feature channels corresponding to the image features, based on spatial domain information corresponding to the image features of the image to be encoded; or by reducing the spatial size corresponding to the image features and / or increasing the number of feature channels corresponding to the image features, based on frequency domain information corresponding to the image features of the image to be encoded; or by reducing the spatial size corresponding to the image features and / or increasing the number of feature channels corresponding to the image features, based on a preset convolutional layer.

[0199] In one possible embodiment of this embodiment, the residual calculation module 300 is further used to construct a residual calculation sequence based on the plurality of extended feature groups, to perform residual calculations for each of the plurality of extended feature groups based on the residual calculation sequence, and to obtain image residual data corresponding to each extended feature group.

[0200] In one possible embodiment of this embodiment, the residual calculation module 300 is further used to traverse the residual calculation sequence, acquire the current extended feature group, acquire auxiliary information output from the auxiliary coding network, construct prior information based on the auxiliary information, perform residual calculation on the current extended feature group based on the prior information, acquire image residual data corresponding to the current extended feature group, and, once the traverse is complete, acquire image residual data corresponding to each extended feature group.

[0201] In one possible embodiment of this embodiment, the data grouping module 200 further extends Image features This is used to group the extended image features based on the corresponding feature channels and to obtain multiple extended feature groups.

[0202] In one possible embodiment of this example, the bitstream generation module 400 further includes an image residual corresponding to each extended feature group data to perform spatial domain resolution enlargement processing, obtain image residual data corresponding to the image to be encoded, where the spatial domain resolution enlargement processing is the reverse process of the spatial domain resolution reduction processing, generate an image bitstream based on the image residual data corresponding to the image to be encoded, and use the image bitstream for transmission to the image decoding side.

[0203] Note that the above is merely an example and does not limit the technical solution of the present invention. In specific applications, those skilled in the art can set it as needed, and the present invention is not limited thereto.

[0204] Note that the above-described flow is merely exemplary and does not limit the protection scope of the present invention. In actual applications, those skilled in the art can select some or all of them according to actual needs to achieve the purpose of the technical solution of this example, and are not limited here.

[0205] Also, for the technical details not described in detail in this example, reference may be made to the image decoding method or image encoding method according to any embodiment of the present invention, and the description is omitted here.

[0206] Note that in this specification, the term "comprising", "containing" or any other variation thereof is intended to include non-exclusive inclusion, so that a process, method, article or system containing a series of elements includes not only those elements but also other elements not explicitly listed, or also includes elements specific to such a process, method, article or system. Without more limitations, an element limited by "comprising one..." does not exclude the existence of other same elements in the process, method, article or system containing the element.

[0207] The numbers of the above embodiments of the present invention are only for the purpose of explanation and do not represent the superiority of the embodiments.

[0208] From the above description of the embodiments, those skilled in the art will clearly understand that the methods of the above embodiments can be implemented by software and the necessary general-purpose hardware platform, and of course by hardware, but in many cases the former is a more preferred embodiment. Based on this understanding, the technical proposal of the present invention may be embodied in the form of a software product, the essential or prior art-contributing portion of which is stored in a storage medium (e.g., read-only memory (ROM) / RAM, magnetic disk, optical disk) and includes several instructions for a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to perform the method of each embodiment of the present invention.

[0209] The foregoing are merely preferred embodiments of the present invention and do not limit the scope of the present invention. Equivalent structural and process transformations, or direct or indirect applications to other related technical fields, made using the contents of the specification and drawings of the present invention are all equally within the scope of protection of the present invention. [Explanation of Symbols]

[0210] 10-bit stream decoding module 20 Residual Recovery Module 30 modules 40 Image Reconstruction Modules 100 Feature Extraction Module 200 Data Grouping Modules 300 Residual Calculation Module 400 bitstream generation module 1001 Processor 1002 Communications Bus 1003 User Interface 1004 Network Interface 1005 memory

Claims

1. The steps include: extracting image residual data or extended residual data from an image bitstream, and obtaining a plurality of extended residual groups based on the extracted image residual data or extended residual data; The steps include performing residual reconstruction for each of the aforementioned multiple extended residual groups and obtaining image reconstruction features corresponding to each extended residual group, The steps include performing spatial domain resolution expansion processing on the image reconstruction features corresponding to each expanded residual group and obtaining the reconstruction feature data, The process includes the step of performing image reconstruction based on the aforementioned reconstruction feature data and obtaining a reconstructed image block. An image decoding method characterized by the following:

2. The steps of extracting image residual data or extended residual data from the image bitstream, and obtaining a plurality of extended residual groups based on the extracted image residual data or extended residual data, are as follows: The steps include extracting the image residual data from the image bitstream, A step of performing a spatial domain resolution reduction process on the aforementioned image residual data and obtaining the aforementioned expanded residual data, wherein the spatial domain resolution expansion process is the reverse process of the spatial domain resolution reduction process, The steps include: grouping the extended residual data to obtain a plurality of extended residual groups, The image decoding method according to feature 1.

3. The step of performing a spatial domain resolution reduction process on the aforementioned image residual data and obtaining the aforementioned expanded residual data is: The step of acquiring the expanded residual data by reducing the spatial size corresponding to the image residual data and / or increasing the number of feature channels corresponding to the image residual data, The image decoding method according to feature 2.

4. The step of acquiring the expanded residual data by reducing the spatial size corresponding to the image residual data and / or increasing the number of feature channels corresponding to the image residual data is: The step of acquiring the expanded residual data is to reduce the spatial size corresponding to the image residual data and / or increase the number of feature channels corresponding to the image residual data, based on spatial region information corresponding to the image residual data. The image decoding method according to feature 3.

5. The step of performing residual reconstruction for each of the aforementioned multiple extended residual groups and obtaining image reconstruction features corresponding to each extended residual group is: The steps include constructing a residual recovery sequence based on the aforementioned multiple extended residual groups, The process includes the step of performing residual reconstruction for each of the plurality of extended residual groups based on the residual reconstruction sequence, and obtaining image reconstruction features corresponding to each extended residual group. The image decoding method according to feature 1.

6. The step of performing residual reconstruction for each of the multiple extended residual groups based on the residual reconstruction sequence and obtaining image reconstruction features corresponding to each extended residual group is: The steps include traversing the residual restoration sequence and obtaining the current expanded residual group, The steps include obtaining auxiliary information output from the auxiliary coding network, The steps include constructing preliminary information based on the aforementioned auxiliary information, The steps include: performing residual reconstruction on the current extended residual group based on the prior information and obtaining image reconstruction features corresponding to the current extended residual group; The traverse is completed, and the step includes obtaining the image reconstruction features corresponding to each extended residual group, The image decoding method according to feature 5.

7. The step of constructing preliminary information based on the aforementioned auxiliary information is: Steps to obtain extended auxiliary information, The steps include detecting whether the current expanded residual group is the first element in the residual restoration sequence, If it is the first element, the step is to construct preliminary information based on the aforementioned extended auxiliary information, If it is not the first element, the process includes the steps of combining the extended auxiliary information and the convolutional processing result corresponding to the image reconstruction features of the restored extended residual group to obtain combined auxiliary information, and constructing prior information based on the combined auxiliary information, The image decoding method according to feature 6.

8. The step of performing image reconstruction based on the aforementioned reconstructed feature data and obtaining a reconstructed image block is: The process includes the steps of performing a synthetic transformation process on the reconstructed feature data using a pre-constructed synthetic transformation network to achieve the image reconstruction and obtaining the reconstructed image block, The aforementioned synthetic transformation network is a network constructed based on deep learning or a neural network. The image decoding method according to feature 1.

9. The step of grouping the aforementioned extended residual data to obtain multiple extended residual groups is: The step includes uniformly grouping the expanded residual data based on the feature channels corresponding to the expanded residual data, The image decoding method according to feature 2.

10. The steps include: performing spatial domain resolution reduction on image features corresponding to the image to be encoded to obtain augmented image features; The steps include grouping the aforementioned extended image features and obtaining multiple extended feature groups, The steps include performing residual calculations for each of the aforementioned multiple extended feature groups and obtaining image residual data corresponding to each extended feature group, The process includes the steps of generating an image bitstream based on the image residual data and transmitting the image bitstream to the image decoding side. An image encoding method characterized by the following.

11. A bitstream decoding module for extracting image residual data or extended residual data from an image bitstream, and for obtaining multiple extended residual groups based on the extracted image residual data or extended residual data, A residual reconstruction module for performing residual reconstruction on each of the aforementioned multiple extended residual groups and obtaining image reconstruction features corresponding to each extended residual group, A data combination module for obtaining reconstructed feature data by performing spatial domain resolution expansion processing on image reconstruction features corresponding to each expanded residual group, Includes an image reconstruction module for performing image reconstruction based on the aforementioned reconstruction feature data and obtaining a reconstructed image block, An image decoding device characterized by the following features.

12. The bitstream decoding module further extracts the image residual data from the image bitstream, performs a spatial domain resolution reduction process on the image residual data, obtains the expanded residual data, and the spatial domain resolution expansion process is the reverse process of the spatial domain resolution reduction process, and is used to group the expanded residual data to obtain multiple expanded residual groups. The image decoding device according to feature 11.

13. The bitstream decoding module is further used to acquire the expanded residual data by reducing the spatial size corresponding to the image residual data and / or increasing the number of feature channels corresponding to the image residual data. The image decoding device according to feature 12.

14. The bitstream decoding module is further used to acquire the expanded residual data by reducing the spatial size corresponding to the image residual data and / or increasing the number of feature channels corresponding to the image residual data, based on spatial region information corresponding to the image residual data. The image decoding device according to feature 13.

15. The bitstream decoding module is further used to construct a residual reconstruction sequence based on the plurality of extended residual groups, perform residual reconstruction for each of the plurality of extended residual groups based on the residual reconstruction sequence, and obtain image reconstruction features corresponding to each extended residual group. The image decoding device according to feature 11.

16. The residual reconstruction module further traverses the residual reconstruction sequence, obtains the current extended residual group, obtains auxiliary information output from the auxiliary coding network, constructs prior information based on the auxiliary information, performs residual reconstruction on the current extended residual group based on the prior information, obtains image reconstruction features corresponding to the current extended residual group, and, upon completion of the traverse, is used to obtain image reconstruction features corresponding to each extended residual group. The image decoding device according to feature 15.

17. The residual reconstruction module further acquires augmented auxiliary information, detects whether the current augmented residual group is the first element in the residual reconstruction sequence, constructs prior information based on the augmented auxiliary information if it is the first element, and if it is not the first element, combines the augmented auxiliary information with the convolutional processing results corresponding to the image reconstruction features of the reconstructed augmented residual group to acquire combined auxiliary information, and uses the combined auxiliary information to construct prior information. The image decoding device according to feature 16.

18. The image reconstruction module is further used to perform a synthetic transformation process on the reconstructed feature data using a pre-constructed synthetic transformation network to achieve the image reconstruction and to obtain the reconstructed image block, wherein the synthetic transformation network is a network constructed based on deep learning or a neural network. The image decoding device according to feature 11.

19. The bitstream decoding module is further used to uniformly group the augmented residual data based on the feature channels corresponding to the augmented residual data. The image decoding device according to feature 12.

20. A feature extraction module for obtaining augmented image features by performing spatial domain resolution reduction processing on image features corresponding to the image to be encoded, A data grouping module for grouping the aforementioned extended image features and obtaining multiple extended feature groups, A residual calculation module for performing residual calculations for each of the aforementioned multiple extended feature groups and obtaining image residual data corresponding to each extended feature group, A bitstream generation module for generating an image bitstream based on the image residual data and transmitting the image bitstream to the image decoding side is included. An image coding device characterized by the following:

21. A decoding device comprising a processor, memory, and an image decoding program stored in the memory and executable on the processor, wherein when the image decoding program is executed by the processor, the image decoding method according to any one of claims 1 to 9 is performed. A decoding device characterized by the following features.

22. An encoding device comprising a processor, memory, and an image decoding program and / or image encoding program stored in the memory and executable on the processor, wherein when the image decoding program is executed by the processor, the image decoding method according to any one of claims 1 to 9 is performed, and when the image encoding program is executed by the processor, the image encoding method according to claim 10 is performed. An encoding device characterized by the following features.

23. A storage medium storing an image decoding program and / or an image encoding program, wherein when the image decoding program is executed, the image decoding method described in any one of claims 1 to 9 is performed, and when the image encoding program is executed, the image encoding method described in claim 10 is performed. A storage medium characterized by the following features.

24. A computer program, configured such that when executed by a processor having memory, it performs the image decoding method described in any one of claims 1 to 9, or the image encoding method described in claim 10. A computer program characterized by the following features.

25. A computer program product including computer program instructions, wherein when the computer program instructions are executed by a processor having memory, the image decoding method described in any one of claims 1 to 9 or the image encoding method described in claim 10 is performed. A computer program product characterized by the following features.

Citation Information

Patent Citations

  • Inter prediction method and apparatus

    CN113556567A

  • Image reconstruction method, image coding and decoding method and related equipment

    CN114943643A

  • Video reproducing method, video reproducing device, and video reproducing program

    JP2012238947A

  • Super-resolution image reconstruction method based on deep convolutional sparse coding

    US20220284547A1