Image compression reconstruction method, system and device using packet token mixer, medium and product

A group token mixer-based entropy model addresses the computational inefficiencies of existing image compression models by decomposing global attention into cross-group and intra-group components, enhancing accuracy and efficiency in entropy estimation.

CN120321400APending Publication Date: 2025-07-15HARBIN INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510465981.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing entropy model has high computational complexity and is difficult to achieve efficient processing in practical applications, which affects the efficiency of image compression.

Method used

The entropy model is constructed using the group token mixer, and the inter-group and in-group information are extracted through the cross-group token mixer and the in-group token mixer respectively, and the entropy encoding and entropy decoding processes are optimized to reduce the computational complexity.

Benefits of technology

It improves the accuracy and processing efficiency of entropy estimation, reduces the computational complexity, and improves the speed of image compression and parameter utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120321400A_ABST
    Figure CN120321400A_ABST
Patent Text Reader

Abstract

The invention discloses an image compression reconstruction method, system and device using a packet token mixer, a medium and a product, and relates to the field of image processing.The method comprises the steps that an analysis transformation network is adopted to map an original image, and hidden layer space features are obtained; quantizing the hidden layer space features to obtain quantized features; constructing an entropy model based on the packet token mixer; entropy coding and entropy decoding are carried out on the quantized features based on the entropy model, and reconstructed features are obtained; and based on the reconstruction features, reconstructing an image by using a generative network. According to the method, the multi-level dependency relationship in the image data can be effectively captured and modeled, the entropy estimation accuracy is improved, and the calculation complexity is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing, and in particular to an image compression and reconstruction method, system, device, medium and product using a grouped token mixer. Background Art

[0002] With the rapid development of deep learning, lossy image compression with end-to-end optimization has made significant progress in the past few years. Currently, the state-of-the-art learning-based methods have surpassed traditional compression methods such as VVC in terms of performance, indicating a bright future for learning-based image compression. Most learning-based image compression methods are based on the autoencoder framework. The original input image is mapped into a latent representation through a non-linear transformation and then quantized into discrete values. These discrete hidden layer variables are encoded using an arithmetic coding tool based on the probability distribution estimated by an entropy model. The final bitrate is the cross-entropy between the prior distribution and the estimated distribution of the hidden layer variables. This shows that more accurate estimation can reduce the number of bits required to compress the image, which has motivated researchers to design more powerful entropy models. Existing entropy models have high computational complexity, and with the increase in data dimension and complexity, the computational cost of existing entropy models has increased significantly, making it difficult to achieve efficient processing in practical applications. Summary of the Invention

[0003] To solve the above problems, the present application provides an image compression and reconstruction method, system, device, medium and product using a grouped token mixer.

[0004] To achieve the above object, the present application provides the following solutions:

[0005] In a first aspect, the present application provides an image compression and reconstruction method using a grouped token mixer, including:

[0006] Mapping the original image using an analysis transformation network to obtain hidden layer space features;

[0007] Quantizing the hidden layer space features to obtain quantized features;

[0008] Constructing an entropy model based on a grouped token mixer; the entropy model includes a grouping module, a mapping network, a plurality of grouped token mixers and a parameter network, and the grouped token mixer includes a cross-group token mixer and an intra-group token mixer; the cross-group token mixer is used to extract inter-group context information, and the intra-group token mixer is used to extract intra-group information;

[0009] Performing entropy encoding and entropy decoding on the quantized features based on the entropy model to obtain reconstructed features;

[0010] Reconstructing an image using a generation network based on the reconstructed features.

[0011] In a second aspect, the present application provides an image compression and reconstruction system using a grouped token mixer, including:

[0012] A mapping unit, configured to map an original image using an analysis transformation network to obtain hidden layer space features;

[0013] A quantization unit, configured to quantize the hidden layer space features to obtain quantized features;

[0014] An entropy model construction unit, configured to construct an entropy model based on a grouped token mixer; the entropy model includes a grouping module, a mapping network, a plurality of grouped token mixers, and a parameter network, and the grouped token mixer includes a cross-group token mixer and an intra-group token mixer; the cross-group token mixer is configured to extract inter-group context information, and the intra-group token mixer is configured to extract intra-group information;

[0015] An entropy encoding and entropy decoding unit, configured to perform entropy encoding and entropy decoding on the quantized features based on the entropy model to obtain reconstructed features;

[0016] An image reconstruction unit, configured to reconstruct an image using a generation network based on the reconstructed features.

[0017] In a third aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the computer program to implement the above-mentioned image compression and reconstruction method using a grouped token mixer.

[0018] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the above-mentioned image compression and reconstruction method using a grouped token mixer.

[0019] In a fifth aspect, the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the above-mentioned image compression and reconstruction method using a grouped token mixer.

[0020] According to the specific embodiments provided by the present application, the present application has the following technical effects:

[0021] The present application provides an image compression and reconstruction method, system, device, medium, and product using a grouped token mixer. Based on the entropy model constructed by the grouped token mixer, it can effectively capture and model multi-level dependencies in image data, improving the accuracy of entropy estimation. At the same time, the design of the grouped token mixer optimizes the calculation process, reduces the computational complexity, and improves the processing efficiency and parameter utilization rate of the entropy model. Description of the Drawings

[0022] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0023] Figure 1 FIG. is a schematic flowchart of an image compression and reconstruction method using a grouped token mixer provided in an embodiment of the present application;

[0024] Figure 2 FIG. is a schematic block diagram of an image compression and reconstruction method using a grouped token mixer provided in an embodiment of the present application;

[0025] Figure 3 FIG. is a schematic diagram of a spatial division example;

[0026] Figure 4 FIG. is a schematic diagram of an entropy model structure;

[0027] Figure 5 FIG. is a schematic diagram of a cross-group token mixer structure;

[0028] Figure 6 FIG. is a schematic diagram of an intra-group token mixer structure;

[0029] Figure 7 FIG. is a schematic diagram of a computer device provided in an embodiment of the present application. Detailed implementation manners

[0030] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0031] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.

[0032] In an exemplary embodiment, as Figure 1 - Figure 2 shown, an image compression and reconstruction method using a grouped token mixer is provided. This method is executed by a computer device, and specifically can be executed alone by a computer device such as a terminal or a server, or can be jointly executed by a terminal and a server. In the embodiments of the present application, taking this method applied to a server as an example for illustration, it includes the following steps 201 to step 208. Among them:

[0033] S1: Map the original image using an analysis transformation network to obtain hidden layer space features.

[0034] As Figure 2 shown, at the encoding end, after the original image is mapped by the analysis transform network, hidden layer space features with downsampling in the spatial dimension and expansion in the channel dimension are obtained.

[0035] S2: Quantize the hidden layer space features to obtain quantized features.

[0036] S3: Construct an entropy model based on a grouped token mixer.

[0037] As Figure 4 shown, the entropy model includes a grouping module, a mapping network, multiple grouped token mixers, and a parameter network. The grouped token mixer includes a cross-group token mixer and an intra-group token mixer. The cross-group token mixer is used to extract inter-group context information, and the intra-group token mixer is used to extract intra-group information.

[0038] As Figure 5 shown, the cross-group token mixer includes a three-dimensional relative position embedding generator, a masked self-attention module, a feed-forward network, and two layers of normalization layers. As Figure 6 shown, the intra-group token mixer includes a position embedding generator, a self-attention module, a feed-forward network, and two layers of normalization layers.

[0039] S4: Perform entropy encoding and entropy decoding on the quantized features based on the entropy model to obtain reconstructed features. Specifically, it includes: dividing the quantized features into multiple groups of hidden layer variables through the grouping module; mapping the multiple groups of hidden layer variables to the same shape through the mapping network; extracting features from the multiple groups of hidden layer variables through multiple grouped token mixers; predicting probability distribution parameters for the extracted features through the parameter network; encoding the multiple groups of hidden layer variables based on the predicted probability distribution parameters to obtain a binary code stream; and decoding and recombining the binary code stream to obtain reconstructed features.

[0040] In this implementation, the entropy model uses context information to predict probability distribution parameters, and encodes according to the predicted probability distribution parameters into a binary code stream. As Figure 2 shown, at the decoding end, the decoder also executes the entropy model to predict probability distribution parameters, and decodes from the binary code stream based on the probability distribution parameters to obtain reconstructed hidden layer features.

[0041] In this implementation, the grouping module in the entropy model first divides the quantized features into g groups of hidden layer variables where Denotes the index set of the i-th group of hidden layer variables, and each group represents satisfying where h i , w i , c i are the height, width of the group, and the channel size of the i-th group respectively. The entropy model of this embodiment decomposes the probability distribution parameters of the hidden layer variables through inter-group autoregression, and its expression is as follows:

[0042]

[0043] That is, by using the known context to complete the prediction of the probability distribution parameters of the current group of hidden layer variables . During training, a total of g - 1 groups of hidden layer variables are input into the network. First, the hidden layer variables are mapped to the same shape through the mapping network Then, after passing through N grouped token mixers to globally extract context information and perform inter-group autoregressive prediction. Each grouped token mixer is divided into a cross-group token mixer and an intra-group token mixer, where the cross-group token mixer realizes the extraction of inter-group context information, and the intra-group token mixer realizes the extraction of intra-group information. After passing through multiple grouped token mixers, the extracted feature is obtained and concatenated with another learnable feature to obtain the autoregressive prior feature, and then this feature is fed into the parameter network to obtain the probability distribution parameters For example, if it is assumed that the hidden layer conforms to a single Gaussian distribution, then The encoding and decoding processes are different from the training process. For the entropy model based on inter-group autoregression, during encoding and decoding, it is not possible to complete one-step inference like during training, and it is necessary to predict the distribution of the hidden layer representation group by group. When predicting , its corresponding autoregressive prior feature directly uses e, and inputting it into the parameter network can obtain the distribution.

[0044] Regarding the grouping process, the process is to divide the quantized features along the channel and spatial dimensions respectively. An example of a simple channel segmentation scheme is to divide the quantized features into k c equal-sized groups. This embodiment introduces two examples of spatial segmentation schemes, which respectively support 2-step and 4-step spatial autoregressive processes. In the 2-step segmentation scheme, the quantized features are segmented in a checkerboard pattern, which means k h = 1, k w = 2, as shown in (a) of Figure 3 . In the 4-step segmentation scheme, k h = 2, k w = 2 is selected, as shown in Figure 3Perform the segmentation of the quantization features as shown in (b) therein. Through these spatial segmentation schemes, the entropy model can utilize the bidirectional spatial context information to form an effective guidance of the hidden layer variables of the previous group to the subsequent group. In the following part, k c ×k h k w is consistently used to represent the method of quantization feature segmentation. This means that the quantization features are first segmented into k c slices along the channel dimension, and then each slice is further grouped through k h k w steps of spatial segmentation.

[0045] Regarding the grouped token mixer, in this embodiment, the attention calculation is not directly performed on all the decoded values in the previous group because this is computationally very intensive and usually infeasible. Instead, the interaction with the previous values is decomposed into two token mixers: the cross-group token mixer and the intra-group token mixer, which are used to model the global dependencies across groups and the intra-group dependencies respectively. Given the grouped and mapped input features where g' is the number of feature groups of the input network, a grouped token mixer can be expressed as:

[0046] Y l = Cross-group token mixer 1(X l )

[0047] Z1 = Intra-group token mixer 1(Y l )

[0048] where Y l and Z l are the intermediate feature and the output feature in the l-th layer module respectively. The entire grouped token mixer architecture is constructed by repeatedly stacking the cross-group token mixer and the intra-group token mixer.

[0049] The intra-group token mixer is designed to model the intra-group dependencies. Given the grouped representation of the input Apply the attention mechanism in the spatial dimension and share the weights in the group dimension. Specifically, Y is fed into the self-attention module to mix the information along the spatial dimension and generate the output feature In the attention mechanism, position embeddings are required to provide the information of the current position. Considering the difference in the image resolution between the training and test phases, a position embedding generator is introduced, which uses depth convolution to introduce the translational invariance inductive bias in the training phase. The process is as follows:

[0050] Y P = DWConv(Y) + Y

[0051] where DWConv is the depth convolution.

[0052] Then, the output Y of the position perception P is linearly projected onto the query, key, and value, and they are split into m representations for each head, generating where d h is the dimension of each head. The self-attention process in the spatial dimension can be expressed as follows:

[0053]

[0054] where, σ represents the softmax function. Then Y o is fed into the feed-forward network, and two linear layers with an expansion ratio of 4 are used to mix information in the channel dimension and generate the output features In this way, the intra-group token mixer in each group can fully mix in the spatial and channel dimensions to generate the output features.

[0055] The cross-group token mixer is used to exchange information across groups, enabling the entropy model to aggregate global spatial-channel information from the previously decoded values. The cross-group token mixer is implemented by applying a masked attention mechanism in the group dimension. During implementation, given the input features first, the group and spatial dimensions are permuted to obtain the rearranged intermediate features The intermediate feature X r undergoes a masked attention process to integrate the information in the group dimension while keeping the weights in the spatial dimension the same. The self-attention process is rewritten as:

[0056]

[0057] where, represents the relative position embedding, is the mask matrix, where the lower triangle is zero and the rest is -∞.

[0058] To encode and distinguish the position information of each group, the group dimension is regarded as a three-dimensional space of size (k c , k h , k w ). Each group index i is assigned a three-dimensional coordinate (x i , y i , z i ). For indices u, v, the relative position vector r u,v is calculated as r u,v =(x u -x v , y u -y v , z u -z v ), and each component of the vector (x u -xv , y u , -y v , z u , -z v ) range from [-k h + 1, k h - 1], [-k w + 1, k w - 1], [-k c + 1, k c - 1], and then mapped to a positive scalar s:

[0059]

[0060] The relative position embedding between u and v can be represented by an index as:

[0061]

[0062] where is a two-dimensional learnable parameter, initialized with a truncated normal distribution, n p = (2k c - 1)·(2k h - 1)·(2k w - 1). Similar to the intra-group token mixer, the output features are fed into a feed-forward network to further mix information along the channel dimension, generating output features Note that the cross-group process means that information is aggregated among accessible groups, but restricted to the same spatial location. By alternately stacking intra-group and cross-group token mixers, the entropy model can fully mix information in the spatial, group, and channel dimensions to capture global information.

[0063] Regarding cache acceleration during encoding and decoding, during the inference process, intra-group autoregression is used to estimate the probability distribution parameters of the hidden layer variables, which requires g inferences of the entropy model. However, due to the large scale of the entropy model and the computationally intensive autoregressive process, the inference speed is still a challenge. When performing inference on the autoregressive model, the activation values of the previous group have been calculated during the previous inference and can be further eliminated. Considering that only the cross-group token mixer interacts with the previously decoded group, the corresponding activation values need to be cached and used in subsequent inference stages. Therefore, cache the activation values of each group's keys and values generated by each cross-group token mixer during each inference.

[0064] During encoding and decoding When 1 < i ≤ g, instead of directly inputting the context of group i - 1 only the hidden layer features of the previous group are input Before performing the attention process, the keys and values With the activation value of the cache Cache_K t-1 , Cache_V t-1 are concatenated. Note that in the attention mechanism, the mask is cancelled because all concatenated decoded keys and values can be accessed in the current step. The self-attention process of the cross-group token mixer with context cache optimization in the t-th step is implemented as follows:

[0065]

[0066] Among them, the concatenated key-value satisfies With context cache optimization, only a single set of latent variables needs to be input to the network during each inference instead of all decoded groups Therefore, the inference process of the network is significantly accelerated, thereby improving the speed of image encoding and decoding.

[0067] Complexity analysis of the grouped token mixer:

[0068] In this application, a complete self-attention process is not directly applied to all tokens, and its computational complexity is:

[0069]

[0070] The global attention is decomposed into two parts: the intra-group token mixer and the cross-group token mixer. The computational complexity of the intra-group token mixer is The computational complexity of the cross-group token mixer is Through this decomposition, the total computational complexity is reduced to:

[0071]

[0072] Among them, the main computational and memory overhead comes from the self-attention process in the intra-group token mixer. This decomposition also enables this application to be extended to larger G values, configured as 10 and 40 in the implementation.

[0073] To better understand the computational complexity, assume that the image size is 768×512, reduced to a latent code of 32×48, and divided into 40 groups (10×4 configuration), resulting in h = 16, w = 24. The overall computational complexity of the module is:

[0074]

[0075] This shows the efficiency of the method in this application compared to directly applying the complete attention process, whose complexity is:

[0076]

[0077] S5: Based on the reconstructed features, use a generative network to reconstruct the image.

[0078] The reconstructed features are upsampled layer by layer after passing through the synthesis transform network to obtain the final reconstructed image.

[0079] The image compression and reconstruction method using a grouped token mixer provided in this embodiment has the following advantages:

[0080] (1) Sufficient dependency modeling: Make full use of spatial-channel redundancy, establish an entropy model with spatial channels, and apply cross-group and intra-group token mixers to achieve a global receptive field and model long-range dependencies.

[0081] (2) Low computational complexity, which is reflected in two aspects. On the one hand, for the design of the entropy model, that is, the hidden layer representation is entropy-coded layer by layer. There are a total of g groups, which is much smaller than the number of pixels hw in the middle hidden layer, and thus much smaller than the autoregressive steps of Entroformer and Contextformer. The fewer the autoregressive steps, the fewer the network inference times, the less the hidden layer entropy estimation time, and the faster the encoding and decoding speed. On the other hand, through decomposition, the global attention is decomposed into a cross-group token mixer and an intra-group token mixer, achieving a reduction in complexity.

[0082] (3) High parameter utilization rate. Multiple groups of distribution predictions share the same network. During the learning process, the commonalities between different group predictions are utilized to make full use of the parameters and avoid parameter redundancy.

[0083] This application also provides an application scenario that applies the above-mentioned image compression and reconstruction method using a grouped token mixer. Specifically: The image compression and reconstruction method using a grouped token mixer provided in this embodiment can be applied to the face image reconstruction scenario. Given a face image x, first pass the image through the analysis transform network to obtain the hidden layer representation y, and then quantize the hidden layer to obtain In the entropy model, first divide the quantized features into groups, Encode each of the g groups of hidden layer variables one by one. When encoding use e as the autoregressive prior feature and send it into the parameter network to obtain the probability distribution parameters and use an entropy coding tool such as arithmetic coding to perform lossless coding of into the binary code stream; when encoding map through the mapping network to send it into the grouped token mixer to extract the autoregressive prior feature. During this process, the intermediate features in the cross-group token mixer, including keys and values, need to be cached. Then, the generated by the grouped token mixer is sent into the parameter network to obtain the probability distribution parameters Subsequently, an entropy coding tool is used to encode it into a binary bitstream. Then the bitstream is transmitted to the decoding end.

[0084] At the decoding end, the entropy coding runs the same process as at the decoding end to obtain the probability prediction of the corresponding group and applies an entropy decoding tool to decode it. After decoding all groups recombine them to obtain Finally, is sent to the generative transformation network to obtain a lossy reconstructed face image

[0085] Based on the same inventive concept, the embodiment of the present application also provides a system for implementing the above-mentioned image compression and reconstruction method using a grouped token mixer. The solution provided by this system to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the image compression and reconstruction system using a grouped token mixer provided below can refer to the limitations on the image compression and reconstruction method using a grouped token mixer in the above text, and will not be elaborated here.

[0086] In an exemplary embodiment, an image compression and reconstruction system using a grouped token mixer is provided, including:

[0087] A mapping unit for mapping the original image using an analysis transformation network to obtain hidden layer space features.

[0088] A quantization unit for quantizing the hidden layer space features to obtain quantization features.

[0089] An entropy model construction unit for constructing an entropy model based on a grouped token mixer; the entropy model includes a grouping module, a mapping network, multiple grouped token mixers, and a parameter network, and the grouped token mixer includes a cross-group token mixer and an intra-group token mixer; the cross-group token mixer is used to extract inter-group context information, and the intra-group token mixer is used to extract intra-group information.

[0090] An entropy coding and entropy decoding unit for performing entropy coding and entropy decoding on the quantization features based on the entropy model to obtain reconstructed features.

[0091] An image reconstruction unit for reconstructing an image based on the reconstructed features using a generative network.

[0092] In an exemplary embodiment, a computer device is provided, which includes a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented. The computer device can be a server or a terminal, and its internal structure diagram can be as shown in Figure 7 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data to be processed. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements an image compression and reconstruction method using a packet token mixer.

[0093] Those skilled in the art can understand that Figure 7 the structure shown in is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements. In an exemplary embodiment, a computer device is provided, which includes a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.

[0094] In an exemplary embodiment, a computer-readable storage medium is provided, which stores a computer program, and when the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0095] In an exemplary embodiment, a computer program product is provided, which includes a computer program, and when the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0096] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of the relevant data need to comply with the relevant regulations.

[0097] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tapes, floppy disks, flash memories, optical memories, high-density embedded non-volatile memories, resistive random access memories (ReRAM), magnetoresistive random access memories (MRAM), ferroelectric random access memories (FRAM), phase change memories (PCM), graphene memories, etc. Volatile memories can include random access memory (RAM) or external cache memories, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0098] The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logics, data processing logics based on quantum computing, etc., without limitation.

[0099] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the various technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.

[0100] In this article, specific examples are used to elaborate on the principles and implementation manners of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. At the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manners and application scopes. To sum up, the content of this specification should not be construed as a limitation to this application.

Claims

1. An image compression and reconstruction method using a grouped token mixer, characterized in that, Including: Using an analysis transformation network to map the original image to obtain hidden layer space features; Quantizing the hidden layer space features to obtain quantized features; Constructing an entropy model based on a grouped token mixer; the entropy model includes a grouping module, a mapping network, multiple grouped token mixers, and a parameter network, and the grouped token mixer includes a cross-group token mixer and an intra-group token mixer; The cross-group token mixer is used to extract inter-group context information, and the intra-group token mixer is used to extract intra-group information; Performing entropy encoding and entropy decoding on the quantized features based on the entropy model to obtain reconstructed features; Based on the reconstructed features, using a generation network to reconstruct the image.

2. The image compression and reconstruction method using a grouped token mixer according to claim 1, wherein Performing entropy encoding and entropy decoding on the quantized features based on the entropy model to obtain reconstructed features, specifically including: Dividing the quantized features into multiple groups of hidden layer variables through the grouping module; Mapping multiple groups of the hidden layer variables to the same shape through the mapping network; Performing feature extraction on multiple groups of the hidden layer variables through multiple grouped token mixers; Predicting probability distribution parameters of the extracted features through a parameter network; Encoding multiple groups of the hidden layer variables based on the predicted probability distribution parameters to obtain a binary code stream; Decoding and reorganizing the binary code stream to obtain reconstructed features.

3. The image compression and reconstruction method using a grouped token mixer according to claim 2, wherein The grouping module divides the quantized features along the channel and spatial dimensions respectively; when dividing in the spatial dimension, the quantized features are divided into multiple groups of hidden layer variables in a checkerboard pattern.

4. The image compression and reconstruction method using a grouped token mixer according to claim 1, characterized in that The cross-group token mixer includes a two-dimensional relative position embedding generator, a masked self-attention module, a feed-forward network, and two layers of normalization layers.

5. The image compression and reconstruction method using a grouped token mixer according to claim 1, wherein, The intra-group token mixer includes a position embedding generator, a self-attention module, a feed-forward network, and two layers of normalization layers.

6. The method for image compression and reconstruction using a grouped token mixer according to claim 2, wherein The entropy model decomposes the probability distribution parameters of the hidden layer variables in an inter-group autoregressive manner.

7. An image compression and reconstruction system using a grouped token mixer, characterized in that, Including: A mapping unit for using an analysis transformation network to map the original image to obtain hidden layer space features; A quantization unit for quantizing the hidden layer space features to obtain quantized features; An entropy model construction unit for constructing an entropy model based on a grouped token mixer; the entropy model includes a grouping module, a mapping network, multiple grouped token mixers, and a parameter network, and the grouped token mixer includes a cross-group token mixer and an intra-group token mixer; The cross-group token mixer is used to extract inter-group context information, and the intra-group token mixer is used to extract intra-group information; An entropy encoding and entropy decoding unit for performing entropy encoding and entropy decoding on the quantized features based on the entropy model to obtain reconstructed features; An image reconstruction unit for reconstructing the image based on the reconstructed features using a generation network.

8. A computer device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the image compression and reconstruction method using a grouped token mixer according to any one of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the image compression and reconstruction method using a grouped token mixer according to any one of claims 1-6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the image compression and reconstruction method using a packet token mixer according to any one of claims 1-6.