Frequency decomposition attention joint optimization image compression method, device and equipment

By employing a frequency decomposition attention joint optimization method, the problem of inaccurate global frequency component optimization in image compressed sensing is solved, achieving high-fidelity reconstruction of high-frequency details and low-frequency textures, and improving the visual effect of image compression.

CN121120835APending Publication Date: 2025-12-12Chinese People's Liberation Army Cyberspace Force Information Engineering University
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511313645.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-15
Publication Date
2025-12-12

AI Technical Summary

Technical Problem

In existing image compressed sensing technology, multi-scale implicit modeling is difficult to accurately optimize global frequency components, resulting in poor high-frequency detail reconstruction quality. The frequency division framework has problems with frequency component enhancement and image fusion.

Method used

A frequency decomposition attention joint optimization method is adopted. Initial reconstructed image feature information is generated through image sampling matrix. Multi-scale optimization network is used to independently optimize full-frequency, low-frequency and high-frequency feature information. The fusion is enhanced by frequency decomposition interactive attention module. A three-optimization network framework of full-frequency, low-frequency and high-frequency is designed, including principal component enhanced gradient descent module and U-shaped proximal mapping module.

Benefits of technology

It achieves explicit attention enhancement and good frequency division fusion of low-frequency and high-frequency components of the image, recovers high-quality image detail texture, and improves the visual effect of image compression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121120835A_ABST
    Figure CN121120835A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an image compression method, device and equipment based on frequency decomposition attention joint optimization. A specific embodiment of the method comprises the steps of performing convolution processing on an original image through an image sampling matrix; generating an initial reconstruction image and initial reconstruction image feature information according to the image sampling information and the image sampling matrix; based on the initial reconstructed image feature information, executing the following feature reconstruction steps: performing frequency domain feature extraction processing on the initial reconstructed image feature information; performing feature optimization processing on the image full-frequency feature information, the image high-frequency feature information and the image low-frequency feature information through a multi-scale optimization network; performing joint enhancement optimization on the optimized full-frequency feature information, the optimized high-frequency feature information and the optimized low-frequency feature information; and performing feature aggregation processing on the reconstructed image feature information to generate a reconstructed image. According to the embodiment, detail textures of the image can be recovered, so that the image compression visual effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present disclosure relate to the field of computer technology, in particular to an image compression method, device and equipment based on frequency decomposition attention joint optimization. BACKGROUND

[0002] Image compressive sensing (ICS) technology uses undersampled measurements to recover high-fidelity images. Because of its short measurement time, low data transmission and storage costs, and low sensor requirements, it is widely used in single-pixel imaging, THz imaging, underwater imaging, medical imaging, video compressive sensing, and remote sensing. At present, when performing image compressive sensing, the commonly used method is to realize implicit modeling of different scale components through multi-scale or attention mechanism network structure design, so as to realize the reconstruction of complex image texture. Or use a frequency division image reconstruction framework, which aims to decompose different frequency components of the image, so as to reconstruct the image specifically, and finally improve the image reconstruction fidelity.

[0003] However, the inventors have found that when using the above method for image compressive sensing, the following technical problems often exist: The implicit modeling of multi-scale is difficult to accurately and specifically optimize the reconstruction of global frequency components, thereby potentially limiting the reconstruction of high-quality high-frequency details. In addition, the frequency division framework still has the problem of how to enhance different frequency components, and the problem of image fusion inadaptability caused by frequency decomposition without complete full-frequency image guidance.

[0004] The above information disclosed in this Background section is only for the purpose of enhancing the understanding of the background of the present inventive concepts, and therefore, it can contain information that does not form the prior art that is already known to those skilled in the art. SUMMARY

[0005] The summary of the present disclosure is intended to introduce the concepts in a simplified form, which will be described in detail in the specific embodiments section. The summary of the present disclosure is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.

[0006] Some embodiments of the present disclosure propose an image compression method, device and equipment based on frequency decomposition attention joint optimization to solve the technical problems mentioned in the background section.

[0007] In a first aspect, some embodiments of this disclosure provide an image compression method using frequency decomposition and attention joint optimization. The method includes: convolving an original image using an image sampling matrix to generate image sampling information; generating an initial reconstructed image and initial reconstructed image feature information based on the image sampling information and the image sampling matrix; and performing the following feature reconstruction steps based on the initial reconstructed image feature information: performing frequency domain feature extraction processing on the initial reconstructed image feature information to generate full-frequency image feature information, high-frequency image feature information, and low-frequency image feature information; performing feature optimization processing on the full-frequency image feature information, high-frequency image feature information, and low-frequency image feature information using a multi-scale optimization network to obtain optimized full-frequency image feature information, optimized high-frequency image feature information, and optimized low-frequency image feature information; performing joint enhancement optimization on the optimized full-frequency image feature information, optimized high-frequency image feature information, and optimized low-frequency image feature information to generate reconstructed image feature information; and performing feature aggregation processing on the reconstructed image feature information to generate a reconstructed image.

[0008] Secondly, some embodiments of this disclosure provide an image compression apparatus with frequency decomposition attention joint optimization. The apparatus includes: a convolution unit configured to perform convolution processing on an original image using an image sampling matrix to generate image sampling information; a generation unit configured to generate an initial reconstructed image and initial reconstructed image feature information based on the image sampling information and the image sampling matrix; an execution unit configured to perform the following feature reconstruction steps based on the initial reconstructed image feature information: performing frequency domain feature extraction processing on the initial reconstructed image feature information to generate full-frequency image feature information, high-frequency image feature information, and low-frequency image feature information; performing feature optimization processing on the full-frequency image feature information, high-frequency image feature information, and low-frequency image feature information through a multi-scale optimization network to obtain optimized full-frequency image feature information, optimized high-frequency image feature information, and optimized low-frequency image feature information; performing joint enhancement optimization on the optimized full-frequency image feature information, optimized high-frequency image feature information, and optimized low-frequency image feature information to generate reconstructed image feature information; and a feature aggregation unit configured to perform feature aggregation processing on the reconstructed image feature information to generate a reconstructed image.

[0009] Thirdly, some embodiments of this disclosure provide an electronic device, including: one or more processors; and a storage device having one or more programs stored thereon, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation of the first aspect above.

[0010] Fourthly, some embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the method described in any of the implementations of the first aspect above.

[0011] The above embodiments of this disclosure have the following beneficial effects: the image compression method using frequency decomposition and attention joint optimization of some embodiments of this disclosure can recover image detail texture, thereby improving the visual effect of image compression. Specifically, the reason for the poor visual effect of related image compression is that multi-scale implicit modeling makes it difficult to accurately and specifically optimize and reconstruct global frequency components, thus potentially limiting high-quality high-frequency detail reconstruction. In addition, the frequency division framework still faces problems such as how to enhance different frequency components and the incompatibility of image fusion caused by frequency decomposition without a complete full-frequency image. Based on this, the image compression method using frequency decomposition and attention joint optimization of some embodiments of this disclosure can achieve explicit attention enhancement of low-frequency and high-frequency components of the image and good frequency division fusion. Specifically, a three-optimization network framework of full-frequency, low-frequency, and high-frequency (FLH) is established. The optimization of each component independently adopts an optimized unfolded multi-scale network (OM-Net), including a principal component enhanced gradient descent module (PCAGDM) and a U-shaped proximal mapping module (UPMM). PCAGDM supplements the optimization of the minimum-dimensional principal component complementary features while optimizing the principal component image to efficiently optimize FLH features. UPMM can perform multi-scale proximal mapping of FLH features, providing a large receptive field. A Frequency Decomposition Interactive Attention Module (FIAM) was designed to enhance the fusion of FLH features, particularly enhancing the high- and low-frequency components related to the full-frequency features, to reduce the unwanted influence of frequency decomposition. Ultimately, high-fidelity reconstruction of high-frequency details and low-frequency textures is achieved, reconstructing richer texture details than other methods… Attached Figure Description

[0012] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.

[0013] Figure 1 This is a schematic diagram comparing the frequency decomposition attention joint optimization image compression method of some embodiments of this disclosure with existing image compression sensing frameworks; Figure 2 This is a flowchart illustrating the framework of the frequency decomposition attention joint optimization image compression method according to this disclosure; Figure 3 This is a schematic diagram of the structure of the frequency decomposition interactive attention module in the image compression method of frequency decomposition attention joint optimization according to the present disclosure; Figure 4This is a visual comparison of the reconstruction results of the frequency decomposition attention joint optimization image compression method disclosed in this paper, using the Barbara image in the Set11 dataset at a sampling rate of 0.10 and the imag092 image in the Urban100 dataset at a sampling rate of 0.25. Figure 5 This is a visual comparison of the reconstruction results of the frequency decomposition and attention joint optimization image compression method disclosed in this paper, using the Butterfly image in the Set5 dataset at a sampling rate of 0.01 and the image at index 8 in the McM18 dataset at a sampling rate of 0.04. Figure 6 This is a flowchart of some embodiments of the frequency decomposition attention joint optimization image compression method of the present disclosure; Figure 7 This is a schematic diagram of the structure of some embodiments of the frequency decomposition attention joint optimization image compression apparatus of the present disclosure; Figure 8 This is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. Detailed Implementation

[0014] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0015] It should also be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.

[0016] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0017] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0018] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0019] This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.

[0020] Figure 1 This is a schematic diagram comparing the frequency decomposition attention joint optimization image compression method of some embodiments of this disclosure with existing image compression sensing frameworks.

[0021] For a long time, image compressed sensing network methods have focused on reconstructing high-fidelity image textures from a limited number of measurements. The main challenge of image compressed sensing lies in recovering complex image textures, i.e., high-frequency components, under conditions of significant undersampling. Most previous methods have utilized multi-scale or attention-based network architectures to implicitly model components at different scales, achieving good reconstruction of complex image textures. However, this multi-scale implicit modeling makes it difficult to accurately and specifically optimize the reconstruction of global frequency components, thus potentially limiting the reconstruction of high-quality high-frequency details, such as… Figure 1 As shown in (a). Furthermore, some ICS methods and other low-level vision tasks employ a frequency-division image reconstruction framework, aiming to decompose different frequency components of an image to specifically reconstruct the image, ultimately improving image reconstruction fidelity, and achieving great success. However, this frequency-division framework still faces long-standing problems such as how to enhance different frequency components and the inadequacy of image fusion caused by frequency decomposition without a complete full-frequency image as guidance, such as... Figure 1 As shown in (b).

[0022] Figure 1 A comparative diagram of the frequency decomposition-attention joint optimization image compression method and previous image compressed sensing frameworks. (a) The general form of a traditional image compressed sensing (ICS) network; yellow arrows indicate the loss of high-frequency details and blurred boundaries in low-frequency regions. (b) An ICS network using frequency decomposition; light blue arrows indicate the maladaptive issues in fused image reconstruction caused by frequency decomposition optimization. (c) The proposed frequency decomposition-attention joint optimization network (FAJO-Net). Reconstruction results show that this network can simultaneously achieve high-fidelity restoration of high-frequency and low-frequency textures and eliminate the maladaptive issues caused by frequency decomposition. All networks in the diagram use the same optimized unfolded multi-scale network structure (OM-Net).

[0023] Continue to refer to Figure 6 The flowchart 600 illustrates some embodiments of the frequency decomposition attention joint optimization image compression method according to the present disclosure. This frequency decomposition attention joint optimization image compression method includes the following steps: Step 601: Convolve the original image using the image sampling matrix to generate image sampling information.

[0024] In some embodiments, the implementer (e.g., a computing device) of the frequency decomposition attention joint optimization image compression method can utilize a learnable sampling matrix to non-overlappingly convolve the original image in blocks to generate image sampling information. In practice, the implementer can generate image sampling information using the following expression: .

[0025] in, It can be a sampling matrix. It can be the original image. It can be pattern sampling information.

[0026] Step 602: Generate the initial reconstructed image and initial reconstructed image feature information based on the image sampling information and the image sampling matrix.

[0027] In some embodiments, the execution entity can generate an initial reconstructed image and initial reconstructed image feature information based on the image sampling information and the image sampling matrix. The initial reconstructed image is typically of poor quality due to block artifacts and reconstruction noise, so deep reconstruction is usually required to optimize image quality. Deep networks typically process images in a characteristic manner, so convolution is used to expand the initial reconstructed image dimensionality into C-channel initial reconstructed image feature information for deep reconstruction.

[0028] In some optional implementations of certain embodiments, the execution entity may generate an initial reconstructed image and initial reconstructed image feature information based on the image sampling information and the image sampling matrix through the following steps: The first step is to generate an initial reconstructed image based on the image sampling information described above. In practice, the execution entity can generate the initial reconstructed image using the following expression: .

[0029] The second step involves expanding the dimension of the initial reconstructed image to generate its feature information. In practice, the executing entity can use the following expression to expand the dimension of the initial reconstructed image, thereby generating its feature information: .

[0030] in, It can be the initial reconstructed image. It can be the transpose of the sampling matrix. It can be the feature information of the initial reconstructed image. This can refer to the convolution operation from the image domain to the feature domain, specifically, it can be an increase in dimensionality to the number of channels C through 3×3 convolution.

[0031] Step 603: Based on the initial reconstructed image feature information, perform the following feature reconstruction steps: Step 6031: Perform frequency domain feature extraction processing on the initial reconstructed image feature information to generate full-frequency image feature information, high-frequency image feature information, and low-frequency image feature information.

[0032] In some embodiments, the aforementioned execution entity can perform frequency domain feature extraction processing on the initial reconstructed image feature information to generate full-frequency image feature information, high-frequency image feature information, and low-frequency image feature information. The frequency domain features are then processed by a high-pass filter and a low-frequency filter to generate high-frequency and low-frequency image feature information, respectively. In execution, the high-pass and low-pass filters are constructed by multiplying the frequency domain features by high-pass and low-pass masks of the same size as the features, as shown below. Figure 2 As shown in (b) and (c).

[0033] It should be noted that, Figure 2 The flowchart of the frequency decomposition attention joint optimization image compression method shown includes the following parts: (a) the overall structure of FAJO-Net; (b) the details of the K-th stage depth reconstruction; (c) the low-pass filter (LPF); (d) the high-pass filter (HPF); (e) the residual block (RB); (f) the principal component enhanced gradient descent module (PCAGDM); and (g) the U-shaped proximal mapping module (UPMM).

[0034] In some optional implementations of certain embodiments, the aforementioned execution entity may perform frequency domain feature extraction processing on the initial reconstructed image feature information through the following steps to generate full-frequency feature information, high-frequency feature information, and low-frequency feature information of the image: The first step is to perform a two-dimensional transformation on the initial reconstructed image feature information to generate full-frequency feature information. In practice, the aforementioned execution entity can perform a two-dimensional Fourier transform on the initial reconstructed image feature information to generate full-frequency feature information.

[0035] The second step involves performing high-frequency filtering on the full-frequency feature information of the image to generate high-frequency feature information. In practice, the aforementioned execution entity can generate high-frequency feature information using the following expression: .

[0036] The third step involves performing low-frequency filtering on the full-frequency feature information of the image to generate low-frequency feature information. In practice, the aforementioned execution entity can generate low-frequency feature information using the following expression: .

[0037] in, and These represent the high-frequency and low-frequency features of the image in the k-th stage (i.e., the k-th execution of the above feature reconstruction steps), respectively. Indicates the first The initial reconstructed image feature information (i.e., reconstructed image feature information) generated in the stage (i.e., the (k-1)th execution of the above feature reconstruction steps). and These represent the two-dimensional Fourier transform and the inverse two-dimensional Fourier transform, respectively. and These can represent low-pass masks (i.e., low-frequency filters) and high-pass masks (i.e., high-pass filters), respectively.

[0038] Step 6032: Through a multi-scale optimization network, feature optimization processing is performed on the full-frequency feature information, high-frequency feature information, and low-frequency feature information of the image to obtain optimized full-frequency feature information, optimized high-frequency feature information, and optimized low-frequency feature information.

[0039] In some embodiments, the aforementioned execution entity can perform feature optimization processing on the full-frequency feature information, high-frequency feature information, and low-frequency feature information of the image through a multi-scale optimization network to obtain optimized full-frequency feature information, optimized high-frequency feature information, and optimized low-frequency feature information. The joint optimization of the full-frequency feature information, high-frequency feature information, and low-frequency feature information first involves independent optimization by three optimized multi-scale networks (OM-Net). The optimization-inspired multi-scale network includes a Principal Component Supplemented Gradient Descent Module (PCAGDM) and a U-shaped multi-scale proximal mapping module (UPMM), which respectively expand the gradient descent and proximal mapping operations in the proximal gradient descent algorithm.

[0040] In practice, the aforementioned executing entity can generate optimized full-frequency feature information, optimized high-frequency feature information, and optimized low-frequency feature information using the following expressions: .

[0041] in, and These represent the Principal Component Supplemented Gradient Descent Module (PCAGDM) and the U-Shaped Multi-Scale Proximal Mapping Module (UPMM), respectively. The PCAGDM, building upon the physical input of gradient calculations in the principal component domain of the image, adds channel-wise gradient calculations for supplementary dimensions. The supplementary dimensions are compressed to a minimum to achieve high-quality image reconstruction with minimal gradient computation. Principal component images are obtained through convolutional feature aggregation, while supplementary principal component features are obtained through convolutional compression and residual block transformation. High-frequency feature information after stage optimization Optimized low-frequency feature information and optimized full-frequency feature information Generated according to the same structure, such as Figure 1 As shown. The optimized full-frequency feature information... Taking generation as an example, the network structure of the Principal Component Supplemented Gradient Descent Module (PCAGDM) is as follows: Figure 2 As shown in (d), it can be executed using the following expression: ; ; .

[0042] in, and They are respectively The gradient of the physical injection of principal component images and supplementary features in the stage. , , and These represent convolution operations from the feature domain to the image domain, from the image domain to the feature domain, from the feature domain to the complementary domain, and from the complementary domain to the feature domain, respectively. Indicates the first The features optimized by the gradient descent module supplemented by the principal component stage. The proximal mapping operation of the U-shaped multi-scale proximal mapping module is implemented through a U-shaped multi-scale network, such as... Figure 2 The structure shown in (e) is executed, and the expression can be represented as: .

[0043] Step 6033: Perform joint enhancement optimization on the optimized full-frequency feature information, optimized high-frequency feature information, and optimized low-frequency feature information to generate reconstructed image feature information.

[0044] In some embodiments, the aforementioned execution entity may perform joint enhancement optimization on the optimized full-frequency feature information, optimized high-frequency feature information, and optimized low-frequency feature information to generate reconstructed image feature information.

[0045] In practice, optimized high-frequency feature information Optimized low-frequency feature information and optimized full-frequency feature information The joint enhancement optimization of the frequency division interactive attention module yields the reconstructed image feature information for the Kth stage, such as... Figure 2 As shown, the frequency-division interactive attention module is a dual-cross attention structure. It performs cross-attention with the optimized high-frequency and low-frequency feature information, respectively, and the optimized full-frequency feature information to enhance the low-frequency and high-frequency components related to the full-frequency features and reduce the impact of artifacts added due to image decomposition. In the dual-cross attention, firstly, high-frequency queries... and low-frequency queries The features optimized from the full frequency are processed by layer normalization LN and then dimensionality-upgraded 1×1 convolution. and depthwise convolution After embedding, the channel is split in half, and finally reshaped into the following shape: Generate, then key Sum ,key Sum The same operation is used to generate features from both the low-frequency optimized features and the high-frequency optimized features, such as... Figure 3 As shown.

[0046] Low-frequency attention maps and high-frequency attention maps are generated by high-frequency queries. and low-frequency queries Respectively with low-frequency and high-frequency keys and The multiplication is followed by a softmax operation to generate the low-frequency and high-frequency attention maps, which are then multiplied by the low-frequency and high-frequency values ​​respectively, transforming their shapes as follows: Finally, channel interaction is performed through 1×1 convolution to obtain the optimized low-frequency feature information, optimized high-frequency feature information, and attention terms of the optimized full-frequency feature information. and .

[0047] Finally, the optimized low-frequency feature information and the optimized high-frequency feature information, along with their attention terms, are added and fused into a high-low frequency attention fusion feature. This feature is then concatenated with the optimized full-frequency feature information and aggregated through convolution to obtain the reconstructed image feature information for the Kth stage. (i.e., the initial reconstructed image feature information in the feature reconstruction step).

[0048] Specifically, the aforementioned execution entity can generate reconstructed image feature information using the following expression: ; ; ; ; ; .

[0049] in, and These represent the attention terms for optimized low-frequency feature information, optimized high-frequency feature information, and optimized full-frequency feature information, respectively. and These represent the feature shape transformation (reshape) operation and the feature separation (slipt) operation, respectively. This represents a convolution operation that aggregates features from channel 2C to C. This indicates a splicing operation. This can be the feature information of the reconstructed image in the Kth stage.

[0050] Optionally, the above feature reconstruction step may further include the following steps: in response to determining that the number of times the above feature reconstruction step has been executed is less than or equal to a preset number of iterations, the reconstructed image feature information is used as the initial reconstructed image feature information, and the above feature reconstruction step is executed again. The specific setting of the preset number of iterations is not specifically limited here.

[0051] Step 604: Perform feature aggregation processing on the reconstructed image feature information to generate a reconstructed image.

[0052] In some embodiments, the execution entity described above can perform feature aggregation processing on the reconstructed image feature information to generate a reconstructed image. In practice, the execution entity can perform feature aggregation processing on the reconstructed image feature information using the following expression to generate a reconstructed image: .

[0053] in, This indicates the reconstructed image. This is represented as a convolution operation from features to an image.

[0054] Following the methodological steps, the frequency decomposition attention joint optimization image compression method of this disclosure was experimentally evaluated. Experimental data and model parameters were set as follows: image block size was set to 32; the number of iteration stages was 8; and the default number of channels was 32. The training data consisted of 40,000 unlabeled COCO2017 images, with each training image size being 128×128 and a batch size of 16. The model was trained using the Adam optimizer, with initial learning rates set to 0.0002, 0.0002, 0.0002, 0.0002, and 0.0001 at sampling rates of 0.01, 0.04, 0.10, 0.25, and 0.50, respectively. The learning rate was halved when the average PSNR on the Set11 and Urban100 datasets no longer improved after five consecutive training epochs. The model was evaluated using the widely used Set5, Set11, McM18, and Urban100 datasets. All evaluations were performed on the Y channel of the YCbCr space of the images. All evaluation experiments were conducted in a PyTorch-1.11.0 environment on an NVIDIA Quadro RTX 6000 GPU.

[0055] The image details reconstructed by the proposed method were visualized and compared with those of various other methods at sampling rates of 0.10 and 0.25. Figure 4 As shown. Figure 4This demonstrates that the proposed method achieves better reconstruction of high-frequency image details and textures that are easily lost in compressed sampling, and can recover some details and textures that other methods cannot recover, thus greatly improving the visual effect of image reconstruction. Furthermore, we visualized and compared the reconstructed image details of the proposed method with other methods at two low sampling rates of 0.01 and 0.04, such as... Figure 5 As shown. Figure 5 This demonstrates that the proposed method achieves sharper contour boundaries while preserving low-frequency image contours even at extremely low sampling rates, thanks to the excellent reconstruction of high-frequency components in the image. In summary, the proposed method maintains a consistent advantage over competing methods in terms of visual performance, regardless of whether the sampling rate is low or high.

[0056] It should be noted that, Figure 4 and Figure 5 The best and second-best PSNR / SSIM results are highlighted in red and blue, respectively.

[0057] Further reference Figure 7 As an implementation of the methods shown in the above figures, this disclosure provides some embodiments of a frequency decomposition attention joint optimization image compression apparatus, which are similar to... Figure 6 Corresponding to the method embodiments shown, this frequency decomposition attention joint optimization image compression device can be specifically applied to various electronic devices.

[0058] like Figure 7 As shown, the frequency decomposition attention joint optimization image compression apparatus 700 in some embodiments includes: a convolution unit 701, a generation unit 702, an execution unit 703, and a feature aggregation unit 704. The convolution unit 701 is configured to perform convolution processing on the original image using an image sampling matrix to generate image sampling information; the generation unit 702 is configured to generate an initial reconstructed image and initial reconstructed image feature information based on the image sampling information and the image sampling matrix; the execution unit 703 is configured to perform the following feature reconstruction steps based on the initial reconstructed image feature information: perform frequency domain feature extraction processing on the initial reconstructed image feature information to generate full-frequency image feature information, high-frequency image feature information, and low-frequency image feature information; perform feature optimization processing on the full-frequency image feature information, high-frequency image feature information, and low-frequency image feature information through a multi-scale optimization network to obtain optimized full-frequency image feature information, optimized high-frequency image feature information, and optimized low-frequency image feature information; perform joint enhancement optimization on the optimized full-frequency image feature information, optimized high-frequency image feature information, and optimized low-frequency image feature information to generate reconstructed image feature information; and the feature aggregation unit 704 is configured to perform feature aggregation processing on the reconstructed image feature information to generate a reconstructed image.

[0059] It is understandable that the units described in the frequency decomposition attention joint optimization image compression apparatus 700 and the reference Figure 6 The steps in the described method correspond accordingly. Therefore, the operations, features, and beneficial effects described above for the method also apply to the frequency decomposition attention joint optimization image compression apparatus 700 and the units contained therein, and will not be repeated here.

[0060] The following is for reference. Figure 8 It shows a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. Figure 8 The electronic device shown is merely an example and should not be construed as limiting the functionality or scope of the embodiments of this disclosure. Figure 8 As shown, the computer device includes a processor, memory, and a network interface connected via a system bus. The memory may include a non-volatile storage medium and internal memory. The non-volatile storage medium may store an operating system and a computer program. The computer program includes program instructions that, when executed, cause the processor to perform any of the methods described above. The processor provides computational and control capabilities to support the operation of the entire computer device. The internal memory provides an environment for the execution of the computer program in the non-volatile storage medium; when executed by the processor, the computer program causes the processor to perform any of the methods described above. The network interface is used for network communication, such as sending assigned tasks. Those skilled in the art will understand that... Figure 8 The structure shown is merely a block diagram of a portion of the structure related to the present disclosure and does not constitute a limitation on the computer device to which the present disclosure is applied. A specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0061] It should be understood that the processor can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among these, a general-purpose processor can be a microprocessor or any conventional processor.

[0062] In one embodiment, the processor is configured to run a computer program stored in a memory to perform the following steps: convolving the original image using an image sampling matrix to generate image sampling information; generating an initial reconstructed image and initial reconstructed image feature information based on the image sampling information and the image sampling matrix; and performing the following feature reconstruction steps based on the initial reconstructed image feature information: performing frequency domain feature extraction processing on the initial reconstructed image feature information to generate full-frequency image feature information, high-frequency image feature information, and low-frequency image feature information; performing feature optimization processing on the full-frequency image feature information, high-frequency image feature information, and low-frequency image feature information using a multi-scale optimization network to obtain optimized full-frequency image feature information, optimized high-frequency image feature information, and optimized low-frequency image feature information; performing joint enhancement optimization on the optimized full-frequency image feature information, optimized high-frequency image feature information, and optimized low-frequency image feature information to generate reconstructed image feature information; and performing feature aggregation processing on the reconstructed image feature information to generate a reconstructed image.

[0063] This disclosure also provides a computer-readable storage medium storing a computer program, the computer program including program instructions, and the method implemented when the program instructions are executed can be referred to the various embodiments of the methods described above.

[0064] The aforementioned computer-readable storage medium may be an internal storage unit of the computer device described in the foregoing embodiments, such as a hard disk or memory of the computer device. Alternatively, the aforementioned computer-readable storage medium may be an external storage device of the computer device, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card equipped on the computer device.

[0065] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.

[0066] The above description is merely a selection of preferred embodiments of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in the embodiments of this disclosure.

Claims

1. A frequency decomposition attention joint optimization image compression method, characterized in that, include: The original image is convolved using an image sampling matrix to generate image sampling information; Based on the image sampling information and the image sampling matrix, an initial reconstructed image and initial reconstructed image feature information are generated; Based on the initial reconstructed image feature information, the following feature reconstruction steps are performed: Frequency domain feature extraction is performed on the initial reconstructed image feature information to generate full-frequency feature information, high-frequency feature information, and low-frequency feature information of the image; By using a multi-scale optimization network, feature optimization processing is performed on the full-frequency feature information, high-frequency feature information, and low-frequency feature information of the image to obtain optimized full-frequency feature information, optimized high-frequency feature information, and optimized low-frequency feature information. The optimized full-frequency feature information, optimized high-frequency feature information, and optimized low-frequency feature information are jointly enhanced and optimized to generate reconstructed image feature information; Feature aggregation processing is performed on the feature information of the reconstructed image to generate the reconstructed image.

2. The method according to claim 1, characterized in that, The step of generating an initial reconstructed image and initial reconstructed image feature information based on the image sampling information and the image sampling matrix includes: An initial reconstructed image is generated based on the image sampling information; The initial reconstructed image is subjected to dimensional expansion processing to generate initial reconstructed image feature information.

3. The method according to claim 2, characterized in that, The step of performing frequency domain feature extraction on the initial reconstructed image feature information to generate full-frequency image feature information, high-frequency image feature information, and low-frequency image feature information includes: The initial reconstructed image feature information is subjected to two-dimensional transformation processing to generate full-frequency feature information of the image; High-frequency filtering is applied to the full-frequency feature information of the image to generate high-frequency feature information of the image; Low-frequency filtering is applied to the full-frequency feature information of the image to generate low-frequency feature information.

4. The method according to claim 3, characterized in that, The feature reconstruction step further includes: In response to determining that the number of times the feature reconstruction step is executed is less than or equal to the preset number of iterations, the reconstructed image feature information is used as the initial reconstructed image feature information, and the feature reconstruction step is executed again.

5. An image compression apparatus with frequency decomposition and attention joint optimization, characterized in that, include: The convolutional unit is configured to convolve the original image through the image sampling matrix to generate image sampling information; The generation unit is configured to generate an initial reconstructed image and initial reconstructed image feature information based on the image sampling information and the image sampling matrix; The execution unit is configured to perform the following feature reconstruction steps based on the initial reconstructed image feature information: perform frequency domain feature extraction processing on the initial reconstructed image feature information to generate full-frequency feature information, high-frequency feature information, and low-frequency feature information of the image; By using a multi-scale optimization network, feature optimization processing is performed on the full-frequency feature information, high-frequency feature information, and low-frequency feature information of the image to obtain optimized full-frequency feature information, optimized high-frequency feature information, and optimized low-frequency feature information. Joint enhancement optimization is then performed on the optimized full-frequency feature information, optimized high-frequency feature information, and optimized low-frequency feature information to generate reconstructed image feature information. The feature aggregation unit is configured to perform feature aggregation processing on the feature information of the reconstructed image to generate the reconstructed image.

6. An electronic device, characterized in that, include: One or more processors; A storage device on which one or more programs are stored; When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1 to 4.

7. A computer-readable medium, characterized in that, It stores a computer program thereon, wherein the computer program, when executed by a processor, implements the method as described in any one of claims 1 to 4.