A method and system for contrast enhancement of infrared images that preserves spatial uniformity

By constructing an infrared image enhancement model that maintains spatial consistency and utilizing the interaction of spatial and channel information, the problems of insufficient spatial consistency and detail representation in infrared image enhancement are solved, achieving high-quality enhancement effects for infrared images.

CN120355638BActive Publication Date: 2026-02-06ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510501398.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2026-02-06
Estimated Expiration
2045-04-21

AI Technical Summary

Technical Problem

Existing infrared image enhancement technologies are insufficient in terms of spatial consistency and detail representation, resulting in blurring and artifacts in the enhanced infrared images.

Method used

An infrared image enhancement model based on spatial consistency preservation is constructed. Through spatial information interaction and channel information interaction, a feature extraction subnet, a contrast enhancement subnet, and a detail enhancement subnet are used for training. The model is trained by combining a global enhancement loss function, a multi-scale structure loss function, and a histogram loss function. The enhancement curve is adjusted pixel by pixel to inject detail information.

Benefits of technology

It effectively improves the spatial consistency and detail quality of infrared image enhancement, eliminates pixel intensity inconsistencies, artifacts and detail blurring in the enhancement results, and improves visual effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355638B_ABST
    Figure CN120355638B_ABST
Patent Text Reader

Abstract

The application discloses an infrared image contrast enhancement method and system for maintaining spatial consistency, to solve the problem of low image contrast and fuzzy details caused by factors such as severe weather and environmental temperature interference during infrared image shooting. The method uses spatial information interaction and channel information interaction to construct an infrared image enhancement model based on spatial consistency maintenance, extracts spatial consistent expression on low-resolution features, adjusts and enhances the curve in the feature domain through pixel-by-pixel curve, and injects detailed information through local interaction on multi-scale expression. The method is trained using a global enhancement loss function, a multi-scale structure loss function and a histogram loss function. Compared with existing infrared enhancement methods, the method can effectively enhance low-contrast infrared images, effectively improve the spatial consistency and detail quality of infrared image enhancement, eliminate the problems of inconsistent pixel intensity in uniform areas of the enhanced result, artifacts and fuzzy details, and improve the visual effect of the enhanced result.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of image processing, and particularly relates to an infrared image contrast enhancement method and system keeping spatial consistency. BACKGROUND

[0002] In recent years, infrared image enhancement technology has important application value in the fields of security monitoring, industrial detection, etc. However, due to the fact that infrared images are often collected under extreme environmental conditions such as low light, fog, overexposure, etc., the imaging quality is low, the details are fuzzy, and the contrast is insufficient. For this reason, researchers have proposed many image enhancement techniques to improve image visibility. Traditional image enhancement methods mainly include histogram equalization and methods based on Retinex theory. The histogram equalization method improves the contrast by stretching the pixel value distribution, but it is easy to produce noise and artifacts, and the parameters need to be manually adjusted for different scenes, so the generalizability is poor. The Retinex method assumes that the image can be decomposed into illumination and reflection components, and the image quality is improved by enhancing the reflection component, but this assumption is often not true under complex lighting conditions, and it is also difficult to effectively introduce regularization constraints.

[0003] With the development of deep learning technology, more and more learning-based enhancement methods are proposed. This kind of method can be mainly divided into three categories: end-to-end network structure, network structure based on Retinex theory, and plug-and-play enhancement module. LLNet (LLNet: A deep autoencoder approach to natural low-light image enhancement) and LightenNet (Lightennet: A convolutional neural network for weakly illuminated image enhancement) represented by end-to-end network first use convolutional neural network to directly output enhanced images, but it is difficult to achieve fine detail enhancement due to the simple single-stage network structure. Subsequently, scholars introduced local kernel prediction module (LEDNet: Joint low-light enhancement and deblurring in the dark) or long-distance modeling module (Ultra-high definition low-light image enhancement: A benchmark and transformer based method) to further improve the image details and global information. But this kind of method usually has high computational complexity, and the model lacks interpretability. In addition, Retinex deep network represented by RetinexNet (Sparse gradient regularized deep Retinex network for robust low-light image enhancement) has good interpretability, but it still cannot effectively deal with complex infrared imaging conditions. At the same time, image enhancement methods based on the concept of plug-and-play, such as ZeroDCE (Zero-reference deep curve estimation for low-light image enhancement) and ChebyLighter (Chebylighter: Optimal curve estimation for low-light image enhancement), adjust the brightness of the image through parameterized curves. These methods are simple and effective, but mainly for visible light images, and lack of special design for the unique degradation of infrared images, so it is difficult to well handle the problems such as noise and spatial structure inconsistency commonly existing in infrared images.

[0004] For the particularity of infrared image, there are also special repair and enhancement methods. Some methods such as TherSuRNet (TherISuRNet-a computationally efficient thermal image super-resolution network), ChaSNet (Channel split convolutional neural network (ChaSNet) for thermal image super-resolution) and other single-stage networks try to balance the performance and computational complexity, while other works such as PSRGAN (Infrared image super-resolution via transfer learning and PSRGAN) process more challenging conditions through multi-stage networks. But the existing methods pay more attention to overall brightness enhancement, and the details are limited, and the spatial consistency is insufficient, resulting in blurred and artifact problems in the enhanced infrared image. But the method can effectively enhance the low-contrast infrared image, effectively improve the spatial consistency and detail quality of infrared image enhancement, eliminate the problem of inconsistent pixel intensity in uniform areas of the enhanced result, artifact, detail blur, and improve the visual effect of the enhanced result. SUMMARY

[0005] In view of the technical defects of the existing infrared image enhancement technology, such as insufficient spatial consistency and insufficient detail enhancement, the present application provides an infrared image contrast enhancement method for maintaining spatial consistency, which uses spatial information interaction and channel information interaction to construct an infrared image enhancement model based on spatial consistency maintenance, extracts spatially consistent expression on low-resolution features, and adjusts the enhancement curve in the feature domain through the pixel-by-pixel curve; injects detail information through local interaction on multi-scale expression; and uses global enhancement loss function, multi-scale structure loss function and histogram loss function for training.

[0006] In order to achieve the above purpose, the technical scheme adopted by the present application is as follows:

[0007] In the first aspect, the present application provides an infrared image contrast enhancement method for maintaining spatial consistency, which enhances the contrast of infrared image and improves the detail quality of image by using spatial information interaction and channel information interaction, including the following steps:

[0008] An infrared image enhancement model is trained, and the infrared image enhancement model includes a feature extraction subnetwork, a contrast enhancement subnetwork and a detail enhancement subnetwork.

[0009] The feature extraction subnetwork is used to encode a given low-contrast, detail-loss-degraded infrared image into a low-resolution feature representation. The contrast enhancement subnetwork is used to enhance the low-resolution feature representation in the feature domain to obtain an enhanced feature representation. The detail enhancement subnetwork is used to progressively enhance the detail structure of the enhanced feature representation and reconstruct the original resolution image.

[0010] The trained infrared image enhancement model is used to enhance the contrast of a given degraded infrared image and output the enhancement result.

[0011] Furthermore, the feature extraction subnet employs nonlinear, inactive network blocks.

[0012] Furthermore, the contrast enhancement subnet is composed of several cascaded global enhancement modules (GEMs). Each GEM maintains the generation of a spatially consistent descriptor through long-distance spatial information interaction and cross-channel information interaction, and performs global enhancement with consistency preservation on the feature domain.

[0013] Furthermore, the GEM calculation process is as follows:

[0014] Let F be the input of the nth GEM. n ∈R h×w×c h, w, and c are the length, width, and number of channels of the input image, respectively. A two-layer self-attention mechanism is adopted, namely the spatial interaction layer and the channel interaction layer to realize long-distance spatial information interaction and cross-channel information interaction, respectively. The attention map dimension of the spatial interaction layer is hw×hw, and the attention map dimension of the channel interaction layer is c×c. Then the outputs of the two layers are added together to obtain the global consistency descriptor P.

[0015] Using a curve generator to extract the input F from the GEM n The globally consistent descriptor P generates an enhanced curve tensor C∈R. h ×w×K K represents the number of curve adjustment iterations;

[0016] Using the enhancement curve tensor C on F n Perform iterative curve adjustments, and use the result of the final adjustment as the output.

[0017] Furthermore, during iterative curve adjustment, the k-th curve adjustment uses the k-th channel slice in the enhancement curve tensor C as the enhancement curve parameter, denoted as C. k ∈R h×w×1 Perform the following curve adjustments:

[0018]

[0019] in, This represents the result of the k-th feature field enhancement, and ⊙ represents element-wise multiplication.

[0020] Furthermore, the detail enhancement subnet is composed of several cascaded local enhancement modules (LEMs), each LEM consisting of a resolution reconstructor and a detail injector, progressively enhancing and reconstructing high-resolution details; the enhancement features of the last LEM are decoded into the enhanced result of the original resolution through a convolutional layer.

[0021] Furthermore, the LEM calculation process is as follows:

[0022] Enhanced features from low resolution using a resolution reconstructor Preliminary reconstruction of high-resolution features, denoted as F content ,Depend on F obtained from reconstruction content Dimensions and The dimensions are the same; the resolution reconstructor is an upsampled convolutional layer. These are the input and output of the m-th LEM, respectively;

[0023] The detail injector comprises a dual attention layer and a weighted fusion operation. First, in the dual attention layer, F... content and F detail The detail-aware enhancement features are calculated from the query features Q, which are mutually neighboring attention features, and the key features K and V. and contrast perception enhancement features Where F detail It is extracted from infrared degraded images by convolutional layers and is related to F. content First, detailed features of the same dimension are generated; second, detail-aware enhanced features are generated through a weight generator. and contrast perception enhancement features The fusion weights are then calculated, and the two are weighted and fused to output high-resolution enhanced features containing rich detail information.

[0024] Furthermore, the loss functions used during the training of the infrared image enhancement model include the global enhancement loss function, the local enhancement loss function, and the histogram loss function;

[0025] Global augmentation loss function L GE for:

[0026]

[0027] in, This indicates an enhancement result derived from low-resolution feature representation. Y represents a convolutional decoding layer. ↓M This means downsampling the ground truth image to half the size of the original image. M The size is doubled to match the size of the low-resolution enhancement result; ||.|1 is the L1 norm;

[0028] Set multi-scale features The size is 1 / 2 of the original picture M-m , local enhancement loss function L LE is:

[0029]

[0030] Wherein, M is the number of LEM, Indicates the enhanced result generated from the multi-scale features, Y ↓M Indicates that the true value image is down-sampled to 1 / 2 of the original size m Times to match the size of the multi-scale enhanced result, SSIM indicates the structural similarity;

[0031] Histogram loss function L hist is:

[0032]

[0033] Wherein, h indicates the histogram of the enhanced result, Indicates the histogram of the true value image.

[0034] Further, it also includes a reconstruction loss function:

[0035]

[0036] Wherein, Y is the true value image, Is the enhanced result, λ per Is the weight coefficient of the perceptual loss, Is the multi-layer feature extracted using the pre-trained VGG network, L rec Is the reconstruction loss.

[0037] The second aspect, the present application provides a kind of to keep spatial consistency infrared image contrast enhancement system, for realizing the spatial consistency infrared image contrast enhancement method described above.

[0038] The beneficial effects of the present application are:

[0039] The method uses spatial information interaction and channel information interaction to construct an infrared image enhancement model based on spatial consistency preservation, extracts spatially consistent expressions on low-resolution features, adjusts the enhancement curve in the feature domain through the pixel-by-pixel curve, and injects detailed information through local interaction on multi-scale expressions. The method is trained using a global enhancement loss function, a multi-scale structure loss function, and a histogram loss function. Compared with existing infrared enhancement methods, the method can effectively enhance low-contrast infrared images, effectively improve the spatial consistency and detail quality of infrared image enhancement, eliminate the problem of inconsistent pixel intensity in uniform areas of the enhanced result, artifacts, and blurred details, and improve the visual effect of the enhanced result. BRIEF DESCRIPTION OF DRAWINGS

[0040] Figure 1 is the overall framework diagram of the present application.

[0041] Figure 2 is a schematic diagram of a global enhancement module (GEM).

[0042] Figure 3 is a schematic diagram of a local enhancement module (LEM).

[0043] Figure 4 is a schematic diagram of a long-distance spatial information interaction layer and a cross-channel information interaction layer in the global enhancement module of the method of the present application.

[0044] Figure 5 is a schematic diagram of a detail injector in the local enhancement module of the method of the present application.

[0045] Figure 6 is a display diagram of infrared degraded images and enhanced results of an embodiment of the present application. (a) input image, i.e., infrared degraded image, (b) result of an existing comparative method SNRLLE (published in CVPR2022), (c) result of the method of the present application, and (d) true value image. DETAILED DESCRIPTION

[0046] The present application will be further described and explained with reference to the specific embodiments. The embodiments are only exemplary and do not circumscribe the scope of the present disclosure. The technical features of each embodiment of the present application can be combined accordingly without conflict.

[0047] The accompanying drawings are merely schematic illustrations of the present application and are not necessarily drawn to scale. Some of the block diagrams shown in the drawings are functional entities, which do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0048] The flowchart shown in the drawing is only an exemplary illustration, and is not necessarily required to include all steps. For example, some steps can be further decomposed, and some steps can be combined or partially combined, so the actual execution order can be changed according to actual conditions.

[0049] The present application provides an infrared image contrast enhancement method for maintaining spatial consistency, which enhances the contrast of infrared images by using spatial information interaction and channel information interaction, and improves the quality of image details, and the specific steps are as follows:

[0050] S1: Construct an infrared image enhancement model for maintaining spatial consistency, including a feature extraction subnetwork, a contrast enhancement subnetwork, and a detail enhancement subnetwork, and use a global enhancement loss function, a multi-scale structure loss function, and a histogram loss function to train the constructed model until convergence;

[0051] S2: For a given infrared image with low contrast and detail loss (hereinafter referred to as a degraded image), use the infrared image enhancement model described in S1, as shown in Figure 1 , the specific steps are as follows:

[0052] S21: Use the feature extraction subnetwork to encode the image into a low-resolution feature representation;

[0053] S22: Use the global contrast enhancement subnetwork to enhance the low-resolution feature representation in the feature domain;

[0054] S23: Use the multi-scale detail enhancement subnetwork to enhance the details step by step and reconstruct the original resolution image.

[0055] The infrared image contrast enhancement model for maintaining spatial consistency described in S1 uses a nonlinear activation-free network block (NAF BLOCK) as a basic unit to construct the feature extraction subnetwork described in S1 for extracting a low-resolution feature representation. The feature extraction subnetwork maps the degraded image X to a low-resolution feature representation F L , and inputs it into the contrast enhancement subnetwork described in S1.

[0056] The contrast enhancement subnetwork described in S1 is composed of a plurality of global enhancement modules (GEMs) connected in series, which generates an enhanced feature representation F E . Each GEM maintains a spatially consistent descriptor by long-distance spatial information interaction and cross-channel information interaction, and performs global enhancement for consistency maintenance in the feature domain, as shown in Figure 2 , the specific calculation process of GEM is as follows:

[0057] (1) Let the input of the nth GEM be F n ∈R h×w×c, two layers of self-attention mechanisms, i.e. a spatial interaction layer and a channel interaction layer, are adopted to realize long-distance spatial information interaction and cross-channel information interaction respectively, wherein the attention map dimension of the spatial interaction layer is hw x hw, and the attention map dimension of the channel interaction layer is c x c, then the outputs of the two layers are added to obtain the global consistency descriptor P.

[0058] As shown in (a) of FIG. 1, Figure 4 The spatial interaction layer mainly focuses on the relationship between different spatial positions in the input feature map. In one specific implementation of the present application, for the input feature map F n ∈R h×w×c , its dimension is converted into a form suitable for attention calculation, i.e. flattened from the feature map dimension h x w x c to hw x c, and then a fully connected layer f s is used for spatial interaction encoding, and the encoded feature map is mapped through linear transformation matrices W QS , W KS , and W VS to obtain a query matrix Q S , a key matrix K S , and a value matrix V S , with a dimension of hw x d s , wherein d s is the projection dimension; the query matrix Q S and the key matrix K S are used to calculate an attention map with a dimension of hw x hw, and the attention map is multiplied with the value matrix V S to output a spatial interaction expression, which is reshaped back to a dimension of h x w x c, and then residual concatenated with the input feature map, and then a convolution layer is used to map the dimension back to h x w x c.

[0059] As shown in (b) of FIG. 1, Figure 4 The channel interaction layer mainly focuses on the relationship between different channels in the input feature map. In one specific implementation of the present application, for the input feature map F n ∈R h×w×c , its dimension is converted into a form suitable for attention calculation, i.e. flattened from the feature map dimension h x w x c to hw x c, and then a fully connected layer f c is used for spatial interaction encoding, and the encoded feature map is mapped through linear transformation matrices W QC , W KC , and W VC to obtain a query matrix Q C , a key matrix K C , and a value matrix V C , with a dimension of d c x hw, wherein d c is the projection dimension; the query matrix Q C and the key matrix K CCompute attention map, dimension d c ×d c , use attention map and value matrix V C Multiply output channel interaction representation, reshape channel interaction representation back to dimension h×w×c, and perform residual concatenation with input feature map, and then use convolution layer to map dimension back to h×w×c.

[0060] (2) Use curve generator to generate enhanced curve tensor C∈R n from input F h and global consistency descriptor P of GEM. ×w×K , for iterative curve adjustment, K is the number of curve adjustment iterations.

[0061] In this embodiment, the curve generator uses a learnable neural network, which concatenates F n and P channel dimensions as the input of the curve generator, and outputs the enhanced curve tensor C.

[0062] (3) In the kth curve adjustment, use the kth channel slice in the enhanced curve tensor C as the enhanced curve parameter, denoted as C k ∈R h×w×1 , perform the following curve adjustment:

[0063]

[0064] where, is the kth feature domain enhancement result, and is the element-wise multiplication;

[0065] The detail enhancement subnetwork described in S1 is composed of a plurality of local enhancement modules (LEMs) cascaded. Each LEM is composed of a resolution reconstructor and a detail injector, which gradually enhances and reconstructs high-resolution details. The enhanced features of the last LEM are decoded into enhanced results of the original resolution through a convolution layer. As shown in Figure 3 , the specific calculation process of each LEM is as follows:

[0066] (1) Use the resolution reconstructor to preliminarily reconstruct high-resolution features from low-resolution enhanced features , denoted as F content . In this embodiment, F content reconstructed by has the same dimension as , and the resolution reconstructor is an up-sampling convolution layer.

[0067] (2) The detail injector includes a dual attention layer and a weighted fusion operation, and injects the detail structure expression F detail extracted by the detail extraction layer from the original image into F content , and the specific steps are as follows:

[0068] First, in the dual attention layer, F content and F detail Query features Q and key-value features K and V that serve as mutual neighborhood attention are used to compute detail-aware enhancement features. and contrast perception enhancement features Here, F detail It is extracted from infrared degraded images by convolutional layers and is related to F. content Detail features of the same dimension, multiple features used to extract F at different scales detail The convolutional layer is called the detail extraction layer;

[0069] like Figure 5 As shown, the neighborhood attention mechanism will select the element Q at a spatial location (x, y) of the query feature Q. L The key features K and V are centered at their corresponding spatial locations (x, y). Neighboring element K within the window L and V L Performing matrix multiplication yields a spatial dimension of... The neighborhood attention map is then multiplied by the value feature to obtain the neighborhood-aware feature at position (x, y). This operation is performed on all positions in the spatial dimension of Q, and the resulting neighborhood-aware feature is the attention output.

[0070] Secondly, a weight generator is used to generate fusion weights for detail-aware enhancement features and contrast-aware enhancement features. Then, the two are weighted and fused to output high-resolution enhancement features containing rich detail information.

[0071] In this embodiment, the weight generator uses a learnable single-layer neural network to... and The channel dimensions are concatenated and used as input to the weight generator, which outputs weight W.

[0072] The loss functions used in training infrared image enhancement models include global enhancement loss function, local enhancement loss function, and histogram loss function.

[0073] The reconstruction loss function used for training the model described in S1 is defined as follows: Where Y and λ represents the reference ground truth image and the enhancement result of this algorithm, respectively. per These represent the weighting coefficients of the perceptual loss, which is defined as follows: in This indicates that a pre-trained VGG network is used to extract multi-layer features.

[0074] The global augmentation loss function L used for training the model described in S1 GESupervise low-resolution enhancement features, set low-resolution features F E The size is 1 / 2 of the original image M The global enhancement loss function calculation method is:

[0075]

[0076] Among them, Indicates the enhancement result generated from the low-resolution feature, Indicates a convolutional decoding layer, Y ↓M Indicates that the true value image is down-sampled to 1 / 2 of the original image size M times to match the size of the low-resolution enhancement result.

[0077] The local enhancement loss function L used for the model training of S1 LE Multi-scale supervision is performed on the features of the pair, and multi-scale features The size is 1 / 2 of the original image M-m The local enhancement loss function calculation method is:

[0078]

[0079] Among them, Indicates the enhancement result generated from the multi-scale feature, Indicates a convolutional decoding layer, Y ↓M Indicates that the true value image is down-sampled to 1 / 2 of the original size m times to match the size of the multi-scale enhancement result.

[0080] The histogram loss function L used for the model training of S1 hist The histogram supervision is performed on the enhancement result, and the soft histogram operator is used to calculate the image histogram in the present application. The histogram loss function calculation method is:

[0081]

[0082] Among them, h indicates the histogram of the model enhancement result, Indicates the histogram of the true value image.

[0083] Figure 6 is the running result of a specific embodiment based on the method of the present application. Compared with the existing leading infrared enhancement method SNRLLE published in CVPR2022, the present method can effectively enhance low-contrast infrared images, effectively improve the spatial consistency and detail quality of infrared image enhancement, eliminate the problems of inconsistent pixel intensity in uniform areas of the enhancement result, artifacts, and blurred details, and effectively improve the visual effect of the enhancement result.

[0084] The embodiment also provides an infrared image contrast enhancement system for maintaining airspace consistency, comprising:

[0085] An infrared image enhancement model module, comprising a feature extraction subnetwork, a contrast enhancement subnetwork and a detail enhancement subnetwork, the feature extraction subnetwork is used for encoding a given infrared degraded image with low contrast and detail loss into a low-resolution feature expression, the contrast enhancement subnetwork is used for enhancing the low-resolution feature expression in a feature domain to obtain an enhanced feature representation, and the detail enhancement subnetwork is used for enhancing the enhanced feature representation in a hierarchical manner and reconstructing an original resolution image.

[0086] A model training module, used for training an infrared image enhancement model.

[0087] A contrast enhancement module, used for enhancing the contrast of a given infrared degraded image by using the trained infrared image enhancement model, and outputting an enhanced result.

[0088] For the system embodiment, since it basically corresponds to the method embodiment, the relevant part is described in the method embodiment, and the implementation method of the remaining modules is not described here. The system embodiment described above is only illustrative, and the units described as separate components can be or can not be physically separated, and the components displayed as units can be or can not be physical units, that is, they can be located in one place or distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the present application according to actual needs. Those skilled in the art can understand and implement it without creative labor.

[0089] The system embodiment of the present application can be applied to any device with data processing capability, which can be a device or apparatus such as a computer. The system embodiment can be realized by software, hardware or a combination of software and hardware. Taking software realization as an example, as a logical device, it is formed by reading the corresponding computer program instructions in the non-volatile memory into the memory and running by the processor of the device with data processing capability.

[0090] The above-described embodiments only express several embodiments of the present application, which are described in detail and specifically, but cannot be understood as limitations on the scope of the present application. For those skilled in the art, without departing from the concept of the present application, some modifications and improvements can be made, which are all within the protection scope of the present application.

Claims

1. A method for enhancing the contrast of infrared images while maintaining spatial consistency, characterized in that, Includes the following steps: Train an infrared image enhancement model, which includes a feature extraction subnetwork, a contrast enhancement subnetwork, and a detail enhancement subnetwork. The feature extraction subnetwork is used to encode a given low-contrast, detail-loss-degraded infrared image into a low-resolution feature representation, and the contrast enhancement subnetwork is used to enhance the low-resolution feature representation in the feature domain to obtain an enhanced feature representation. The detail enhancement subnet is used to progressively enhance the detail structure of the enhanced feature representations and reconstruct the original resolution image; The contrast enhancement subnet is composed of several cascaded global enhancement modules (GEMs). Each GEM maintains a spatially consistent descriptor through long-distance spatial information interaction and cross-channel information interaction, and performs global enhancement with consistency preservation on the feature domain. The GEM calculation process is as follows: Let the input of the nth GEM be... , These represent the length, width, and number of channels of the input image, respectively. A two-layer self-attention mechanism is employed: a spatial interaction layer and a channel interaction layer to achieve long-distance spatial information interaction and cross-channel information interaction, respectively. The attention map dimension of the spatial interaction layer is [dimension missing]. The attention graph dimension of the channel interaction layer is Then, the outputs of the two layers are added together to obtain the globally consistent descriptor. ; Using a curve generator from GEM input and globally consistent descriptors Generate enhancement curve tensor K represents the number of curve adjustment iterations; Using Enhanced Curve Tensor right Perform iterative curve adjustments, with the final adjustment result as the output; during iterative curve adjustments, the k-th adjustment uses the enhanced curve tensor. The k-th channel slice in the curve is used as the enhancement curve parameter, denoted as . Perform the following curve adjustments: ; in, This is the result of the k-th feature domain enhancement. This is element-wise multiplication; = ; The trained infrared image enhancement model is used to enhance the contrast of a given degraded infrared image and output the enhancement result.

2. The infrared image contrast enhancement method for maintaining spatial consistency according to claim 1, characterized in that, The feature extraction subnet uses nonlinear, inactive network blocks.

3. The infrared image contrast enhancement method for maintaining spatial consistency according to claim 1, characterized in that, The detail enhancement subnet consists of several cascaded Local Enhancement Modules (LEMs). Each LEM consists of a resolution reconstructor and a detail injector, progressively enhancing and reconstructing high-resolution details. The enhancement features of the last LEM are decoded into the enhanced result at the original resolution through a convolutional layer.

4. The infrared image contrast enhancement method for maintaining spatial consistency according to claim 3, characterized in that, The LEM calculation process is as follows: Enhanced features from low resolution using a resolution reconstructor Preliminary reconstruction of high-resolution features, denoted as ,Depend on Reconstructed Dimensions and The dimensions are the same; the resolution reconstructor is an upsampled convolutional layer. These are the input and output of the m-th LEM, respectively; The detail injector includes a dual attention layer and a weighted fusion operation. First, in the dual attention layer, and The detail-aware enhancement features are calculated from the query features Q, which are mutually neighboring attention features, and the key features K and V. and contrast perception enhancement features ;in It is extracted from infrared degraded images by convolutional layers and First, detailed features of the same dimension are generated; second, detail-aware enhanced features are generated through a weight generator. and contrast perception enhancement features The fusion weights are then calculated, and the two are weighted and fused to output high-resolution enhanced features containing rich detail information. .

5. The infrared image contrast enhancement method for maintaining spatial consistency according to claim 4, characterized in that, The loss functions used in training infrared image enhancement models include the global enhancement loss function, the local enhancement loss function, and the histogram loss function. Global augmentation loss function for: ; in, This indicates an enhancement result derived from low-resolution feature representation. This represents a convolutional decoding layer. This indicates downsampling the ground truth image to the original image size. Double the size to match the low-resolution enhanced results; It is an L1 norm; Set multi-scale features Size is the same as the original image Local augmentation loss function for: ; in, It is the number of LEMs. This indicates the enhancement result generated from multi-scale features. This indicates downsampling the ground truth image to its original size. The size is doubled to match the size of the multi-scale augmentation results. Indicates structural similarity; Histogram loss function for: ; in, Histogram representing the enhancement results, A histogram representing the truth image.

6. The infrared image contrast enhancement method for maintaining spatial consistency according to claim 1, characterized in that, It also includes the reconstruction loss function: ; in, It is a truth image. It enhances the results. These are the weighting coefficients for perceived loss. It uses a pre-trained VGG network to extract multi-layer features. It is a reconstruction loss.

7. An infrared image contrast enhancement system that maintains spatial consistency, for implementing the method of claim 1, Its features are, The system includes: The infrared image enhancement model module includes a feature extraction subnetwork, a contrast enhancement subnetwork, and a detail enhancement subnetwork. The feature extraction subnetwork is used to encode a given low-contrast, detail-loss-degraded infrared image into a low-resolution feature representation. The contrast enhancement subnetwork is used to enhance the low-resolution feature representation in the feature domain to obtain an enhanced feature representation. The detail enhancement subnetwork is used to progressively enhance the detail structure of the enhanced feature representation and reconstruct the original resolution image. The model training module is used to train an infrared image enhancement model. The contrast enhancement module is used to enhance the contrast of a given degraded infrared image using a trained infrared image enhancement model and output the enhancement result.