Infrared image enhancement method and system based on deep learning convolutional neural network
By constructing a feature transformation space and bidirectional feature mapping of a deep learning convolutional neural network, the problem of poor enhancement reliability in infrared image enhancement is solved, and high-quality infrared image enhancement effect is achieved.
Patent Information
- Application Number
- CN202511070733.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-07-31
Smart Images

Figure CN120912459A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of infrared image enhancement, and particularly relates to an infrared image enhancement method and system based on a deep learning convolutional neural network. BACKGROUND
[0002] With the development of computer vision technology, infrared imaging is increasingly widely applied in night vision monitoring, security and warning, industrial detection and other fields. As a key technology for improving the quality of infrared imaging, the processing effect of infrared image enhancement directly affects the performance of downstream vision tasks.
[0003] In related technologies, the quality of infrared images is mainly improved through image gray scale transformation and local information enhancement. Specifically, first, multi-scale gray scale histogram analysis is performed on the infrared image to obtain the brightness distribution characteristics and gray scale statistical information of the image at different spatial scales. Then, a non-linear gray scale mapping function is designed based on the global gray scale characteristics, and adaptive weight coefficients are constructed combining the statistical characteristics of local regions to dynamically adjust the parameters of the mapping function. Next, the image is processed in zones, and the enhancement intensity is determined according to the local characteristics of each zoned image. Then, the gray scale transformation and contrast enhancement processing of each zoned image are performed based on the adjusted mapping function. Finally, the processed zoned images are fused, and the edge regions of the image are smoothed to obtain the enhanced infrared image.
[0004] However, due to the inherent limitations of infrared imaging, there are obvious deficiencies in the enhancement processing relying only on the gray scale information of the image itself. On the one hand, the multi-scale gray scale analysis and local feature extraction in related technologies cannot accurately identify the key target structures in the image, which leads to the destruction of the original features of the image in the enhancement process. On the other hand, the adaptive weight adjustment based on preset rules lacks understanding of the semantic information of the image, and it is difficult to achieve effective enhancement while maintaining the true features of the image, especially under complex background and low signal-to-noise ratio conditions, which often leads to problems such as blurred image contours and lost details, thereby resulting in poor reliability of infrared image enhancement based on single modal processing in related technologies. SUMMARY
[0005] The present application provides an infrared image enhancement method and system based on a deep learning convolutional neural network, which is used to improve the reliability of infrared image enhancement.
[0006] In a first aspect, the application provides an infrared image enhancement method based on a deep learning convolutional neural network, applied to the infrared image enhancement system. The method comprises: in the case of obtaining an infrared image and a visible light image of a target monitoring area, constructing a feature transformation space, wherein the infrared image and the visible light image are spatio-temporally aligned registered images of the target monitoring area, the feature transformation space includes a first feature domain for storing the visible light image, a second feature domain for storing the infrared image, and an intermediate feature domain for storing common information of the first feature domain and the second feature domain; using a target feature mapping network to map the infrared image from the second feature domain to the intermediate feature domain to obtain a first intermediate feature, mapping the first intermediate feature to the first feature domain to obtain a first mapped image, and using the target feature mapping network to map the visible light image from the first feature domain to the intermediate feature domain to obtain a second intermediate feature, and mapping the second intermediate feature to the second feature domain to obtain a second mapped image; determining a first feature difference between the first mapped image and the visible light image, and determining a second feature difference between the second mapped image and the infrared image; and performing enhancement processing on the infrared image according to the first feature difference and the second feature difference to obtain a target enhanced infrared image.
[0007] By adopting the above technical solution, the visible light image and the infrared image are used to construct a feature transformation space including a first feature domain, a second feature domain and an intermediate feature domain, which provides a basic framework for feature mapping of different modal images. On this basis, the target feature mapping network can realize bidirectional feature mapping conversion between the infrared image and the visible light image, and through calculation of the feature difference between the mapped image and the original image, reliable evaluation basis can be provided for subsequent enhancement processing. This enhancement method based on dual modal information and feature mapping can not only make full use of the detailed information in the visible light image to make up for the deficiency of the infrared image, but also can ensure the controllability of the enhancement process through calculation of the feature difference, so as to realize high-quality enhancement of the infrared image, and can effectively avoid the loss of details and over-enhancement in the traditional single modal enhancement method. Further, the technical problem of poor enhancement reliability of the infrared image based on single modal processing in the related art is solved, and the technical effect of improving the enhancement reliability of the infrared image is achieved.
[0008] Optionally, in the case of obtaining the infrared image and the visible light image of the target monitoring area, a feature transformation space is constructed, specifically including: inputting the infrared image and the visible light image into a preset feature extraction network to obtain infrared deep layer feature maps and visible light deep layer feature maps output by the preset feature extraction network; performing texture feature processing on the visible light deep layer feature maps to obtain a first feature domain, wherein the first feature domain has visible light texture detail characteristics; performing heat intensity feature processing on the infrared deep layer feature maps to obtain a second feature domain, wherein the second feature domain has infrared heat radiation distribution characteristics; performing feature channel splicing and convolution fusion processing on the infrared deep layer feature maps and the visible light deep layer feature maps to extract cross-modal common structure features to obtain an intermediate feature domain, wherein the intermediate feature domain includes common structure characteristics of the infrared image and the visible light image.
[0009] By using the above technical solution, the infrared image and the visible light image are subjected to deep layer feature extraction by the preset feature extraction network, and then texture feature processing and heat intensity feature processing are performed, respectively, so that the first feature domain with visible light texture detail characteristics and the second feature domain with infrared heat radiation distribution characteristics can be obtained. In particular, by performing feature channel splicing and convolution fusion processing on the infrared deep layer feature maps and the visible light deep layer feature maps, cross-modal common structure features can be extracted to obtain the intermediate feature domain containing common structure characteristics of the two kinds of images. This domain construction method based on deep features can effectively separate and retain the feature attributes of different modal images, and at the same time, the association between the two modalities can be established through the intermediate feature domain, providing more accurate feature expression for subsequent feature mapping.
[0010] Optionally, the texture feature processing on the visible light deep layer feature maps to obtain the first feature domain specifically includes: performing multi-scale feature extraction on the visible light deep layer feature maps by using a plurality of convolution layers of different scales to obtain a texture feature map, wherein the texture feature map includes global texture structure features and local texture detail features; performing texture gradient analysis on the global texture structure features to generate a texture gradient statistical histogram; determining a texture segmentation threshold according to the peak value distribution of the texture gradient statistical histogram; dividing the texture feature map into M texture regions according to the texture segmentation threshold and the global texture structure features, wherein M is a positive integer; performing first boundary gradient detection on the M texture regions to determine a texture transition zone region between each two texture regions in the M texture regions, wherein the texture transition zone region is a fuzzy boundary region caused by texture gradient; comparing and enhancing the first texture feature of each texture transition zone region with the second texture feature of a first adjacent texture region and the third texture feature of a second adjacent texture region, and fusing the local texture detail features to generate a texture boundary enhanced feature map; performing hierarchical division on the texture boundary enhanced feature map according to the texture segmentation threshold to generate a texture hierarchical feature map; and determining the texture hierarchical feature map as the first feature domain.
[0011] By adopting the technical solution, the visible light deep feature map is subjected to multi-scale feature extraction by multiple convolution layers with different scales to obtain a texture feature map containing global texture structure features and local texture detail features. Through texture gradient analysis and segmentation processing, the texture feature map can be divided into multiple texture regions, and the boundary of the texture transition zone is enhanced. This multi-level texture feature processing method can comprehensively capture the texture information in the visible light image and provide rich texture detail reference for infrared image enhancement.
[0012] Optionally, the infrared deep feature map is subjected to heat intensity feature processing to obtain a second feature domain, specifically including: multi-scale feature extraction of the infrared deep feature map by multiple convolution layers with different expansion rates to obtain a heat radiation feature map, wherein the heat radiation feature map includes global heat radiation distribution features and local heat radiation detail features; heat intensity range analysis of the global heat radiation distribution features to generate a heat intensity statistical histogram; determination of a heat intensity segmentation threshold according to the peak value distribution of the heat intensity statistical histogram; division of the heat radiation feature map into N heat intensity regions according to the heat intensity segmentation threshold and the global heat radiation distribution features, wherein N is a positive integer; second boundary gradient detection of the N heat intensity regions to determine a heat intensity transition zone between each two of the N heat intensity regions, wherein the heat intensity transition zone is a blurred boundary region caused by heat diffusion; comparison and reconstruction of first heat radiation features of each heat intensity transition zone with second heat radiation features of a first adjacent heat intensity region and third heat radiation features of a second adjacent heat intensity region, and fusion of local heat radiation detail features to generate a heat boundary enhanced feature map; temperature layering of the heat boundary enhanced feature map according to the heat intensity segmentation threshold to generate a heat intensity layered feature map; and determination of the heat intensity layered feature map as the second feature domain.
[0013] By adopting the technical solution, the infrared deep feature map is subjected to multi-scale feature extraction by multiple convolution layers with different expansion rates to obtain a heat radiation feature map containing global heat radiation distribution features and local heat radiation detail features. Through heat intensity range analysis and segmentation processing, the heat radiation feature map can be divided into multiple heat intensity regions, and the heat intensity transition zone is subjected to comparison and reconstruction. This multi-level heat radiation feature processing method can accurately retain the temperature distribution characteristics of the infrared image and provide reliable heat radiation feature reference for feature mapping.
[0014] Optionally, the infrared image is mapped from the second feature domain to the intermediate feature domain by using the target feature mapping network to obtain first intermediate features, the first intermediate features are mapped to the first feature domain to obtain a first mapped image, and the visible light image is mapped from the first feature domain to the intermediate feature domain by using the target feature mapping network to obtain second intermediate features, the second intermediate features are mapped to the second feature domain to obtain a second mapped image, and the method specifically comprises the following steps: a multi-layer convolutional neural network is used to construct a first feature extraction encoder and a second feature extraction encoder, wherein the first feature extraction encoder and the second feature extraction encoder share convolutional layer parameters of the multi-layer convolutional neural network; a multi-layer deconvolutional neural network is used to construct a first feature reconstruction decoder and a second feature reconstruction decoder, wherein the first feature reconstruction decoder and the second feature reconstruction decoder share deconvolutional layer parameters of the multi-layer deconvolutional neural network; the first feature extraction encoder and the first feature reconstruction decoder are connected in series to construct a first mapping subnetwork, wherein the first mapping subnetwork is used to realize feature mapping from the second feature domain to the first feature domain; the second feature extraction encoder and the second feature reconstruction decoder are connected in series to construct a second mapping subnetwork, wherein the second mapping subnetwork is used to realize feature mapping from the first feature domain to the second feature domain; the infrared image is mapped from the second feature domain to the intermediate feature domain by using the first mapping subnetwork to obtain the first intermediate features, and the first intermediate features are mapped to the first feature domain by using the first mapping subnetwork to obtain the first mapped image; the visible light image is mapped from the first feature domain to the intermediate feature domain by using the second mapping subnetwork to obtain the second intermediate features, and the second intermediate features are mapped to the second feature domain by using the second mapping subnetwork to obtain the second mapped image.
[0015] By adopting the above technical solution, the multi-layer convolutional neural network and the deconvolutional neural network sharing parameters are used to construct the feature extraction encoder and the feature reconstruction decoder respectively, and the first mapping subnetwork and the second mapping subnetwork are constructed by being connected in series. The parameter sharing mechanism of the feature extraction encoder and the feature reconstruction decoder can ensure the consistency of the bidirectional feature mapping. The mapping structure design based on the deep neural network can establish the nonlinear mapping relationship between different feature domains, and can realize accurate conversion of the infrared image and the visible light image features.
[0016] Optionally, the first feature difference between the first mapping image and the visible light image is determined, and the second feature difference between the second mapping image and the infrared image is determined, specifically comprising: extracting a first shared feature vector corresponding to the first mapping image from the intermediate feature domain, and extracting a second shared feature vector corresponding to the visible light image from the intermediate feature domain; calculating the cosine similarity of the first shared feature vector and the second shared feature vector in the feature channel dimension to generate a visible light domain feature difference map; determining a first dynamic threshold according to the local variance distribution of the visible light domain feature difference map; marking the region lower than the first dynamic threshold in the visible light domain feature difference map as a first feature difference region; extracting a pseudo-infrared feature vector corresponding to the second mapping image from the second feature domain, and extracting a real infrared feature vector corresponding to the infrared image from the second feature domain; calculating the temperature distribution KL divergence of the pseudo-infrared feature vector and the real infrared feature vector to generate a thermal radiation fidelity difference map; determining a second dynamic threshold according to the thermal intensity distribution of the thermal radiation fidelity difference map; marking the region higher than the second dynamic threshold in the thermal radiation fidelity difference map as a second feature difference region; performing connected component analysis on the first feature difference region to determine the area distribution and local difference intensity of the first feature difference region; determining the first feature difference according to the area distribution and the local difference intensity; performing thermal radiation intensity gradient analysis on the second feature difference region to determine the gradient change rate and the thermal intensity difference amplitude of the second feature difference region; determining the second feature difference according to the gradient change rate and the thermal intensity difference amplitude.
[0017] By adopting the above technical solution, the feature vectors are extracted from the intermediate feature domain and the second feature domain respectively, and the feature difference region is determined by calculating the cosine similarity and the KL divergence. The connected component analysis and the thermal radiation intensity gradient analysis are performed on the feature difference region, and accurate feature difference information can be obtained. This multi-dimensional feature difference analysis method can comprehensively evaluate the quality of feature mapping and provide accurate difference guidance for enhancement processing.
[0018] Optionally, the infrared image is enhanced according to the first feature difference and the second feature difference to obtain a target enhanced infrared image, specifically including: constructing a visible light detail compensation template according to the first feature difference, wherein the visible light detail compensation template includes edge contour enhancement information and texture structure enhancement information; constructing an infrared fidelity correction template according to the second feature difference, wherein the infrared fidelity correction template includes thermal radiation distribution preservation information and temperature gradient correction information; performing weighted fusion on the visible light detail compensation template and the infrared fidelity correction template to generate a comprehensive enhancement template, wherein the weight coefficient of the weighted fusion is determined according to the relative intensity of the first feature difference and the second feature difference; performing spatial domain block processing on the infrared image to divide the infrared image into a plurality of image subblocks; determining the enhancement intensity value of the corresponding region of each image subblock in the comprehensive enhancement template, and determining the local enhancement parameter of each image subblock according to the enhancement intensity value; performing adaptive enhancement processing on each image subblock by using the local enhancement parameter to obtain an enhanced image subblock set; seamlessly splicing all the enhanced image subblocks in the enhanced image subblock set, and performing smoothing processing on the splicing boundary to obtain an enhanced splicing image; performing global tone mapping and dynamic range compression on the enhanced splicing image to generate the target enhanced infrared image.
[0019] By adopting the above technical solution, the visible light detail compensation template and the infrared fidelity correction template are constructed according to the feature differences, and the comprehensive enhancement template is generated by weighted fusion. The infrared image is processed by spatial domain block processing, and adaptive enhancement processing is performed according to the comprehensive enhancement template. This local adaptive enhancement method based on the template can effectively fuse the detail information of the visible light image while maintaining the thermal radiation characteristics of the infrared image, thereby improving the visual quality of the infrared image.
[0020] In a second aspect, the embodiments of the present application provide an infrared image enhancement system, which includes one or more processors and a memory. The memory is coupled with the one or more processors, and the memory is configured to store computer program codes including computer instructions. The one or more processors invoke the computer instructions to enable the infrared image enhancement system to perform the method described in the first aspect and any possible implementation manner of the first aspect.
[0021] In a third aspect, the embodiments of the present application provide a computer program product containing instructions, which, when the computer program product is executed on an infrared image enhancement system, enables the infrared image enhancement system to perform the method described in the first aspect and any possible implementation manner of the first aspect.
[0022] In a fourth aspect, the embodiments of the present application provide a computer readable storage medium, including instructions, when the instructions run on an infrared image enhancement system, causing the infrared image enhancement system to perform the method described in the first aspect and any possible implementation manner of the first aspect.
[0023] The one or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: 1. The infrared image enhancement method based on the deep learning convolutional neural network provided in the present application constructs a feature transformation space including a first feature domain, a second feature domain and an intermediate feature domain using the visible light image and the infrared image, to provide a basic framework for feature mapping of different modal images. On this basis, the target feature mapping network can realize bidirectional feature mapping conversion between the infrared image and the visible light image, and through calculation of feature differences between the mapped image and the original image, reliable evaluation basis can be provided for subsequent enhancement processing. This enhancement method based on dual modal information and feature mapping can not only make full use of the detail information in the visible light image to make up for the deficiency of the infrared image, but also can ensure the controllability of the enhancement process through calculation of feature differences, so that high-quality enhancement of the infrared image can be realized, and detail loss and over-enhancement in the traditional single modal enhancement method can be effectively avoided.
[0024] 2. The infrared image enhancement method based on the deep learning convolutional neural network provided in the present application uses a preset feature extraction network to perform deep feature extraction on the infrared image and the visible light image, and then performs texture feature processing and thermal intensity feature processing respectively, so that a first feature domain with visible light texture detail characteristics and a second feature domain with infrared thermal radiation distribution characteristics can be obtained. In particular, by performing feature channel splicing and convolution fusion processing on the infrared deep feature map and the visible light deep feature map, common structure feature of the two images can be extracted to obtain an intermediate feature domain containing common structure characteristics of the two images. This domain construction method based on deep features can effectively separate and retain the feature attributes of different modal images, and at the same time can establish the correlation between the two modalities through the intermediate feature domain, to provide more accurate feature expression for subsequent feature mapping.
[0025] 3. The infrared image enhancement method based on the deep learning convolutional neural network provided in the present application uses multiple convolution layers of different scales to perform multi-scale feature extraction on the visible light deep feature map, to obtain a texture feature map containing global texture structure features and local texture detail features. Through texture gradient analysis and segmentation processing, the texture feature map can be divided into multiple texture regions, and the boundary of the texture transition zone can be enhanced. This multi-level texture feature processing method can comprehensively capture the texture information in the visible light image, and can provide rich texture detail reference for infrared image enhancement. BRIEF DESCRIPTION OF DRAWINGS
[0026] Figure 1 is a flowchart of an infrared image enhancement method based on a deep learning convolutional neural network in an embodiment of the present application. Figure 2 is a schematic diagram of an entity device structure of an infrared image enhancement system in an embodiment of the present application. DETAILED DESCRIPTION
[0027] The terms used in the following embodiments of the present application are only for the purpose of describing specific embodiments and are not intended to be limiting of the present application. As used in the specification and the appended claims of the present application, the singular forms "a," "an" and "the" are intended to include both singular and plural forms, unless the context clearly indicates otherwise. It will be further understood that the terms "and / or" as used herein refers to any or all possible combinations of one or more of the associated listed items.
[0028] Hereinafter, the terms "first" and "second" are only for the purpose of description and cannot be understood as implying or suggesting relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined with "first" and "second" can explicitly or implicitly include one or more of the features, and in the description of the embodiments of the present application, the meaning of "a plurality of" is two or more, unless otherwise specified.
[0029] The present application provides an infrared image enhancement method based on a deep learning convolutional neural network, referring to Figure 1 , Figure 1 is a flowchart of an infrared image enhancement method based on a deep learning convolutional neural network in an embodiment of the present application, comprising the following steps: Step S101, in the case of obtaining an infrared image and a visible light image of a target monitoring area, a feature transformation space is constructed, wherein the infrared image and the visible light image are spatio-temporally aligned registered images of the target monitoring area, the feature transformation space includes a first feature domain for storing the visible light image, a second feature domain for storing the infrared image, and an intermediate feature domain for storing common information of the first feature domain and the second feature domain; Step S102, mapping the infrared image from the second feature domain to the intermediate feature domain using a target feature mapping network to obtain a first intermediate feature, mapping the first intermediate feature to the first feature domain to obtain a first mapped image, and mapping the visible light image from the first feature domain to the intermediate feature domain using the target feature mapping network to obtain a second intermediate feature, and mapping the second intermediate feature to the second feature domain to obtain a second mapped image; Step S103, determining a first feature difference between the first mapped image and the visible light image, and determining a second feature difference between the second mapped image and the infrared image; In step S104, the infrared image is enhanced according to the first feature difference and the second feature difference to obtain a target enhanced infrared image.
[0030] In the above embodiments, the target monitoring area refers to a specific scene range that needs image acquisition and processing; the infrared image refers to an image reflecting the thermal radiation characteristics of the target acquired by an infrared sensor; the visible light image refers to an image reflecting the visual features of the target acquired by a visible light camera; the spatio-temporally aligned registered image refers to an image pair that has undergone spatial position and time synchronization processing; the feature transformation space refers to a three-domain structure for storing different modal image features, including a first feature domain for storing visible light image features, a second feature domain for storing infrared image features, and an intermediate feature domain for storing common information of the two features; the first feature domain and the second feature domain respectively represent feature representation structures with specific attributes; the intermediate feature domain refers to a transitional feature space for storing shared information between different feature domains; the feature mapping refers to a feature conversion relationship established between different feature domains; and the feature difference refers to the degree of difference between different feature representations.
[0031] In the above embodiments, the problem of insufficient infrared image details in night monitoring scenes is solved. First, an infrared camera and a visible light camera installed in the monitoring area are used to capture an image pair, wherein the infrared camera captures a thermal imaging image with a resolution of 640x480, and the visible light camera captures a color image with a resolution of 1920x1080. A registration algorithm based on SIFT features is used to perform spatio-temporal alignment processing on the two images, and the visible light image is down-sampled to the same resolution as the infrared image to obtain a registered image pair. To achieve effective conversion of different modal image features, a three-domain feature transformation space is constructed. Specifically, a pre-trained VGG16 network is used to extract the conv4_3 layer features of the infrared image and the visible light image respectively to obtain 256-channel feature maps. The visible light feature map is subjected to 3x3 convolution and ReLU activation processing to obtain the first feature domain, and the infrared feature map is subjected to 3x3 convolution and ReLU activation to obtain the second feature domain. By concatenating the two feature maps in the channel dimension and then performing 1x1 convolution, 128-channel common structural features are extracted as the intermediate feature domain.
[0032] In the above embodiment, a target feature mapping network with an encoder-decoder structure is used to achieve feature transformation. The encoder consists of four 3×3 convolutional layers, each followed by BatchNorm and ReLU; the decoder uses four 3×3 deconvolutional layers, also configured with BatchNorm and ReLU. Infrared features are mapped to an intermediate feature domain by the encoder to obtain the first intermediate feature, and then mapped back to the first feature domain by the decoder to obtain the first mapped image. Similarly, visible light features are mapped through the same network to obtain the second intermediate feature and the second mapped image. When calculating the feature difference between the mapped image and the original image, the features of the conv4_3 layer are extracted and the cosine similarity is calculated to obtain a difference map. Adaptive threshold segmentation is performed on the difference map to mark the regions with significant differences. Combining statistical features such as the area ratio of the difference regions and the average difference degree, the feature difference value is quantified. Finally, an adaptive enhancement strategy is designed based on the feature differences: enhancing details and contrast in regions with significant differences, and maintaining the original features in regions with smaller differences, thereby improving the quality of the infrared image.
[0033] Through the above steps, a feature transformation space is constructed using visible light and infrared images, including a first feature domain, a second feature domain, and an intermediate feature domain, providing a basic framework for feature mapping of images of different modalities. Based on this, the target feature mapping network can achieve bidirectional feature mapping transformation between infrared and visible light images. By calculating the feature differences between the mapped image and the original image, a reliable evaluation basis can be provided for subsequent enhancement processing. This enhancement method based on dual-modal information and feature mapping not only fully utilizes the detailed information in the visible light image to compensate for the deficiencies of the infrared image, but also ensures the controllability of the enhancement process through the calculation of feature differences, thereby achieving high-quality enhancement of infrared images and effectively avoiding detail loss and over-enhancement in traditional single-modal enhancement methods. This solves the technical problem of poor reliability in infrared image enhancement based on single-modal processing in related technologies, achieving the technical effect of improving the reliability of infrared image enhancement.
[0034] The entity performing the above steps may be a system with infrared image enhancement capabilities, such as an infrared image enhancement system, or a device with infrared image enhancement capabilities, or a controller or processor in the device or system, or a separate controller or processor, or other processing devices or processing units with similar processing functions, but is not limited to these.
[0035] In an optional embodiment, in the case that the infrared image and the visible light image of the target monitoring area are acquired, a feature transformation space is constructed, specifically including: inputting the infrared image and the visible light image into a preset feature extraction network to acquire infrared deep layer feature maps and visible light deep layer feature maps output by the preset feature extraction network; performing texture feature processing on the visible light deep layer feature maps to obtain a first feature domain, wherein the first feature domain has visible light texture detail characteristics; performing heat intensity feature processing on the infrared deep layer feature maps to obtain a second feature domain, wherein the second feature domain has infrared heat radiation distribution characteristics; performing feature channel splicing and convolution fusion processing on the infrared deep layer feature maps and the visible light deep layer feature maps to extract cross-modal common structure features to obtain an intermediate feature domain, wherein the intermediate feature domain includes common structure characteristics of the infrared image and the visible light image.
[0036] In the above embodiments, the preset feature extraction network represents a neural network model that can extract deep layer features of an image after large-scale data training; the infrared deep layer feature map refers to a high-dimensional feature representation of an infrared image extracted by the preset feature extraction network; the visible light deep layer feature map refers to a high-dimensional feature representation of a visible light image extracted by the preset feature extraction network; the texture feature processing refers to feature extraction and preliminary enhancement operations on image texture structures; the heat intensity feature processing represents an extraction and analysis process of heat radiation intensity features; the feature channel splicing refers to combination of channel dimensions of different feature maps; the convolution fusion processing represents mixing and integration of features through convolution operations; the cross-modal common structure features refer to structural features commonly existing in different modal images.
[0037] In the above embodiments, in order to construct an effective feature transformation space, the following specific implementation scheme is adopted: first, an ImageNet pre-trained ResNet50 network is selected as a feature extraction network, and registered infrared images and visible light images with a size of 640x480 are input into the network. Deep layer features are extracted from the conv4_x layer of the network to obtain respective 40x30x1024-dimensional feature maps as infrared deep layer feature maps and visible light deep layer feature maps. Multi-scale convolution structure is used for texture feature processing of the visible light deep layer feature maps, including three parallel branches: the first branch uses 1x1 convolution to extract point-level features, the second branch uses 3x3 convolution to extract local texture features, and the third branch uses 5x5 convolution to extract texture structure features in a larger range. The output feature maps of the three branches are spliced in the channel dimension, and then fused through 1x1 convolution to obtain a 512-channel feature map. The original features and the fused features are added through residual connection, and finally activated through ReLU to obtain the first feature domain with rich texture details.
[0038] In the above embodiment, the heat intensity feature processing of the infrared deep feature map uses a multi-branch structure with dilated convolution: the first branch uses a standard 3x3 convolution, the second branch uses a 3x3 convolution with an expansion rate of 2, and the third branch uses a 3x3 convolution with an expansion rate of 4. This design can expand the receptive field without increasing the number of parameters, and better capture the distribution characteristics of thermal radiation. Similarly, the feature maps of the three branches are spliced and fused into 512 channels using a 1x1 convolution, and after a residual connection, a ReLU activation is performed to obtain a second feature domain that preserves the distribution characteristics of thermal radiation. To extract cross-modal common structural features, the infrared deep feature map and the visible light deep feature map are spliced in the channel dimension to obtain a 2048-channel feature map. Two layers of 3x3 convolution are used for feature fusion, the first layer outputs 512 channels, and the second layer outputs 256 channels, and BatchNorm and ReLU are used after each layer. The final 256-channel feature map, which contains the structural information common to both modal images, such as target outline, scene layout, etc., is the intermediate feature domain. This feature domain construction method can effectively separate and preserve the feature attributes of different modalities, and at the same time, it establishes a connection between the two modalities through the intermediate feature domain, providing a good foundation for subsequent feature mapping.
[0039] In an optional embodiment, the visible light deep feature map is processed for texture feature processing to obtain a first feature domain, specifically including: performing multi-scale feature extraction on the visible light deep feature map using multiple convolution layers of different scales to obtain a texture feature map, wherein the texture feature map includes global texture structure features and local texture detail features; performing texture gradient analysis on the global texture structure features to generate a texture gradient statistical histogram; determining a texture segmentation threshold according to the peak distribution of the texture gradient statistical histogram; dividing the texture feature map into M texture regions according to the texture segmentation threshold and the global texture structure features, wherein M is a positive integer; performing first boundary gradient detection on the M texture regions to determine a texture transition zone between each two texture regions in the M texture regions, wherein the texture transition zone is a blurred boundary region caused by texture gradient; comparing and enhancing the first texture feature of each texture transition zone with the second texture feature of the first adjacent texture region and the third texture feature of the second adjacent texture region, and fusing the local texture detail features to generate a texture boundary enhancement feature map; performing hierarchical division on the texture boundary enhancement feature map according to the texture segmentation threshold to generate a texture hierarchical feature map; and determining the texture hierarchical feature map as the first feature domain.
[0040] In the above embodiments, multi-scale feature extraction represents the process of feature extraction at different spatial scales; texture feature map refers to a feature representation that includes image texture information; global texture structure feature represents the texture organization feature of the whole image; local texture detail feature refers to the fine texture feature of the local region of the image; texture gradient analysis represents the analysis of the degree of texture change; texture gradient statistical histogram refers to a statistical histogram representing the distribution of texture gradient; texture segmentation threshold represents a threshold for distinguishing different texture regions; texture transition zone refers to the transition region between different texture regions; boundary gradient detection represents the detection of regional boundary change; texture boundary enhanced feature map refers to a texture feature representation enhanced by boundary; texture hierarchical feature map represents a hierarchical texture feature representation.
[0041] In the above embodiments, in order to realize effective texture feature processing of the visible light image, the following specific implementation scheme is adopted: first, a multi-scale feature extraction network is constructed, which contains three parallel convolution branches, using 1x1, 3x3 and 5x5 convolution kernels respectively, with a convolution step of 1 and output channel numbers of 128, 256 and 128 respectively. The output feature map of each branch is spliced in the channel dimension after BatchNorm and ReLU activation to obtain a texture feature map with 512 channels, of which the first 256 channels correspond to global texture structure features and the last 256 channels correspond to local texture detail features. When performing gradient analysis on the global texture structure features, Sobel gradients in horizontal and vertical directions are calculated respectively to obtain gradient amplitude and direction maps. The histogram of gradient amplitude is counted, and 32 evenly distributed histogram intervals are used to obtain the texture gradient statistical histogram. The peak value distribution of the histogram is analyzed by Otsu's method (OTSU) algorithm to determine three key thresholds as the texture segmentation thresholds, which are 0.15, 0.35 and 0.65. According to these thresholds, the texture feature map is divided into four texture regions (M=4): weak texture region (gradient value<0.15), medium texture region (0.15≤gradient value<0.35), strong texture region (0.35≤gradient value<0.65) and complex texture region (gradient value≥0.65). A 3x3 Sobel operator is used to detect the boundary gradient of each region, and the region with a gradient amplitude greater than 0.1 and less than 0.3 is marked as a texture transition zone.
[0042] In the above embodiment, the feature of each transition zone is enhanced: first, the first texture feature of the region (transition zone feature) is extracted, and the second and third texture features of the adjacent two texture regions are extracted. The mean and standard deviation of the three groups of features are calculated, the mean of the transition zone feature is adjusted to the weighted average of the mean of the adjacent region features, and the standard deviation is adjusted to the maximum of the adjacent region standard deviations. Then the enhanced transition zone feature is weighted and fused with the local texture detail feature, with weights of 0.7 and 0.3 respectively, to obtain a texture boundary enhancement feature map. Finally, according to the previously determined segmentation threshold, the enhanced feature map is divided into four levels to generate a texture layered feature map: the first layer is the basic texture layer (value range 0-0.15), the second layer is the detail texture layer (value range 0.15-0.35), the third layer is the main texture layer (value range 0.35-0.65), and the fourth layer is the fine texture layer (value range 0.65-1.0). The layered feature map is used as the first feature domain for subsequent feature mapping. Experiments show that this multi-level texture feature processing method can effectively preserve and enhance the texture information in the visible light image, providing a reliable texture reference for infrared image enhancement.
[0043] In an optional embodiment, the infrared deep layer feature map is processed for heat intensity feature to obtain a second feature domain, specifically including: a plurality of convolution layers with different expansion rates are used to perform multi-scale feature extraction on the infrared deep layer feature map to obtain a heat radiation feature map, wherein the heat radiation feature map includes global heat radiation distribution features and local heat radiation detail features; the global heat radiation distribution features are analyzed for heat intensity range to generate a heat intensity statistical histogram; a heat intensity segmentation threshold is determined according to the peak distribution of the heat intensity statistical histogram; the heat radiation feature map is divided into N heat intensity regions according to the heat intensity segmentation threshold and the global heat radiation distribution features, wherein N is a positive integer; second boundary gradient detection is performed on the N heat intensity regions to determine heat intensity transition zone regions between every two of the N heat intensity regions, wherein the heat intensity transition zone region is a blurred boundary region caused by heat diffusion; the first heat radiation feature of each heat intensity transition zone region is compared and reconstructed with the second heat radiation feature of the first adjacent heat intensity region and the third heat radiation feature of the second adjacent heat intensity region, and the local heat radiation detail feature is fused to generate a heat boundary enhancement feature map; the heat boundary enhancement feature map is temperature layered according to the heat intensity segmentation threshold to generate a heat intensity layered feature map; and the heat intensity layered feature map is determined as the second feature domain.
[0044] In the above embodiments, the dilation rate represents the dilation coefficient of the convolution kernel; the thermal radiation feature map refers to a feature representation characterizing the thermal radiation distribution; the global thermal radiation distribution feature representation represents the distribution characteristics of the overall temperature field; the local thermal radiation detail feature refers to the thermal radiation detail information of the local region; the thermal intensity range analysis represents the statistical analysis of the temperature intensity range; the thermal intensity statistical histogram refers to a statistical graph characterizing the temperature distribution; the thermal intensity segmentation threshold represents a threshold for distinguishing different temperature regions; the thermal intensity transition zone region refers to the transition region between different temperature regions; the heat diffusion refers to the propagation phenomenon of heat in space; the thermal boundary enhanced feature map represents the thermal feature representation after boundary enhancement; and the thermal intensity layered feature map refers to a feature representation with a temperature hierarchy.
[0045] In the above embodiments, to realize the thermal intensity feature processing of the infrared image, the following specific implementation scheme is adopted: first, a multi-branch dilated convolution network is constructed for multi-scale feature extraction, which includes four parallel branches, each branch using a 3x3 convolution kernel with dilation rates of 1, 2, 4, and 8, respectively, to ensure capturing thermal radiation features under different receptive fields. The output channel numbers of each branch are 128, 128, 128, and 128, respectively, which are spliced after BatchNorm and LeakyReLU (slope 0.2) activation to obtain a 512-channel thermal radiation feature map, of which the first 256 channels represent the global thermal radiation distribution feature and the last 256 channels represent the local thermal radiation detail feature. When analyzing the thermal intensity range of the global thermal radiation distribution feature, the feature values are first normalized to the [0, 1] interval, and then the distribution histogram of the thermal intensity values is counted. A total of 64 evenly distributed histogram intervals are used for statistics to obtain the thermal intensity statistical histogram. An adaptive threshold algorithm is applied to analyze the peak value distribution characteristics of the histogram to identify the main temperature distribution intervals. Four key thermal intensity segmentation thresholds are determined by the K-means clustering method, which are 0.2, 0.4, 0.6, and 0.8. According to these thresholds, the thermal radiation feature map is divided into five thermal intensity regions (N=5): background region (intensity value <0.2), low temperature region (0.2≤intensity value <0.4), medium temperature region (0.4≤intensity value <0.6), high temperature region (0.6≤intensity value <0.8), and extremely high temperature region (intensity value ≥0.8). An improved Canny operator is used to detect the boundary gradient of each thermal intensity region, with double thresholds set to 0.1 and 0.3, and the detected thermal intensity transition zone width is about 5% of the size of the feature map.
[0046] In the above embodiment, the feature reconstruction is performed for each thermal intensity transition zone: the first thermal radiation feature (transition zone feature) of the zone is extracted, and the second and third thermal radiation features of the two adjacent thermal intensity zones are extracted. The feature reconstruction method based on Gaussian mixture model is adopted to adjust the distribution parameters of the transition zone feature to the weighted combination of the distribution parameters of the adjacent zone features. Specifically, the mean value is weighted by distance, and the variance is the minimum value to ensure transition smoothing. The reconstructed transition zone feature is adaptively fused with the local thermal radiation detail feature, the fusion weight is dynamically adjusted according to the local temperature gradient, and the thermal boundary enhancement feature map is generated. Finally, the enhanced feature map is five-layer temperature layered according to the thermal intensity segmentation threshold to obtain the thermal intensity layered feature map: the first layer is the background temperature layer (0-0.2), the second layer is the low temperature feature layer (0.2-0.4), the third layer is the medium temperature feature layer (0.4-0.6), the fourth layer is the high temperature feature layer (0.6-0.8), and the fifth layer is the extremely high temperature feature layer (0.8-1.0). The layered feature map is used as the second feature domain. This multi-level thermal intensity feature processing method can effectively maintain the temperature distribution characteristics of the infrared image, while enhancing the clarity of the thermal target boundary.
[0047] In an optional embodiment, the infrared image is mapped from the second feature domain to the intermediate feature domain by using the target feature mapping network to obtain the first intermediate feature, the first intermediate feature is mapped to the first feature domain to obtain the first mapped image, and the visible light image is mapped from the first feature domain to the intermediate feature domain by using the target feature mapping network to obtain the second intermediate feature, and the second intermediate feature is mapped to the second feature domain to obtain the second mapped image, specifically comprising: a multi-layer convolutional neural network is used to construct a first feature extraction encoder and a second feature extraction encoder, wherein the first feature extraction encoder and the second feature extraction encoder share the convolutional layer parameters of the multi-layer convolutional neural network; a multi-layer deconvolutional neural network is used to construct a first feature reconstruction decoder and a second feature reconstruction decoder, wherein the first feature reconstruction decoder and the second feature reconstruction decoder share the deconvolutional layer parameters of the multi-layer deconvolutional neural network; the first feature extraction encoder and the first feature reconstruction decoder are connected in series to construct a first mapping subnetwork, wherein the first mapping subnetwork is used to realize the feature mapping from the second feature domain to the first feature domain; the second feature extraction encoder and the second feature reconstruction decoder are connected in series to construct a second mapping subnetwork, wherein the second mapping subnetwork is used for feature mapping from the first feature domain to the second feature domain; the infrared image is mapped from the second feature domain to the intermediate feature domain by using the first mapping subnetwork to obtain the first intermediate feature, and the first intermediate feature is mapped to the first feature domain by using the first mapping subnetwork to obtain the first mapped image; the visible light image is mapped from the first feature domain to the intermediate feature domain by using the second mapping subnetwork to obtain the second intermediate feature, and the second intermediate feature is mapped to the second feature domain by using the second mapping subnetwork to obtain the second mapped image.
[0048] In the above embodiments, the multi-layer convolutional neural network represents a deep learning model composed of multiple convolutional layers; the feature extraction encoder refers to a network module that converts input features into latent representations; the feature reconstruction decoder represents a network module that reconstructs latent representations into target features; the convolutional layer parameters refer to the weight and bias parameters in the convolutional neural network; the deconvolutional neural network represents a network that reconstructs features through deconvolution operations; the deconvolutional layer parameters refer to the weight and bias parameters in the deconvolutional network; the mapping sub-network represents a network structure that implements inter-domain mapping of specific features; parameter sharing refers to the use of the same network parameters by different network modules.
[0049] In the above embodiments, to achieve bidirectional mapping conversion between feature domains, an encoder-decoder network structure based on shared parameters is designed. Specifically, a feature extraction encoder is first constructed, containing 5 convolutional layers, with a convolution kernel size of 3x3, a step size of 1, and output channel numbers of 64, 128, 256, 512, and 256 in sequence. Each layer is followed by a BatchNorm and a PReLU activation function. To enhance feature extraction capability, a residual connection is added between the third and fourth convolutional layers. The first feature extraction encoder and the second feature extraction encoder completely share these convolutional layer parameters, ensuring consistency in feature extraction for different feature domains. The feature reconstruction decoder uses a symmetric 5-layer deconvolutional structure, with a deconvolution kernel size of 3x3, a step size of 1, and output channel numbers of 512, 256, 128, 64, and the original feature channel number in sequence. BatchNorm and PReLU activation are also used after each deconvolution layer, and a skip connection is added between the corresponding layers to preserve more feature details. The first feature reconstruction decoder and the second feature reconstruction decoder share all deconvolutional layer parameters, ensuring consistency in the feature reconstruction process.
[0050] In the above embodiment, the first feature extraction encoder and the first feature reconstruction decoder are connected in series to construct the first mapping sub-network, which is used to realize the mapping from the infrared feature domain to the visible light feature domain. A 1x1 convolution layer is added between the encoder and the decoder for feature channel reorganization and alignment. Similarly, the second feature extraction encoder and the second feature reconstruction decoder are connected in series to construct the second mapping sub-network, which realizes the mapping from the visible light feature domain to the infrared feature domain. The two mapping sub-networks adopt a symmetrical structure design to ensure the balance of the bidirectional feature mapping. When performing feature mapping, the second feature domain features of the infrared image are first input into the encoder of the first mapping sub-network to obtain 256-channel first intermediate features, which contain the main structural information of the infrared image. Then the first intermediate features are input into the decoder to reconstruct the first mapping image with the same dimension as the first feature domain through the deconvolution operation, which has the texture features of the visible light image. Similarly, the first feature domain features of the visible light image are input into the second mapping sub-network to obtain 256-channel second intermediate features through the encoder, and then the second mapping image with infrared features is reconstructed through the decoder. To improve the stability of feature mapping, the combination of cycle consistency loss and adversarial loss is used during network training. Cycle consistency ensures the reversibility of feature mapping, and adversarial loss promotes the generated features to conform to the distribution characteristics of the target domain. This bidirectional mapping network structure based on shared parameters can effectively realize the feature conversion between different feature domains and provide a reliable feature basis for subsequent image enhancement.
[0051] In an optional embodiment, the determining the first feature difference between the first mapping image and the visible light image and the determining the second feature difference between the second mapping image and the infrared image specifically comprises: extracting a first shared feature vector corresponding to the first mapping image from the intermediate feature domain and a second shared feature vector corresponding to the visible light image from the intermediate feature domain; calculating a cosine similarity of the first shared feature vector and the second shared feature vector in a feature channel dimension to generate a visible light domain feature difference map; determining a first dynamic threshold according to a local variance distribution of the visible light domain feature difference map; marking a region lower than the first dynamic threshold in the visible light domain feature difference map as a first feature difference region; extracting a pseudo-infrared feature vector corresponding to the second mapping image from the second feature domain and a real infrared feature vector corresponding to the infrared image from the second feature domain; calculating a temperature distribution KL divergence of the pseudo-infrared feature vector and the real infrared feature vector to generate a thermal radiation fidelity difference map; determining a second dynamic threshold according to a thermal intensity distribution of the thermal radiation fidelity difference map; marking a region higher than the second dynamic threshold in the thermal radiation fidelity difference map as a second feature difference region; performing connected component analysis on the first feature difference region to determine an area distribution and a local difference intensity of the first feature difference region; determining the first feature difference according to the area distribution and the local difference intensity; performing thermal radiation intensity gradient analysis on the second feature difference region to determine a gradient change rate and a thermal intensity difference amplitude of the second feature difference region; and determining the second feature difference according to the gradient change rate and the thermal intensity difference amplitude.
[0052] In the above embodiment, the shared feature vector represents a feature representation extracted from the intermediate feature domain; the feature channel dimension refers to a channel direction of the feature representation; the cosine similarity represents an index for measuring the similarity of the feature vectors; the visible light domain feature difference map refers to an image representing the feature difference in the visible light domain; the local variance distribution represents the variation degree of the local region of the image; the first dynamic threshold represents a threshold for feature difference region division dynamically determined according to the local variance distribution of the visible light domain feature difference map; the second dynamic threshold refers to a threshold for feature difference region division dynamically determined according to the thermal intensity distribution of the thermal radiation fidelity difference map; the pseudo-infrared feature vector represents an infrared domain feature generated by mapping; the real infrared feature vector refers to a feature extracted from the original infrared image; the KL divergence represents an index for measuring the difference of probability distribution; the thermal radiation fidelity difference map refers to an image representing the feature difference of thermal radiation; the connected component analysis represents the analysis of the connectivity of the image region; and the gradient change rate refers to the change speed of the feature gradient.
[0053] In the above embodiment, in order to accurately evaluate the effect of feature mapping, an evaluation scheme based on multi-dimensional feature difference analysis is designed. First, the feature vector is extracted from the intermediate feature domain. For the first mapped image and the original visible light image, 256-dimensional first shared feature vector and second shared feature vector are extracted respectively. The cosine similarity of the two vectors is calculated in the feature channel dimension to obtain a visible light domain feature difference map with the same size as the original image. By calculating the local 9x9 window variance of the difference map and using an adaptive threshold method to determine the first dynamic threshold, the lower quartile value of the local variance distribution is about 0.3. The area in the difference map with a similarity lower than the threshold is marked as the first feature difference area. At the same time, the pseudo-infrared feature vector of the second mapped image and the real infrared feature vector of the original infrared image are extracted from the second feature domain, each with a dimension of 512. The temperature distribution KL divergence (Kullback-Leibler divergence) of the two feature vectors is calculated to generate a thermal radiation fidelity difference map. The thermal intensity histogram distribution of the difference map is analyzed, and the upper quartile value is selected as the second dynamic threshold, which is about 0.7. The area in the difference map with a divergence value higher than the threshold is marked as the second feature difference area.
[0054] In the above embodiment, the first feature difference area is subjected to connected component analysis: first, the difference area is binarized, and then an 8-connected region labeling algorithm is used to identify connected regions. The area ratio of each connected region is counted, and the average difference intensity in the region is calculated. The area with an area ratio greater than 5% and an average difference intensity greater than 0.5 is regarded as the main difference area. According to the number of main difference areas, the total area ratio and the average difference intensity, the first feature difference value is calculated by weighted summation. The thermal radiation intensity gradient analysis is performed on the second feature difference area: the Sobel operator is used to calculate the gradient map of the difference area, and the spatial distribution of the gradient change rate is counted. At the same time, the difference amplitude of the thermal intensity value in the difference area and the surrounding area is calculated. Specifically, the gradient change rate takes the local mean of the gradient amplitude, and the thermal intensity difference amplitude takes the mean of the intensity difference between the target area and the background area. The gradient change rate and the thermal intensity difference amplitude are combined through a nonlinear mapping function to obtain the second feature difference value. This multi-dimensional feature difference analysis method can comprehensively evaluate the quality of feature mapping, especially in terms of preserving texture details and thermal radiation characteristics. The first feature difference can reflect the accuracy of texture mapping, and the second feature difference can evaluate the degree of preservation of thermal radiation characteristics, providing a reliable reference for subsequent adaptive enhancement.
[0055] In an optional embodiment, the infrared image is enhanced according to the first feature difference and the second feature difference to obtain a target enhanced infrared image, specifically including: constructing a visible light detail compensation template according to the first feature difference, wherein the visible light detail compensation template includes edge contour enhancement information and texture structure enhancement information; constructing an infrared fidelity correction template according to the second feature difference, wherein the infrared fidelity correction template includes thermal radiation distribution maintaining information and temperature gradient correction information; performing weighted fusion of the visible light detail compensation template and the infrared fidelity correction template to generate a comprehensive enhancement template, wherein the weight coefficient of the weighted fusion is determined according to the relative intensity of the first feature difference and the second feature difference; performing spatial domain block processing on the infrared image to divide the infrared image into a plurality of image subblocks; determining the enhancement intensity value of the corresponding region of each image subblock in the comprehensive enhancement template, and determining the local enhancement parameter of each image subblock according to the enhancement intensity value; performing adaptive enhancement processing on each image subblock by using the local enhancement parameter to obtain an enhanced image subblock set; seamlessly splicing all enhanced image subblocks in the enhanced image subblock set, and performing smoothing processing on the splicing boundary to obtain an enhanced splicing image; performing global tone mapping and dynamic range compression on the enhanced splicing image to generate the target enhanced infrared image.
[0056] In the above embodiments, the detail compensation template represents a reference template for enhancing image details; the edge contour enhancement information refers to feature information for enhancing target edges; the texture structure enhancement information represents feature information for enhancing texture details; the fidelity correction template refers to a reference template for maintaining thermal radiation characteristics; the thermal radiation distribution maintaining information represents feature information for maintaining temperature distribution; the temperature gradient correction information refers to feature information for correcting temperature transition; the weighted fusion represents a feature combination method based on weights; the spatial domain block refers to a processing of dividing an image into sub-regions; the local enhancement parameter represents a parameter for sub-region enhancement; the seamless splicing refers to image fusion for eliminating splicing traces; the tone mapping represents an adjustment method of image tone; and the dynamic range compression refers to a processing of adjusting image dynamic range.
[0057] In the above embodiment, in order to realize adaptive enhancement of infrared images, a multi-level enhancement scheme based on dual-template fusion is designed. First, a visible light detail compensation template is constructed according to the first feature difference. The template contains two parts of information: edge contour enhancement information is extracted by a Laplace operator, and the enhancement intensity is adaptively adjusted according to the local contrast; the texture structure enhancement information is extracted by a local binary pattern (LBP) operator, and the texture enhancement degree is adjusted by wavelet transform. According to the second feature difference, an infrared fidelity correction template is constructed. The template also contains two types of information: heat radiation distribution preservation information is obtained based on heat intensity histogram equalization, which ensures the original temperature distribution characteristics after enhancement; temperature gradient correction information is calculated by an improved gradient operator, which is used to correct the transition characteristics of the thermal target boundary. Both templates are normalized to the [0, 1] interval to ensure the consistency of the numerical range. The two templates are adaptively weighted and fused to generate a comprehensive enhancement template. The weight coefficient is dynamically determined according to the relative strength of the feature difference: when the first feature difference is dominant, the weight of the visible light detail compensation template is larger, with a value range of 0.6-0.8; when the second feature difference is significant, the weight of the infrared fidelity correction template increases, with a value range of 0.5-0.7. The specific weight is obtained by mapping the feature difference ratio through a sigmoid function.
[0058] In the above embodiment, the original infrared image is spatially divided into blocks, and a 32x32 pixel overlapping sliding window is used, with an adjacent window overlap rate of 25%. For each image sub-block, the local enhancement parameter is designed according to the average enhancement intensity value of the corresponding region in the comprehensive enhancement template. The enhancement parameters include contrast gain coefficient (range 1.0-2.0), detail enhancement coefficient (range 0.5-1.5) and smoothing factor (range 0.1-0.3). The local enhancement parameters are used to adaptively enhance each image sub-block: first, the contrast is adjusted, then the detail enhancement is performed, and finally the smoothing processing is applied to suppress noise. The enhanced sub-blocks are seamlessly spliced using a Poisson equation-based image fusion algorithm, and a weighted average strategy is used in the overlapping area to ensure smooth transition. The guide filter is used to smooth the splicing boundary to eliminate possible splicing marks. Finally, the enhanced spliced image is globally optimized: adaptive gamma correction is used for tone mapping, with a gamma value range of 0.6-1.4; an improved local histogram equalization is used to realize dynamic range compression, which ensures the visibility of details while avoiding over-enhancement. This enhancement scheme can effectively improve the visual quality of infrared images, both maintaining the thermal radiation characteristics and enhancing the image details, providing better image input for subsequent target recognition tasks.
[0059] By means of the embodiments of the present application, a feature transformation space is constructed by using the visible light image and the infrared image, including a first feature domain, a second feature domain and an intermediate feature domain, to provide a basic framework for feature mapping of different modal images. On this basis, the target feature mapping network can realize bidirectional feature mapping conversion between the infrared image and the visible light image, and through calculation of feature differences between the mapping image and the original image, reliable evaluation basis can be provided for subsequent enhancement processing. This enhancement method based on dual modal information and feature mapping can not only make full use of the detailed information in the visible light image to make up for the deficiency of the infrared image, but also can guarantee the controllability of the enhancement process through calculation of feature differences, so that high-quality enhancement of the infrared image can be realized, and loss of details and over-enhancement in the traditional single modal enhancement method can be effectively avoided.
[0060] The infrared image enhancement system in the embodiments of the present application is described below from the perspective of hardware processing, with reference to Figure 2 , Figure 2 FIG. 1 is a schematic structural diagram of an entity device of the infrared image enhancement system in the embodiments of the present application.
[0061] It should be noted that Figure 2 the structure of the infrared image enhancement system shown is only an example, and should not bring any limitation to the functions and use range of the embodiments of the present application.
[0062] As shown in Figure 2 , the infrared image enhancement system includes a central processing unit (CPU) 201, which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 202 or loaded into a random access memory (RAM) 203 from a storage part 208, such as the method described in the above embodiments. In the RAM 203, various programs and data required for system operation are also stored. The CPU 201, the ROM 202 and the RAM 203 are connected to each other through a bus 204. An input / output (I / O) interface 205 is also connected to the bus 204.
[0063] The following components are connected to the I / O interface 205: an input section 206 including an audio input device, a push button switch, and the like; an output section 207 including a Liquid Crystal Display (LCD), and an audio output device, a lamp, and the like; a storage section 208 including a hard disk and the like; and a communication section 209 including a network interface card such as a LAN (Local Area Network) card, a modem, and the like. The communication section 209 performs a communication process via a network such as the Internet. A drive 210 is also connected to the I / O interface 205 as necessary. A removable recording medium 211 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, and the like is attached to the drive 210 as necessary so that a computer program read therefrom is installed into the storage section 208 as necessary.
[0064] In particular, the processes described above with reference to the flow charts can be implemented as a computer software program in accordance with embodiments of the present application. For example, an embodiment of the present application includes a computer program product comprising a computer program carried on a computer readable medium, the computer program containing a computer program for executing the methods shown in the flow charts. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 209, and / or installed from the removable recording medium 211. When the computer program is executed by the central processing unit (CPU) 201, various functions defined in the present application are performed.
[0065] Note that specific examples of the computer readable storage medium can include but are not limited to one or more of a conduit with one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read only memory (ROM), an erasable programmable read only memory (EPROM), a flash memory, a fiber optic device, a portable compact disc read only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the present application, a computer readable storage medium can be any tangible medium that can contain, or store a program for use by or in connection with an instruction execution system, apparatus, or device.
[0066] The flowcharts and block diagrams in the drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various embodiments of the present application. In this regard, each block in the flowcharts or block diagrams can represent a module, segment, or portion of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved.
[0067] In particular, the infrared image enhancement system of the embodiment includes a processor and a memory, and the memory stores a computer program. When the computer program is executed by the processor, the infrared image enhancement method based on the deep learning convolutional neural network provided in the above embodiment is implemented.
[0068] As another aspect, the present application further provides a computer readable storage medium. The storage medium can be included in the infrared image enhancement system described in the above embodiments, or can exist independently without being assembled into the infrared image enhancement system. The storage medium carries one or more computer programs. When the one or more computer programs are executed by a processor of the infrared image enhancement system, the infrared image enhancement system implements the infrared image enhancement method based on the deep learning convolutional neural network provided in the above embodiments.
[0069] The above embodiments are only used to illustrate the technical solutions of the present application, but not limit the present application; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements to some technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.
[0070] Those skilled in the art can understand that all or part of the flow of the above-mentioned embodiment method can be instructed by a computer program to relevant hardware to complete, and the program can be stored in a computer readable storage medium. The program can include the flow of each method embodiment as described above when executed. The foregoing storage medium includes ROM or random storage memory RAM, magnetic disk or optical disk and various program code storage media.
Claims
1. An infrared image enhancement method based on a deep learning convolutional neural network, characterized in that, The method comprises the steps of: In the case of obtaining the infrared image and the visible light image of the target monitoring area, Constructing a feature transformation space, wherein the infrared image and the visible light image are spatio-temporally aligned registered images of the target monitoring area, the feature transformation space comprises a first feature domain for storing visible light images, a second feature domain for storing infrared images, and an intermediate feature domain for storing common information of the first feature domain and the second feature domain; Using a target feature mapping network to map the infrared image from the second feature domain to the intermediate feature domain to obtain first intermediate features, mapping the first intermediate features to the first feature domain to obtain a first mapped image, and using the target feature mapping network to map the visible light image from the first feature domain to the intermediate feature domain to obtain second intermediate features, and mapping the second intermediate features to the second feature domain to obtain a second mapped image; Determining the first feature difference between the first mapped image and the visible light image, and determining the second feature difference between the second mapped image and the infrared image; According to the first feature difference and the second feature difference, the infrared image is enhanced to obtain a target enhanced infrared image.
2. The method of claim 1, wherein, In the case of obtaining the infrared image and the visible light image of the target monitoring area, the feature transformation space is constructed, specifically comprising: Inputting the infrared image and the visible light image into a preset feature extraction network to obtain infrared deep feature maps and visible light deep feature maps output by the preset feature extraction network; Performing texture feature processing on the visible light deep feature maps to obtain the first feature domain, wherein the first feature domain has visible light texture detail characteristics; Performing heat intensity feature processing on the infrared deep feature maps to obtain the second feature domain, wherein the second feature domain has infrared heat radiation distribution characteristics; Performing feature channel splicing and convolution fusion processing on the infrared deep feature maps and the visible light deep feature maps to extract cross-modal common structure features to obtain the intermediate feature domain, wherein the intermediate feature domain includes common structure characteristics of the infrared image and the visible light image.
3. The method of claim 2, wherein, The texture feature processing on the visible light deep feature maps to obtain the first feature domain specifically comprises: Using multiple convolution layers of different scales to perform multi-scale feature extraction on the visible light deep feature maps to obtain texture feature maps, wherein the texture feature maps include global texture structure features and local texture detail features; Performing texture gradient analysis on the global texture structure features to generate a texture gradient statistical histogram; Determine the texture segmentation threshold according to the peak distribution of the texture gradient statistical histogram; Divide the texture feature maps into M texture regions according to the texture segmentation threshold and the global texture structure features, wherein M is a positive integer; Performing first boundary gradient detection on the M texture regions to determine the texture transition zone between each two texture regions in the M texture regions, wherein the texture transition zone is a blurred boundary region caused by texture gradient; The first texture feature of each texture transition zone is compared and enhanced with the second texture feature of the first adjacent texture zone and the third texture feature of the second adjacent texture zone, and the local texture detail features are fused to generate a texture boundary enhanced feature map; The texture boundary enhanced feature map is hierarchically divided according to the texture segmentation threshold to generate a texture hierarchical feature map; The texture hierarchical feature map is determined as the first feature domain.
4. The method of claim 2, wherein, The infrared deep layer feature map is processed by heat intensity feature processing to obtain the second feature domain, specifically including: A plurality of convolution layers with different expansion rates are used to perform multi-scale feature extraction on the infrared deep layer feature map to obtain a thermal radiation feature map, wherein the thermal radiation feature map includes global thermal radiation distribution features and local thermal radiation detail features; The global thermal radiation distribution features are analyzed to generate a heat intensity statistical histogram; A heat intensity segmentation threshold is determined according to the peak distribution of the heat intensity statistical histogram; The thermal radiation feature map is divided into N heat intensity regions according to the heat intensity segmentation threshold and the global thermal radiation distribution features, wherein N is a positive integer; Second boundary gradient detection is performed on the N heat intensity regions to determine the heat intensity transition zone between each two heat intensity regions, wherein the heat intensity transition zone is a blurred boundary region caused by heat diffusion; The first thermal radiation feature of each heat intensity transition zone is compared and reconstructed with the second thermal radiation feature of the first adjacent heat intensity region and the third thermal radiation feature of the second adjacent heat intensity region, and the local thermal radiation detail features are fused to generate a thermal boundary enhanced feature map; The thermal boundary enhanced feature map is temperature layered according to the heat intensity segmentation threshold to generate a heat intensity hierarchical feature map; The heat intensity hierarchical feature map is determined as the second feature domain.
5. The method of claim 1, wherein, The infrared image is mapped from the second feature domain to the intermediate feature domain by the target feature mapping network to obtain a first intermediate feature, the first intermediate feature is mapped to the first feature domain to obtain a first mapped image, and the visible light image is mapped from the first feature domain to the intermediate feature domain by the target feature mapping network to obtain a second intermediate feature, and the second intermediate feature is mapped to the second feature domain to obtain a second mapped image, specifically including: A multi-layer convolutional neural network is used to construct a first feature extraction encoder and a second feature extraction encoder, wherein the first feature extraction encoder and the second feature extraction encoder share convolution layer parameters of the multi-layer convolutional neural network; A multi-layer deconvolutional neural network is used to construct a first feature reconstruction decoder and a second feature reconstruction decoder, wherein the first feature reconstruction decoder and the second feature reconstruction decoder share deconvolution layer parameters of the multi-layer deconvolutional neural network; The first feature extraction encoder and the first feature reconstruction decoder are connected in series to construct a first mapping subnetwork, wherein the first mapping subnetwork is used to realize feature mapping from the second feature domain to the first feature domain; The second feature extraction encoder and the second feature reconstruction decoder are connected in series to construct a second mapping subnetwork, wherein the second mapping subnetwork is used for feature mapping from the first feature domain to the second feature domain; The first intermediate feature is obtained by mapping the infrared image from the second feature domain to the intermediate feature domain by using the first mapping subnetwork, and the first mapping image is obtained by mapping the first intermediate feature from the intermediate feature domain to the first feature domain by using the first mapping subnetwork; The second intermediate feature is obtained by mapping the visible light image from the first feature domain to the intermediate feature domain by using the second mapping subnetwork, and the second mapping image is obtained by mapping the second intermediate feature from the intermediate feature domain to the second feature domain by using the second mapping subnetwork.
6. The method of claim 1, wherein, The first feature difference between the first mapping image and the visible light image is determined, and the second feature difference between the second mapping image and the infrared image is determined, specifically including: A first shared feature vector corresponding to the first mapping image is extracted from the intermediate feature domain, and a second shared feature vector corresponding to the visible light image is extracted from the intermediate feature domain; The cosine similarity of the first shared feature vector and the second shared feature vector in the feature channel dimension is calculated to generate a visible light domain feature difference map; A first dynamic threshold is determined according to the local variance distribution of the visible light domain feature difference map; Regions in the visible light domain feature difference map lower than the first dynamic threshold are marked as first feature difference regions; A pseudo-infrared feature vector corresponding to the second mapping image is extracted from the second feature domain, and a real infrared feature vector corresponding to the infrared image is extracted from the second feature domain; The temperature distribution KL divergence of the pseudo-infrared feature vector and the real infrared feature vector is calculated to generate a thermal radiation fidelity difference map; A second dynamic threshold is determined according to the thermal intensity distribution of the thermal radiation fidelity difference map; Regions in the thermal radiation fidelity difference map higher than the second dynamic threshold are marked as second feature difference regions; Connected component analysis is performed on the first feature difference regions to determine the area distribution and local difference intensity of the first feature difference regions; The first feature difference is determined according to the area distribution and the local difference intensity; Thermal radiation intensity gradient analysis is performed on the second feature difference regions to determine the gradient change rate and thermal intensity difference amplitude of the second feature difference regions; The second feature difference is determined according to the gradient change rate and the thermal intensity difference amplitude.
7. The method of claim 1, wherein, The infrared image is enhanced according to the first feature difference and the second feature difference to obtain a target enhanced infrared image, specifically including: A visible light detail compensation template is constructed according to the first feature difference, wherein the visible light detail compensation template includes edge contour enhancement information and texture structure enhancement information; An infrared fidelity correction template is constructed according to the second feature difference, wherein the infrared fidelity correction template includes thermal radiation distribution retention information and temperature gradient correction information; The visible light detail compensation template and the infrared fidelity correction template are fused by weighting to generate a comprehensive enhancement template, wherein a weight coefficient of the weighted fusion is determined according to relative intensities of the first feature difference and the second feature difference; The infrared image is subjected to spatial domain block processing to divide the infrared image into a plurality of image sub-blocks; An enhancement intensity value of a corresponding region of each image sub-block in the comprehensive enhancement template is determined, and a local enhancement parameter of the each image sub-block is determined according to the enhancement intensity value; The each image sub-block is subjected to adaptive enhancement processing by using the local enhancement parameter to obtain a set of enhanced image sub-blocks; All the enhanced image sub-blocks in the set of enhanced image sub-blocks are seamlessly spliced, and a splicing boundary is subjected to smoothing processing to obtain an enhanced splicing image; The enhanced splicing image is subjected to global tone mapping and dynamic range compression to generate the target enhanced infrared image.
8. An infrared image intensifier system characterized by The infrared image enhancement system comprises one or more processors and a memory; the memory is coupled with the one or more processors, the memory is used to store computer program code, the computer program code comprises computer instructions, and the one or more processors invoke the computer instructions to enable the infrared image enhancement system to perform the method according to any one of claims 1-7.
9. A computer-readable storage medium comprising instructions, characterized in that, The instructions enable the infrared image enhancement system to perform the method according to any one of claims 1-7 when the instructions run on the infrared image enhancement system.
10. A computer program product, characterised in that, The computer program product enables the infrared image enhancement system to perform the method according to any one of claims 1-7 when the computer program product runs on the infrared image enhancement system.
Citation Information
Patent Citations
Infrared and visible light image fusion method based on multi-scale generative adversarial network
CN111145131A
Thermal infrared-visible light cross-modal face recognition method
CN114898429A
Charging temperature detection method, system and device of charging gun and storage medium
CN118037674A
Weak light target detection method based on infrared and visible light image fusion
CN118135200A
Low-light image enhancement method and device based on infrared visible light information integration
CN119831912A
Cited By
Power grid equipment infrared image enhancement method based on single-mode scarce sample generation
CN121073858A
Infrared Image Enhancement Method for Power Grid Equipment Based on Single-Modal Sparse Samples
CN121073858B