An image fusion method and apparatus based on infrared polarization imaging

By decomposing and optimizing the features based on the polarization bidirectional reflectance distribution function, a fusion weight map is generated for image fusion. This solves the problem of insufficient information richness and accuracy in traditional infrared polarization imaging fusion methods, and realizes fused images with high information entropy and strong edge fidelity, meeting the needs of high-precision target recognition and complex environment detection.

CN120912453BActive Publication Date: 2026-04-03BEIJING DONGYU HONGDA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Traditional infrared polarization imaging fusion methods produce images with low information richness and accuracy, making it difficult to meet the needs of various application scenarios.

Method used

Based on the bidirectional polarization reflectance distribution function, infrared polarized image data is decomposed into primitive images with three physical attributes: total radiation intensity, degree of linear polarization, and polarization angle. These primitive images are then subjected to feature optimization processing through a preset task mode, and information abundance and complementarity measures are calculated to generate a fusion weight map for weighted fusion.

Benefits of technology

It ensures the physical authenticity and completeness of information representation, significantly improves the saliency of key task features, and the fused image has high information entropy, strong edge fidelity and task-related feature enhancement characteristics, meeting the needs of high-precision target recognition and complex environment detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120912453B_ABST
    Figure CN120912453B_ABST
Patent Text Reader

Abstract

This application discloses an image fusion method and apparatus based on infrared polarization imaging, relating to the field of image processing technology. In this method, infrared polarization image data of a target scene is acquired; it is decomposed into three types of original primitive images based on the bidirectional polarization reflectance distribution function: total radiation intensity, degree of linear polarization, and polarization angle; the primitive images are processed according to a preset task mode to obtain feature primitive images; the regional information abundance metric and the pairwise regional complementarity metric of each feature primitive image are calculated; a fusion weight map is generated using a normalized exponential function, and the fused image is output after weighted fusion of the feature primitive images. This application solves the problems of low image information richness and accuracy in traditional methods by using physical feature decomposition, task-specific processing, and quantified fusion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of image processing, specifically to an image fusion method and apparatus based on infrared polarization imaging. Background Technology

[0002] In the field of image fusion technology, its application scope is becoming increasingly wide with the continuous development of technology. Image fusion technology can integrate information from multiple source images, providing richer and more accurate data support for subsequent analysis and decision-making. In the area of ​​image fusion based on infrared polarization imaging, traditional image fusion is usually based on feature extraction, which extracts feature information from images, such as edges and textures, and then performs image fusion based on these features.

[0003] However, the image information richness and accuracy of the images fused by the above image fusion methods are relatively low, making it difficult to meet the needs of various application scenarios. Summary of the Invention

[0004] To address the aforementioned technical problems, this application provides an image fusion method and apparatus based on infrared polarization imaging.

[0005] The first aspect of this application provides an image fusion method based on infrared polarization imaging:

[0006] Acquire infrared polarization image data of the target scene;

[0007] Based on the bidirectional polarization reflection distribution function, the infrared polarization image data is decomposed into a first feature primitive image, a second feature primitive image, and a third feature primitive image.

[0008] According to the preset task mode, feature optimization processing is performed on the first feature original primitive image, the second feature original primitive image and the third feature original primitive image respectively to obtain the first feature primitive image, the second feature primitive image and the third feature primitive image;

[0009] Calculate the first region information abundance metric of the first feature primitive image, the second region information abundance metric of the second feature primitive image, and the third region information abundance metric of the third feature primitive image, respectively.

[0010] Calculate a first region complementarity measure between the first feature primitive image and the second feature primitive image, a second region complementarity measure between the first feature primitive image and the third feature primitive image, and a third region complementarity measure between the second feature primitive image and the third feature primitive image;

[0011] The information abundance measures of the first region, the second region, the third region, the first region complementarity measure, the second region complementarity measure, and the third region complementarity measure are processed by a normalized exponential function to generate a fusion weight map.

[0012] The first feature primitive image, the second feature primitive image, and the third feature primitive image are weighted and fused using the fusion weight map to output a fused image.

[0013] By employing the aforementioned technical solution, infrared polarization data is decomposed based on a physical model of the bidirectional polarization reflection distribution function. This separates original primitive images from the source, yielding complementary physical attributes of total radiation intensity (first feature), linear polarization degree (second feature), and polarization angle (third feature), ensuring the physical authenticity and completeness of the information representation. Targeted enhancement processing of the primitive images is performed using a preset task mode, significantly improving the saliency of key task features. Quantifying regional information abundance and cross-feature complementarity accurately characterizes the contribution value of each primitive image in different spatial regions. A fusion weight map is dynamically generated using a normalized exponential function, adaptively allocating optimal weights in a data-driven manner. This maximizes the utilization of complementary information between features while preserving high-information-rich regions of each primitive image. The synergistic effect of the technical chain enables the fused image to simultaneously possess high information entropy (richness), strong edge fidelity (accuracy), and task-related feature enhancement characteristics, thereby meeting the stringent image quality requirements of scenarios such as high-precision target recognition and complex environment detection.

[0014] Optionally, the step of decomposing the infrared polarization image data into a first feature primitive image, a second feature primitive image, and a third feature primitive image based on the polarization bidirectional reflectance distribution function includes:

[0015] A physical decomposition model is established using the polarization bidirectional reflection distribution function;

[0016] The total radiation intensity information is extracted from the infrared polarization image data according to the physical decomposition model, and the first feature original primitive image is constructed according to the total radiation intensity information;

[0017] Based on the physical decomposition model, linear polarization parameters are extracted from the linear polarization state component information of the infrared polarization image data, and the second feature original primitive image is constructed based on the linear polarization parameters.

[0018] The polarization angle parameter is extracted from the linear polarization state component information according to the physical decomposition model, and the original primitive image of the third feature is constructed according to the polarization angle parameter.

[0019] By employing the aforementioned technical solution, a physical decomposition model based on the bidirectional polarization reflection distribution function can accurately extract three core polarization features from infrared polarization image data: total radiation intensity, degree of linear polarization, and polarization angle. Total radiation intensity reflects the target's thermal radiation characteristics, while the degree of linear polarization and polarization angle reflect physical properties such as the target's surface material and roughness. This decomposition method ensures that the three types of original primitive images retain the key information of infrared polarization imaging completely, avoiding feature confusion or loss.

[0020] Optionally, according to a preset task mode, feature optimization processing is performed on the first feature primitive image, the second feature primitive image, and the third feature primitive image respectively to obtain the first feature primitive image, the second feature primitive image, and the third feature primitive image:

[0021] According to the preset task mode, a mapping relationship between the original gray values ​​of pixels and the target gray values ​​in the first feature primitive image is established, and the gray values ​​of the target region are adjusted according to the mapping relationship to obtain the first feature primitive image;

[0022] The original primitive image of the second feature is decomposed into multiple feature layers with different spatial resolutions according to a preset scale parameter. Edge features and texture features are extracted from each feature layer. The edge features and texture features are superimposed and combined according to a preset weight ratio to obtain the second feature primitive image.

[0023] The noise pixels in the original primitive image of the third feature are identified by a preset noise judgment condition. All non-noise pixels within a preset range around the noise pixel are selected, the statistical values ​​of all the non-noise pixels are calculated, and the pixel value of the noise pixel is replaced with the statistical values ​​to obtain the third feature primitive image.

[0024] By adopting the above technical solutions, the gray values ​​of the target area in the original primitive image of the first feature can be adjusted in a targeted manner to enhance the distinction between the target area and the background area; through multi-scale feature layer decomposition and weighted combination, the edge and texture features at different scales in the original primitive image of the second feature can be fully preserved and strengthened to improve the clarity of the target contour; by replacing the noise pixels based on the statistical values ​​of non-noise pixels within a preset range, the noise in the original primitive image of the third feature can be effectively removed while retaining the effective feature information to the greatest extent.

[0025] Optionally, calculating the first region information abundance measure of the first feature primitive image, the second region information abundance measure of the second feature primitive image, and the third region information abundance measure of the third feature primitive image respectively includes:

[0026] The first feature primitive image, the second feature primitive image, and the third feature primitive image are divided into multiple image blocks according to a preset size, resulting in a first number of first image blocks of the first feature primitive image, a second number of second image blocks of the second feature primitive image, and a third number of third image blocks of the third feature primitive image.

[0027] The gray-level entropy value and the average gradient magnitude of the first image block are weighted and summed to obtain a first weighted sum. The first number of first weighted sums are arranged according to the spatial coordinate position of the first image block in the first feature primitive image to obtain the first region information abundance measure. The gray-level entropy value represents the complexity of the gray-level distribution within the image block, and the average gradient magnitude represents the intensity of the edge features within the image block.

[0028] The gray-level entropy value and the average gradient magnitude of the second image block are weighted and summed to obtain a second weighted sum. The second number of second weighted sums are arranged according to the spatial coordinate position of the second image block in the second feature primitive image to obtain the second region information abundance measure.

[0029] The gray-level entropy value and the average gradient magnitude of the third target image block are weighted and summed to obtain a third weighted sum. The third number of third weighted sums are arranged according to the spatial coordinate position of the third image block in the third feature primitive image to obtain the third region information abundance measure.

[0030] By employing the aforementioned technical solution, the feature primitive image is divided into image blocks for regional analysis. A weighted calculation combining gray-level entropy and the mean gradient magnitude is used to quantify the information value of local regions within each feature primitive image. This regionalization measurement method preserves spatial correlation while comprehensively characterizing the information abundance of different regions through the fusion of dual indicators. Gray-level entropy highlights the value of information-dense areas, while the mean gradient magnitude emphasizes the importance of regions with significant edge features. The resulting regional information abundance metric accurately reflects the information quality of each local region.

[0031] Optionally, calculating the first region complementarity measure between the first feature primitive image and the second feature primitive image, the second region complementarity measure between the first feature primitive image and the third feature primitive image, and the third region complementarity measure between the second feature primitive image and the third feature primitive image includes:

[0032] Calculate the first gray-level difference entropy between the first target image block of the first feature primitive image and the second target image block of the second feature primitive image, and use the first gray-level difference entropy as the first region complementarity measure. The first target image block is any one of the first number of first image blocks, and the second target image block is the image block in the second number of second image blocks that is in the same spatial coordinate as the first target image block.

[0033] Calculate the second gray-level difference entropy between the first target image block and the third target image block of the third feature primitive image, and use the second gray-level difference entropy as the second region complementarity measure. The third target image block is the image block among the third number of third image blocks that is in the same spatial coordinate range as the first target image block.

[0034] Calculate the third grayscale difference entropy between the second target image block and the third target image block, and use the third grayscale difference entropy as the complementarity measure of the third region.

[0035] By employing the aforementioned technical solution, gray-level difference entropy is calculated on a per-image-block basis with identical spatial coordinates, accurately quantifying the information differences and complementarity of different feature primitive images within the same region. Gray-level difference entropy, through statistical distribution probability of gray-level differences, captures both the numerical differences between features and reflects the complexity of the difference distribution, effectively characterizing the degree of information complementarity of different feature primitives within the same region. A higher difference entropy indicates stronger feature complementarity in that region. This regionalized complementarity measurement method ensures that the fusion process can specifically integrate the complementary advantages of different features in each region, avoiding information redundancy or conflict.

[0036] Optionally, calculating the first gray-level difference entropy between the first target image patch of the first feature primitive image and the second target image patch of the second feature primitive image, and using the first gray-level difference entropy as the complementarity measure of the first region, includes:

[0037] Calculate the difference in grayscale values ​​between each pair of pixels in the first target image block and the second target image block and take the absolute value to obtain a grayscale difference data group. Each pair of pixels includes a first pixel in the first target image block and a second pixel in the second target image block. The position of the first pixel in the first target image block is the same as the position of the second pixel in the second target image block.

[0038] The frequency of each difference value in the grayscale difference data group is counted, and the frequency of each difference value is divided by the total number of pixels in the first image block to obtain the probability of each difference value.

[0039] The first grayscale difference entropy is obtained by multiplying the probability of each difference value by the logarithm of the probability, summing all the intermediate values ​​and taking the negative value.

[0040] The first grayscale difference entropy is determined as the first region complementarity measure between the first target image block and the second target image block.

[0041] By employing the above technical solution, the grayscale difference between two target image blocks is calculated at the pixel level. Absolute value processing preserves the magnitude of the difference, and frequency statistics and probability transformation convert the discrete grayscale difference into quantifiable distribution features. Entropy calculation not only reflects the magnitude of the difference but also accurately characterizes the complexity of the difference distribution. The more diverse and uniform the differences, the higher the entropy value, intuitively reflecting the degree of information complementarity between the two feature primitives in that region.

[0042] Optionally, the step of processing the information abundance measures of the first region, the second region, the third region, the first region complementarity measure, the second region complementarity measure, and the third region complementarity measure using a normalized exponential function to generate a fusion weight map further includes:

[0043] The first region information abundance measure, the first region complementarity measure, and the second region complementarity measure are weighted and fused to obtain the first comprehensive feature weight.

[0044] The second region information abundance measure, the first region complementarity measure, and the third region complementarity measure are weighted and fused to obtain the second comprehensive feature weight.

[0045] The third region information abundance measure, the second region complementarity measure, and the third region complementarity measure are weighted and fused to obtain the third comprehensive feature weight;

[0046] Input the first comprehensive feature weight, the second comprehensive feature weight and the third comprehensive feature weight into the normalized exponential function to calculate the first weight value matrix corresponding to the first comprehensive feature weight, the second weight value matrix corresponding to the second comprehensive feature weight and the third weight value matrix corresponding to the third comprehensive feature weight.

[0047] A fused weight map is generated based on the first weight value matrix, the second weight value matrix, and the third weight value matrix. The fused weight map has the same spatial size as the first feature primitive image, the second feature primitive image, and the third feature primitive image.

[0048] By adopting the above technical solution, the regional information abundance and complementarity measurement are first weighted and fused in a targeted manner to form a comprehensive weight that can comprehensively reflect the information value of each feature primitive and the cross-feature complementarity relationship. Then, it is transformed into a spatial matching weight matrix through a normalized exponential function. This processing integrates the information quality of single features and the complementary advantages of multiple features, and preserves the spatial location correlation through matrix form. The final fused weight map can be accurately aligned with the feature primitive image.

[0049] A second aspect of this application provides an image fusion system based on infrared polarization imaging, specifically comprising:

[0050] The image data acquisition module is used to acquire infrared polarization image data of the target scene;

[0051] The image decomposition module is used to decompose the infrared polarized image data into a first feature original primitive image, a second feature original primitive image, and a third feature original primitive image based on the polarization bidirectional reflectance distribution function.

[0052] The task processing module is used to perform feature optimization processing on the first feature original primitive image, the second feature original primitive image and the third feature original primitive image respectively according to a preset task mode, so as to obtain the first feature primitive image, the second feature primitive image and the third feature primitive image.

[0053] The information abundance measurement calculation module is used to calculate the information abundance measurement of the first region of the first feature primitive image, the information abundance measurement of the second region of the second feature primitive image, and the information abundance measurement of the third region of the third feature primitive image, respectively.

[0054] The complementarity measurement calculation module is used to calculate the first region complementarity measurement between the first feature primitive image and the second feature primitive image, the second region complementarity measurement between the first feature primitive image and the third feature primitive image, and the third region complementarity measurement between the second feature primitive image and the third feature primitive image.

[0055] The fusion weight graph generation module is used to process the information abundance measure of the first region, the information abundance measure of the second region, the information abundance measure of the third region, the complementarity measure of the first region, the complementarity measure of the second region, and the complementarity measure of the third region through a normalized exponential function to generate a fusion weight graph.

[0056] The fused image output module is used to perform weighted fusion of the first feature primitive image, the second feature primitive image, and the third feature primitive image using the fusion weight map, and output the fused image.

[0057] A third aspect of this application provides an electronic device including a processor, a memory, a user interface, and a network interface, wherein the memory is used to store instructions, the user interface and the network interface are both used to communicate with other devices, and the processor is used to execute the instructions stored in the memory to cause the electronic device to perform the method as described in any of the foregoing.

[0058] A fourth aspect of this application provides a computer-readable storage medium storing instructions that, when executed, perform the method described in any of the preceding descriptions.

[0059] In summary, one or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages:

[0060] Starting from the physical essence, infrared polarization data is decomposed into three core feature primitives—total radiation intensity, linear polarization degree, and polarization angle—based on the polarization bidirectional reflectance distribution function. This fully preserves multi-dimensional information such as the target's thermal radiation characteristics and surface physical properties, laying a high-quality foundation for subsequent processing. Simultaneously, feature primitives are differentially enhanced according to preset task modes through operations such as contrast enhancement, edge strengthening, and noise suppression. By quantifying regional information abundance and complementarity, combined with grayscale difference entropy, a fine-grained measurement from pixel to region is achieved, providing an objective basis for weight allocation and avoiding subjective bias. Finally, a fusion weight map is dynamically generated based on these quantization results, balancing the information quality of single features with the complementary advantages of multiple features. This ensures that the fused image possesses both high information richness and accurate information integration, effectively meeting the stringent image quality requirements of scenarios such as high-precision target recognition and complex environment detection. Attached Figure Description

[0061] Figure 1 This is a schematic diagram of the system architecture of an embodiment of an image fusion method or an image fusion system based on infrared polarization imaging that applies this application;

[0062] Figure 2 This is a schematic flowchart of an image fusion method based on infrared polarization imaging disclosed in an embodiment of this application;

[0063] Figure 3 This is a schematic diagram of a module of an image fusion system based on infrared polarization imaging disclosed in an embodiment of this application;

[0064] Figure 4 This is a schematic diagram of the structure of an electronic device disclosed in an embodiment of this application.

[0065] Explanation of reference numerals in the attached figures: 301, Image data acquisition module; 302, Image decomposition module; 303, Task processing module; 304, Information abundance measurement calculation module; 305, Complementarity measurement calculation module; 306, Fusion weight map generation module; 307, Fusion image output module; 401, Processor; 402, Communication bus; 403, User interface; 404, Network interface; 405, Memory. Detailed Implementation

[0066] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.

[0067] like Figure 1 As shown, system architecture 100 may include terminal devices 101, 102, and 103, a network 104, and a server 105. Network 104 serves as the medium for providing communication links between terminal devices 101, 102, and 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.

[0068] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as model training applications, video recognition applications, web browser applications, social platform software, etc.

[0069] Terminal devices 101, 102, and 103 can be either hardware or software. When terminal devices 101, 102, and 103 are hardware, they can be various electronic devices with displays, including but not limited to smartphones, tablets, e-book readers, MP3 (Moving Picture Experts Group Audio Layer III) players, MP4 (Moving Picture Experts Group Audio Layer IV) players, laptops, and desktop computers, etc. When terminal devices 101, 102, and 103 are software, they can be installed in the aforementioned electronic devices. They can be implemented as multiple software programs or software modules (e.g., multiple software programs or software modules used to provide distributed services) or as a single software program or software module. No specific limitations are imposed here.

[0070] This embodiment discloses an image fusion method based on infrared polarization imaging. Figure 2 This is a flowchart illustrating an image fusion method based on infrared polarization imaging disclosed in an embodiment of this application, as shown below. Figure 2 As shown, the method includes the following steps:

[0071] S201. Acquire infrared polarization image data of the target scene;

[0072] Specifically, an infrared polarization imaging camera equipped with a 12μm infrared detector was used. The camera was fixed on a tripod and leveled to ensure that the lens optical axis was aligned with the target scene (specifically including man-made buildings, natural vegetation, and moving or stationary vehicles with a height of 1-5m, in a clear nighttime environment with ambient light mainly consisting of moonlight and ambient diffused light). The spectral response range was set to 8-14μm, the frame rate to 15fps, and the exposure time to 50ms. The camera acquisition program was started, and the polarizer was switched sequentially to 0°, 45°, 90°, and 135° (0° is parallel to the ground and along the horizontal central axis of the camera lens, i.e., parallel to the long side of the camera body; 90° is perpendicular to the 0° direction, i.e., perpendicular to the ground and along the vertical central axis of the camera lens)). Five frames of images were continuously acquired in each direction to reduce random noise. During the acquisition process, the camera and the target scene remained relatively stationary. The acquired image data was stored in 16-bit .raw format on the camera's built-in memory card, and metadata such as acquisition time, ambient temperature, and target distance were recorded. After acquisition, the data was transmitted via USB (Universal). The Serial Bus (Universal Serial Bus) interface transmits the raw data to the server, resulting in an infrared polarization image dataset containing four polarization states.

[0073] S202. Based on the polarization bidirectional reflectance distribution function, the infrared polarization image data is decomposed into a first feature original primitive image, a second feature original primitive image, and a third feature original primitive image;

[0074] Specifically, infrared polarization image data (640×512 pixels, 16-bit .raw format) in four directions (0°, 45°, 90°, 135°) are loaded. A physical decomposition model is established based on the polarization bidirectional reflectance distribution function (combining the diffuse and specular reflection characteristics of the target surface). The decomposition yields three feature primitive images: The first feature primitive image is formed by calculating the average grayscale value of the corresponding pixels in the four directions and arranging them according to the original coordinates; The second feature primitive image first obtains two linear polarization components from the grayscale difference between the images in the 0° and 90° and 45° and 135° directions, calculates the sum of squares of the two linear polarization components, and then divides the square root of the sum of squares by the total radiation intensity (set to 0 where the total radiation is 0) to obtain the degree of polarization value. The degree of polarization value is then arranged according to the coordinates; The third feature primitive image is formed by calculating the polarization angle based on the above two linear polarization components through arctangent (correcting the angle to 0°-180° according to the positive and negative values ​​of the components) and arranging them according to the coordinates.

[0075] Optionally, the step of decomposing the infrared polarization image data into a first feature primitive image, a second feature primitive image, and a third feature primitive image based on the polarization bidirectional reflectance distribution function includes:

[0076] A physical decomposition model is established using the polarization bidirectional reflectance distribution function; total radiation intensity information is extracted from the infrared polarized image data based on the physical decomposition model, and the first feature primitive image is constructed based on the total radiation intensity information; linear polarization degree parameters are extracted from the linear polarization state component information of the infrared polarized image data based on the physical decomposition model, and the second feature primitive image is constructed based on the linear polarization degree parameters; polarization angle parameters are extracted from the linear polarization state component information based on the physical decomposition model, and the third feature primitive image is constructed based on the polarization angle parameters.

[0077] Specifically, based on the characteristics of diffuse reflection (weak polarization) and specular reflection (strong polarization) in the bidirectional polarization reflection distribution function, and their relationship with target material, roughness, and illumination angle, a four-layer model was constructed using the Python programming language, combined with the OpenCV (Open Computer Vision) library and the NumPy (NumPy Numerical Computation Extension Library). First, an input preprocessing layer was built. Code was written to read the raw format data of the four-directional infrared polarization images and convert it into a 640×512 two-dimensional array using the NumPy library. Then, the `medianBlur` function from the OpenCV library was called to perform median filtering on the image, removing salt-and-pepper noise introduced by sensor noise. The processed data was used as the model input. Next, a reflection component separation layer is constructed. Based on the polarization difference characteristics of diffuse reflection and specular reflection in the polarization bidirectional reflection distribution function, a separation algorithm is constructed: First, the preliminary polarization degree of each pixel is calculated (estimated by the gray values ​​of the four-directional image). A polarization degree threshold of 0.3 is set. Regions with polarization degrees below this threshold are identified as diffuse reflection-dominant regions, and their gray values ​​are extracted to form a diffuse reflection component matrix. Regions with polarization degrees higher than or equal to this threshold are identified as specular reflection-dominant regions, and their gray values ​​are extracted to form a specular reflection component matrix. Then, a feature parameter mapping layer is designed, and a conversion formula is written to map the two components to intermediate parameters: the basic value of total radiation intensity = diffuse reflection component × 0.6 + specular reflection component × 0.4 (weights are determined by calibration using 10 sets of standard sample images); the horizontal and vertical components (Q component) of the linear polarization state = gray value of the image in the 0° direction - gray value of the image in the 90° direction; the oblique component (U component) of the linear polarization state = gray value of the image in the 45° direction - gray value of the image in the 135° direction. These parameters are stored as a matrix of the same size as the input image. Finally, an output rule definition layer is constructed, and the calculation logic code for three types of feature parameters is written: the total radiation intensity is calculated by the average gray value of the corresponding pixels in the four-directional image; the degree of linear polarization is calculated by dividing the square root of the sum of the squares of the Q component and the squares of the U component by the total radiation intensity (the degree of linear polarization is set to 0 when the total radiation intensity is 0); the polarization angle is calculated by multiplying the arctangent of the ratio of the U component to the Q component by a fixed coefficient (e.g., 0.5), and the angle is corrected to the range of 0°-180° according to the positive and negative values ​​of the Q and U components. These calculation rules are encapsulated into callable functions to complete the model construction.

[0078] Furthermore, the input preprocessing layer and feature parameter mapping layer in the physical decomposition model are invoked to load the preprocessed four-directional (0°, 45°, 90°, 135°) infrared polarization image data (a two-dimensional array of 640×512 pixels, with pixel values ​​ranging from 0 to 1); for each spatial coordinate (x, y) in the image, x represents the position index of the pixel in the horizontal direction (horizontal direction) of the image, and y represents the position index of the pixel in the vertical direction (vertical direction) of the image. For example, in a 640×512 pixel image, the value of x ranges from 0 to 639 (corresponding to each column of pixels from left to right), and the value of y ranges from 0 to 511 (corresponding to each row of pixels from top to bottom). Following the total radiometric intensity calculation rules defined by the model feature parameter mapping layer, pixel values ​​at the coordinates in the 0°, 45°, 90°, and 135° directions are extracted. Outliers deviating more than 30% from the pixel values ​​in the other three directions are removed (if outliers exist, the average of the remaining three values ​​is used instead). Then, the arithmetic mean of the pixel values ​​in the four directions is calculated, and this arithmetic mean is used as the total radiometric intensity value at (x, y). After traversing all coordinates and completing the calculation, all total radiometric intensity values ​​are arranged in the spatial coordinate order of the original image, forming a 640×512 pixel two-dimensional matrix. The total radiometric intensity values ​​(range 0-1) in the matrix are converted to 16-bit integers (range 0-65535) through linear mapping. The OpenCV imwrite function is used to save this matrix as a label image file, thus obtaining the original primitive image of the first feature.

[0079] Furthermore, the feature parameter mapping layer and output rule definition layer of the physical decomposition model are invoked. Preprocessed four-directional infrared polarization image data is loaded via a Python script. For each pixel coordinate (x, y) (x is the horizontal index of 0-639, y is the vertical index of 0-511), the linear polarization state component information is first extracted. The model defines and calculates the Q component (pixel value in the 0° direction minus the pixel value in the 90° direction) and the U component (pixel value in the 45° direction minus the pixel value in the 135° direction). Then, the linear polarization degree parameter is calculated according to the output rules: using the sqr function of the numerical computation extension library (NumPy). The `t` function calculates the square root of the sum of the squares of the Q and U components (i.e., the linear polarization intensity), and then divides it by the total radiation intensity value of that coordinate (if the total radiation intensity is 0, the linear polarization degree is directly set to 0), obtaining a linear polarization degree value in the range of 0-1. After traversing all pixel coordinates to complete the calculation, all linear polarization degree values ​​are arranged in the original image spatial coordinate order to form a 640×512 pixel two-dimensional matrix. The matrix values ​​(0-1) are converted into 16-bit integers (0-65535) through linear mapping, and saved as a tag image file format using the `imwrite` function, thus obtaining the original primitive image of the second feature.

[0080] Furthermore, the feature parameter mapping layer and output rule definition layer of the physical decomposition model are invoked. The extracted linear polarization state components (Q and U components) and total radiance map data are loaded via a Python script. For each pixel coordinate (x, y) (x is the horizontal index of 0-639, y is the vertical index of 0-511), it is first determined whether the total radiance at that coordinate is 0 (if it is 0, the polarization angle is directly set to 0°). If the total radiance is not 0, the arctangent of the ratio of the U component to the Q component is calculated using NumPy's arctan2 function (the return value ranges from -π to π radians) based on the Q and U components. The result is then multiplied by 0.5 to convert it to a polarization angle in radians, and subsequently converted using the angle conversion formula (radians × 180 / π). The angle is converted to an angle value; quadrant correction is performed based on the signs of the Q and U components: when Q>0 and U>0, the angle remains unchanged; when Q<0, the angle is increased by 90°; when U<0, the angle is increased by 180°, ensuring that the final angle range is limited to 0°-180°; after traversing all pixel coordinates to complete the calculation, all polarization angle values ​​are arranged in the original image space coordinate order to form a 640×512 pixel two-dimensional matrix; the angle values ​​from 0° to 180° are converted to 16-bit integers (0-65535) through linear mapping, and saved as a label image file (TIFF) using OpenCV's imwrite function, thus obtaining the original primitive image of the third feature (polarization angle map), and generating a metadata file that records the angle mapping relationship.

[0081] S203. According to the preset task mode, perform feature optimization processing on the first feature original primitive image, the second feature original primitive image and the third feature original primitive image respectively to obtain the first feature primitive image, the second feature primitive image and the third feature primitive image.

[0082] Specifically, when the preset task mode is target detection mode, for the first feature primitive image, its grayscale value range and the frequency of each grayscale value are first statistically analyzed. Grayscale value intervals for the target region and background region are set, and a linear mapping relationship between the original grayscale values ​​and the corresponding intervals is established. The grayscale value of each pixel is adjusted according to this relationship to enhance the distinction between the target and the background, thus obtaining the first feature primitive image. For the second feature primitive image, it is decomposed into multiple feature layers with different spatial resolutions according to preset scale parameters. For each layer, edge features are extracted by checking whether the grayscale difference between a pixel and its 8 neighboring pixels exceeds a preset threshold. This is achieved by calculating a 3×3... Texture features are extracted from the mean and variance of the neighborhood grayscale. Then, edge and texture features are superimposed and combined at the corresponding pixel positions according to a preset weight ratio (e.g., 40% for edges and textures at the first scale, and 30% for the second and third scales) to obtain the second feature primitive image. For the original primitive image of the third feature, noise pixels are identified by a preset noise judgment condition (e.g., the absolute value of the difference between the pixel grayscale value and the mean of the 3×3 neighborhood exceeds a preset threshold of 50). All non-noise pixels within a 5×5 range around the noise pixel are selected, and their mean grayscale values ​​are calculated as statistical values. The pixel values ​​of the noise pixels are replaced with these statistical values ​​to obtain the third feature primitive image.

[0083] Optionally, according to a preset task mode, feature optimization processing is performed on the first feature primitive image, the second feature primitive image, and the third feature primitive image respectively to obtain the first feature primitive image, the second feature primitive image, and the third feature primitive image:

[0084] When the preset task mode is target detection mode, a mapping relationship is established between the original grayscale values ​​of pixels in the first feature primitive image and the grayscale values ​​of the target. The grayscale values ​​of the target region are adjusted according to the mapping relationship to obtain the first feature primitive image. The second feature primitive image is decomposed into multiple feature layers with different spatial resolutions according to preset scale parameters. Edge features and texture features are extracted from each feature layer. The edge features and texture features are superimposed and combined according to preset weight ratios to obtain the second feature primitive image. Noise pixels in the third feature primitive image are identified through preset noise judgment conditions. All non-noise pixels within a preset range around the noise pixel are selected. The statistical values ​​of all non-noise pixels are calculated, and the pixel values ​​of the noise pixel are replaced with the statistical values ​​to obtain the third feature primitive image.

[0085] Specifically, when the preset task mode is the target detection mode, firstly, by traversing all pixels of the first feature primitive image, the grayscale value range and the frequency of each grayscale value are statistically obtained. Based on the statistical results, a target grayscale value interval is set, wherein the expected grayscale value interval of the target region and the expected grayscale value interval of the background region do not overlap and have a preset difference. Then, a mapping relationship between the original grayscale value and the target grayscale value is established. Specifically, the original grayscale value is linearly mapped to the corresponding target grayscale value interval proportionally. For pixels that have been preliminarily determined to be target regions by previous image analysis, their original grayscale value is mapped to the expected grayscale value interval of the target region, and the remaining pixels are mapped to the expected grayscale value interval of the background region. Finally, the grayscale value of each pixel in the image is converted according to the mapping relationship to complete the adjustment of the grayscale value of the target region and obtain the first feature primitive image.

[0086] Furthermore, the original primitive image of the second feature is decomposed into multiple feature layers with different spatial resolutions according to preset scale parameters (e.g., setting 3 scale levels, which are 1 times, 1 / 2, and 1 / 4 of the original image resolution, respectively). For each feature layer, edge features are extracted by calculating the gray-level difference between a pixel and its neighboring pixels (specifically, when the gray-level difference between a pixel and its 8 surrounding neighboring pixels exceeds a preset threshold, the pixel is determined to be an edge point and marked). Texture features are extracted by statistically analyzing the distribution pattern of gray-level values ​​in the pixel's neighborhood (specifically, the mean and variance of gray-level values ​​in the 3×3 neighborhood of each pixel are calculated as texture feature values). According to preset weight ratios (e.g., edge features and texture features of the first scale layer each account for 40% weight, edge features and texture features of the second scale layer each account for 30% weight, and edge features and texture features of the third scale layer each account for 30% weight), the edge features and texture features extracted from each scale layer are numerically superimposed at the corresponding pixel positions to obtain the second feature primitive image.

[0087] Furthermore, noise pixels in the original primitive image of the third feature are identified by a preset noise judgment condition (the preset noise judgment condition is: for each pixel in the image, calculate the gray value of it and all pixels in the 3×3 neighborhood; if the absolute value of the difference between the gray value of the pixel and the mean exceeds a preset threshold of 50, it is a noise pixel); select all non-noise pixels within a 5×5 range around the noise pixel, calculate the gray value of these non-noise pixels as a statistical value; replace the original pixel value of the noise pixel with the statistical value, and process all identified noise pixels in the image in sequence to obtain the third feature primitive image.

[0088] S204. Calculate the first region information abundance measure of the first feature primitive image, the second region information abundance measure of the second feature primitive image, and the third region information abundance measure of the third feature primitive image, respectively.

[0089] Specifically, the first, second, and third feature primitive images are called. Each image is first divided into non-overlapping image blocks of a preset size of 16×16 pixels. For each first image block, the `calcHist` function is called to calculate the gray-level histogram, and the probability of different gray levels is calculated. The gray-level entropy value, reflecting the complexity of the gray-level distribution, is obtained through probability calculation. Simultaneously, the Sobel operator is used to calculate the gradient of the image block in the horizontal and vertical directions. The gradient values ​​in the two directions are combined to calculate the gradient intensity of each pixel. Finally, the average gradient intensity of all pixels is calculated to obtain the gradient reflecting the intensity of edge features. The mean gradient magnitude is calculated; then, it is weighted and summed in a ratio of 0.6 (grayscale entropy value): 0.4 (gradient magnitude mean) to obtain the first weighted sum. All first weighted sums are arranged according to the spatial coordinates of the original image blocks to form a first region information abundance measurement matrix with the same size as the original image. Using the same process, the weighted sum of grayscale entropy value and gradient magnitude mean is calculated for the second image block with a weight of 0.5:0.5 to obtain the second weighted sum, which is then arranged to form the second region information abundance measurement matrix. The weighted sum of the two is calculated for the third image block with a weight of 0.4:0.6 to obtain the third weighted sum, which is then arranged to form the third region information abundance measurement matrix.

[0090] Optionally, calculating the first region information abundance measure of the first feature primitive image, the second region information abundance measure of the second feature primitive image, and the third region information abundance measure of the third feature primitive image respectively includes:

[0091] The first feature primitive image, the second feature primitive image, and the third feature primitive image are divided into multiple image blocks according to a preset size, resulting in a first number of first image blocks of the first feature primitive image, a second number of second image blocks of the second feature primitive image, and a third number of third image blocks of the third feature primitive image. A first weighted sum is obtained by weighting and summing the gray-level entropy value and the average gradient magnitude of the first image blocks. The first number of first weighted sums are then arranged according to the spatial coordinates of the first image blocks in the first feature primitive image to obtain the first region information abundance measure. The gray-level entropy value characterizes the information abundance within the image block. The complexity of the grayscale distribution, wherein the mean gradient magnitude characterizes the intensity of edge features within an image patch; a second weighted sum is obtained by weighting and summing the grayscale entropy value and the mean gradient magnitude of the second image patch, and the second number of second weighted sums are arranged according to the spatial coordinate position of the second image patch in the second feature primitive image to obtain the second region information abundance measure; a third weighted sum is obtained by weighting and summing the grayscale entropy value and the mean gradient magnitude of the third target image patch, and the third number of third weighted sums are arranged according to the spatial coordinate position of the third image patch in the third feature primitive image to obtain the third region information abundance measure.

[0092] Specifically, the first, second, and third feature primitive images are called, and the preset image block size is set to 16×16 pixels. For the first feature primitive image, starting from the top left pixel (0,0), a 16-pixel wide column segment is taken every 16 pixels horizontally, and a 16-pixel high row segment is taken every 16 pixels vertically. The intersection of the two forms a 16×16 pixel first image block. This is then slid to the right and down to capture, resulting in 40 horizontal (640÷16) and 32 vertical (512÷16) first image blocks, for a total of 1280 first image blocks. If there are areas with less than 16×16 pixels at the image edge (such as non-integer segmentation after resizing), 0 values ​​are used to fill the area to 16×16 pixels to ensure that all first image blocks are the same size. The same method is used to divide the second feature primitive image to obtain 1280 second image blocks, and the third feature primitive image to obtain 1280 third image blocks.

[0093] Furthermore, for each first image block, the OpenCV `calcHist` function is called to calculate the grayscale histogram. The frequency of each grayscale level is counted and divided by 256 (total number of pixels) to obtain the probability. The grayscale entropy value (characterizing the complexity of grayscale distribution) is obtained by accumulating the negative value of the product of each probability and its logarithm. Simultaneously, the Sobel operator is used to calculate the horizontal and vertical gradients. The gradient values ​​of each pixel in both directions are converted into gradient strength by square root of the sum of squares, and then the average value is calculated to obtain the average gradient magnitude (characterizing edge strength). The first weighted sum is obtained by weighting the first image block according to the ratio of 60% of the grayscale entropy value and 40% of the average gradient magnitude. The 1280 first weighted sums are then calculated according to the spatial coordinate position of each first image block in the first feature primitive image (e.g., the image block in row x and column y corresponds to the original image (x×16, y×...). The regions (x+1)×16-1, (y+1)×16-1) are arranged and reconstructed into a 640×512 pixel two-dimensional matrix, which is the first region information abundance measure. For each second image block, the gray-level entropy value and the mean gradient magnitude are calculated using the same method. The gray-level entropy value and the mean gradient magnitude are weighted and summed at a ratio of 50% for the gray-level entropy value and 50% for the mean gradient magnitude to obtain the second weighted sum. The 1280 second weighted sums are arranged according to their corresponding spatial coordinate positions and reconstructed into a matrix of the same size, which is the second region information abundance measure. For each third image block, the gray-level entropy value and the mean gradient magnitude are calculated using the same method. The gray-level entropy value and the mean gradient magnitude are weighted and summed at a ratio of 40% for the gray-level entropy value and 60% for the mean gradient magnitude to obtain the third weighted sum. The 1280 third weighted sums are arranged according to their corresponding spatial coordinate positions and reconstructed into a matrix of the same size, which is the third region information abundance measure.

[0094] S205. Calculate the first region complementarity measure between the first feature primitive image and the second feature primitive image, the second region complementarity measure between the first feature primitive image and the third feature primitive image, and the third region complementarity measure between the second feature primitive image and the third feature primitive image.

[0095] Specifically, load the information abundance metric matrices for the first, second, and third regions. For each pixel coordinate (x, y) (x is the horizontal index from 0 to 639, and y is the vertical index from 0 to 511), calculate the complementarity metric for the first region: first, take the value of the information abundance metric for the first region at (x, y) and the value of the information abundance metric for the second region at (x, y), calculate the absolute difference between the two, and then divide by the maximum value of the two (if the maximum value is 0, the difference is directly set to 0) to obtain the complementarity value for the first region at that location; use the same method to calculate the information abundance of the first region. The absolute difference between the information abundance measure of the first region and the information abundance measure of the third region at (x, y) is divided by the maximum value of the two (set to 0 when the maximum value is 0) to obtain the complementarity value of the second region; the absolute difference between the information abundance measure of the second region and the information abundance measure of the third region at (x, y) is calculated and divided by the maximum value of the two (set to 0 when the maximum value is 0) to obtain the complementarity value of the third region; after traversing all pixel coordinates, the complementarity values ​​of the first region are arranged according to the original coordinates to form the first region complementarity measure matrix with the same size as the original image; similarly, the second and third region complementarity measure matrices are formed respectively.

[0096] Optionally, calculating the first region complementarity measure between the first feature primitive image and the second feature primitive image, the second region complementarity measure between the first feature primitive image and the third feature primitive image, and the third region complementarity measure between the second feature primitive image and the third feature primitive image includes:

[0097] Calculate the first grayscale difference entropy between the first target image block of the first feature primitive image and the second target image block of the second feature primitive image, and use the first grayscale difference entropy as the first region complementarity measure. The first target image block is any one of the first number of first image blocks, and the second target image block is an image block in the second number of second image blocks that is in the same spatial coordinates as the first target image block. Calculate the second grayscale difference entropy between the first target image block and the third target image block of the third feature primitive image, and use the second grayscale difference entropy as the second region complementarity measure. The third target image block is an image block in the third number of third image blocks that is in the same spatial coordinate range as the first target image block. Calculate the third grayscale difference entropy between the second target image block and the third target image block, and use the third grayscale difference entropy as the third region complementarity measure.

[0098] Specifically, when calculating the complementarity measure of the first region, firstly, select any first target image patch from the first feature primitive image and locate its spatial coordinate range in the image (e.g., from the upper left corner coordinate (x1, y1) to the lower right corner coordinate (x2, y2)). Then, select a second target image patch from the second feature primitive image that completely overlaps with this range. Traverse all pixel pairs with the same position in the two image patches (i.e., the pixel at coordinate (i, j) in the first target image patch and the pixel at coordinate (i, j) in the second target image patch), calculate the difference in grayscale values ​​of each pair of pixels and take the absolute value to form a grayscale difference data group. Count the frequency of each difference value in this data group, divide the frequency by the total number of pixels in the image patch (256) to obtain the probability of each difference value, and then use the formula: H diff p represents the gray-level difference entropy, k represents the index of the gray-level difference, and its value ranges from 1 to n (n is the total number of all possible gray-level differences, determined by the gray-level range of the image, such as the gray-level difference range of a 16-bit image being 0-65535), and p k The probability of the k-th gray level difference in the image block is represented by the first gray level difference entropy, which is used as the first region complementarity measure of the corresponding spatial regions of the first target image block and the second target image block.

[0099] Furthermore, the same method is used to calculate the second and third region complementarity measures as described above: For the second region complementarity measure, a third target image block with the same spatial coordinate range as the first target image block is selected. The absolute value of the gray-level difference between the corresponding pixels of the two is calculated and a data set is formed. After frequency statistics and probability calculation, the second gray-level difference entropy is obtained using the same entropy formula, which serves as the second region complementarity measure value. For the third region complementarity measure, the absolute value of the gray-level difference between the corresponding pixels of the second and third target image blocks is calculated. The frequency, probability, and entropy values ​​are calculated using the same process to obtain the third gray-level difference entropy, which serves as the third region complementarity measure value. After traversing all image blocks, the measure values ​​are arranged according to their original spatial coordinate positions to form a matrix of first, second, and third region complementarity measures that is consistent with the size of the feature primitive image.

[0100] Optionally, calculating the first gray-level difference entropy between the first target image patch of the first feature primitive image and the second target image patch of the second feature primitive image, and using the first gray-level difference entropy as the complementarity measure of the first region, includes:

[0101] The difference in grayscale values ​​between each pair of pixels in the first target image block and the second target image block is calculated and its absolute value is taken to obtain a grayscale difference data group. Each pair of pixels includes a first pixel in the first target image block and a second pixel in the second target image block. The position of the first pixel in the first target image block is the same as the position of the second pixel in the second target image block. The frequency of each difference value in the grayscale difference data group is counted, and the frequency of each difference value is divided by the total number of pixels in the first image block to obtain the probability of each difference value. The probability of each difference value is multiplied by the logarithm of the probability to obtain an intermediate value. All intermediate values ​​are then added together and the result is negative to obtain the first grayscale difference entropy. The first grayscale difference entropy is determined as the first region complementarity measure between the first target image block and the second target image block.

[0102] Specifically, the `imread` function of OpenCV is called to load the image files of the first and second target image blocks respectively, and the NumPy library is used to convert them into two-dimensional arrays (data type uint16, pixel value range 0-65535). The array of the first target image block is denoted as `arr1`, and the array of the second target image block is denoted as `arr2`. Both have dimensions of (16, 16), and the index (i, j) in the array corresponds to the pixel in the i-th row and j-th column of the image block (i∈[0, 15], j∈[0, 15]). Next, an empty list `diff_list` is defined to store the grayscale difference data, and a double for loop is used to traverse all pixel coordinates: the outer loop variable `i` ranges from 0 to 15 (traversing rows), and the inner loop variable `j`... Starting from 0 to 15 (traversing columns), in each loop, arr1[i,j] (the grayscale value of the pixel in the i-th row and j-th column of the first target image block) and arr2[i,j] (the grayscale value of the pixel at the same position in the second target image block) are obtained by index. The difference between the two is calculated and the absolute value is taken, i.e., diff=abs(arr1[i,j]-arr2[i,j]). The diff values ​​obtained in each calculation are added to diff_list in sequence. After the loop ends, diff_list will contain 256 elements (16×16=256), and each element is an integer in the range of 0-65535. Finally, diff_list is converted into a one-dimensional array using the array function of the NumPy library. This array is the grayscale difference data group.

[0103] Furthermore, the generated grayscale difference data set is obtained, and the bincount function of the NumPy library is called to count the frequency of each possible difference value in the array (the function parameter minlength is set to 65536), resulting in a frequency array freq of length 65536, where freq[d] represents the number of times the difference value d appears in the data set; the total number of pixels in the first image block is determined to be 256 (16×16), and each element in the frequency array freq is divided by 256 using NumPy's division operation to obtain a probability array prob (data type converted to float32), where prob[d] = freq[d] / 256, representing the probability of the difference value d appearing; all elements in prob that are greater than 0 and their corresponding difference values ​​are filtered out by conditional judgment, and the results are stored as a tuple containing two lists (differences, probability), where differences records all the difference values ​​that have appeared, and probability records the probability of the corresponding difference value; finally, it is verified whether the sum of all elements in probability is close to 1.0 to ensure the correctness of the probability statistics.

[0104] Furthermore, the statistically analyzed difference values ​​and their corresponding probabilities are obtained (stored as a data structure containing two arrays, unique_diffs and probs, where unique_diffs represents the observed difference values ​​and probs represents the corresponding probabilities, i.e., the probability p of each difference value occurring). The calculation method is as follows: for each non-zero probability p, the intermediate value mid = p × log2(p) is first calculated (where p is the probability of a certain difference value occurring); the accumulation variable sum_mid is initialized to 0.0, and a for loop is used to iterate through all non-zero probabilities p: if p > 0 (to avoid errors caused by calculating the logarithm of a 0 probability value), then mid is calculated and accumulated into sum_mid; after the loop ends, the first gray-level difference entropy H = -sum_mid, with a value range of 0 to 8; the value of H is used as the first region complementarity metric between the current first target image patch and the second target image patch, and stored in the matrix position corresponding to the spatial coordinates of the original image patch; it is verified whether H is within the range of 0 ≤ H ≤ 8; after iterating through all image patch pairs, the first region complementarity metric matrix is ​​obtained.

[0105] S206. The information abundance measure of the first region, the information abundance measure of the second region, the information abundance measure of the third region, the complementarity measure of the first region, the complementarity measure of the second region, and the complementarity measure of the third region are processed by a normalized exponential function to generate a fusion weight map.

[0106] Specifically, the information abundance metric matrix and the complementarity metric matrix of the first, second, and third regions are loaded. First, the comprehensive score of each of the three feature primitive images is calculated. The comprehensive score of each feature is its information abundance value multiplied by the sum of its complementarity values ​​with the other two features. Then, the three comprehensive scores are adjusted by subtracting the maximum value of the three scores from each score to obtain a more stable adjusted score. Next, the weight of each feature is calculated using a normalized exponential function. Each weight represents the fusion ratio of the corresponding feature at that pixel position. Finally, the calculated three weights are stored in a three-channel weight map according to the pixel position. Each channel corresponds to the weight of one feature, and the data is saved in 32-bit floating-point multi-channel TIFF format to obtain the fused weight map.

[0107] Optionally, the step of processing the information abundance measures of the first region, the second region, the third region, the first region complementarity measure, the second region complementarity measure, and the third region complementarity measure using a normalized exponential function to generate a fusion weight map further includes:

[0108] The first region information abundance measure, the first region complementarity measure, and the second region complementarity measure are weighted and fused to obtain a first comprehensive feature weight; the second region information abundance measure, the first region complementarity measure, and the third region complementarity measure are weighted and fused to obtain a second comprehensive feature weight; the third region information abundance measure, the second region complementarity measure, and the third region complementarity measure are weighted and fused to obtain a third comprehensive feature weight; the first comprehensive feature weight, the second comprehensive feature weight, and the third comprehensive feature weight are input into a normalized exponential function to calculate a first weight value matrix corresponding to the first comprehensive feature weight, a second weight value matrix corresponding to the second comprehensive feature weight, and a third weight value matrix corresponding to the third comprehensive feature weight; a fused weight map is generated based on the first weight value matrix, the second weight value matrix, and the third weight value matrix, and the fused weight map has the same spatial size as the first feature primitive image, the second feature primitive image, and the third feature primitive image.

[0109] Specifically, load the first region information abundance measurement matrix (A1), the second region information abundance measurement matrix (A2), the third region information abundance measurement matrix (A3), and the first region complementarity measurement matrix (C12), the second region complementarity measurement matrix (C13), and the third region complementarity measurement matrix (C23). Set three sets of fusion weights (e.g., all of which are 1 / 3). For each pixel coordinate (x, y) (x∈[0, 639], y∈[0, 511]), calculate the first comprehensive feature weight W1 = (A1[x, y] × w_a + C12[x, y] × w_c1 + C13[x, y] × w_c2), where w_a, w_c1, and w_c2 are preset weighting coefficients. Calculate the second comprehensive feature weight W2 = (A2[x,y]×w_a + C12[x,y]×w_c1 + C23[x,y]×w_c3), where w_a, w_c1, and w_c3 are preset weighting coefficients; calculate the third comprehensive feature weight W3 = (A3[x,y]×w_a + C13[x,y]×w_c2 + C23[x,y]×w_c3), where w_a, w_c2, and w_c3 are preset weighting coefficients; restrict W1, W2, and W3 to the range of 0-1 using the NumPy clip function; after traversing all pixels, store the three weights as matrices of the same size as the original image, thus obtaining the first, second, and third comprehensive feature weights.

[0110] Further, load the first comprehensive feature weight matrix (W1), the second comprehensive feature weight matrix (W2), and the third comprehensive feature weight matrix (W3); for each pixel coordinate (x, y) in the image (x∈[0,639], y∈[0,511]), perform the following operations: extract the three comprehensive feature weight values, denoted as w1=W1[x, y], w2=W2[x, y], w3=W3[x, y] respectively (where w1 represents the value of the first comprehensive feature weight in that pixel); first calculate the maximum value max_w among the three, and then adjust each weight value: w1'=w1-max_w, w2'=w2-max_w, w3'=w3-max_w; call the exp function to calculate the exponential value: exp1=exp(w1'), exp2=exp(w2'), exp3=exp(w3') (where exp represents the natural exponential function, used for amplification adjustment). The difference between the weights is calculated as follows: exp1 represents the result of natural exponentiation on the adjusted first comprehensive feature weight w1'; the exponents are calculated as sum_exp=exp1+exp2+exp3; finally, the normalized weight values ​​are calculated: the value of the first weight matrix at (x, y) is w1_norm=exp1 / sum_exp, the value of the second weight matrix at (x, y) is w2_norm=exp2 / sum_exp, and the value of the third weight matrix at (x, y) is w3_norm=exp3 / sum_exp (where w1_norm, w2_norm, and w3_norm are the normalized first, second, and third weight values, respectively; after traversing all pixels, the three normalized weight values ​​are stored as 32-bit floating-point matrices of the same size as the original image, thus obtaining the first weight matrix, the second weight matrix, and the third weight matrix.

[0111] Further, load the first, second, and third weight matrix (W1_norm, W2_norm, W3_norm); create a three-channel empty matrix fusion_weights with dimensions 640×512×3 (height×width×number of channels); for each pixel coordinate (x, y) (x∈[0,639], y∈[0,511]), assign the value W1_norm[x, y] of the first weight matrix at that position to the corresponding position of channel 0 of fusion_weights, i.e., fusion_weights[x, y, 0] = W1_norm[x, y], and then assign the value of the second weight matrix... The value W2_norm[x, y] is assigned to the first channel, i.e., fusion_weights[x, y, 1] = W2_norm[x, y]. The value W3_norm[x, y] of the third weight matrix is ​​assigned to the second channel, i.e., fusion_weights[x, y, 2] = W3_norm[x, y]. The three-channel values ​​at each pixel position are summed using the NumPy sum function. After traversing all pixels and completing the assignment, the three-channel matrix fusion_weights is saved as a 32-bit floating-point multi-channel TIFF format, i.e., the fused weight map, using the OpenCV imwrite function.

[0112] S207. The first feature primitive image, the second feature primitive image, and the third feature primitive image are weighted and fused using the fusion weight map to output the fused image.

[0113] Specifically, load the fusion weight map and the first, second, and third feature primitive images, ensuring that all image spatial dimensions are perfectly matched; create an empty matrix `fused_image` with the same size as the feature primitive images to store the fusion result; for each pixel coordinate (x, y) (x∈[0, 639], y∈[0, 511]), extract the three-channel weight values ​​at the corresponding position in the fusion weight map: w1=fusion_weights[x, y, 0] (weight of the first feature primitive image), w2=fusion_weights[x, y, 1] (weight of the second feature primitive image), w3=fusion_weights[x, y, 2] (weight of the third feature primitive image), and simultaneously extract the pixel values ​​of the three feature primitive images at that position: p1=first feature primitive... Image [x, y], p2 = second feature primitive image [x, y], p3 = third feature primitive image [x, y]; calculate the fused pixel value using the weighted summation formula: fused_pixel = round(w1×p1+w2×p2+w3×p3), where the round function is used to convert the floating-point calculation result to an integer; use NumPy's clip function to limit fused_pixel to the range of 0-65535 (which conforms to the value range of a 16-bit image) to avoid pixel value overflow; assign the calculated fused_pixel to fused_image[x, y], and after traversing all pixels to complete the assignment, use the imwrite function to save fused_image as a 16-bit single-channel TIFF format, which is the final output fused image.

[0114] This embodiment also discloses an image fusion system based on infrared polarization imaging. Figure 3 This is a schematic diagram of a module of an image fusion system based on infrared polarization imaging disclosed in an embodiment of this application, as shown below. Figure 3 As shown, the system includes:

[0115] Image data acquisition module 301 is used to acquire infrared polarization image data of the target scene;

[0116] The image decomposition module 302 is used to decompose the infrared polarization image data into a first feature original primitive image, a second feature original primitive image, and a third feature original primitive image based on the polarization bidirectional reflectance distribution function.

[0117] The task processing module 303 is used to perform feature optimization processing on the first feature original primitive image, the second feature original primitive image and the third feature original primitive image respectively according to a preset task mode, so as to obtain the first feature primitive image, the second feature primitive image and the third feature primitive image.

[0118] The information abundance measurement calculation module 304 is used to calculate the information abundance measurement of the first region of the first feature primitive image, the information abundance measurement of the second region of the second feature primitive image, and the information abundance measurement of the third region of the third feature primitive image, respectively.

[0119] The complementarity measurement calculation module 305 is used to calculate the first region complementarity measurement between the first feature primitive image and the second feature primitive image, the second region complementarity measurement between the first feature primitive image and the third feature primitive image, and the third region complementarity measurement between the second feature primitive image and the third feature primitive image.

[0120] The fusion weight graph generation module 306 is used to process the information abundance measure of the first region, the information abundance measure of the second region, the information abundance measure of the third region, the complementarity measure of the first region, the complementarity measure of the second region, and the complementarity measure of the third region through a normalized exponential function to generate a fusion weight graph.

[0121] The fused image output module 307 is used to perform weighted fusion of the first feature primitive image, the second feature primitive image and the third feature primitive image using the fusion weight map, and output the fused image.

[0122] Optionally, the image decomposition module 302 is specifically used for:

[0123] A physical decomposition model is established using the polarization bidirectional reflectance distribution function; total radiation intensity information is extracted from the infrared polarized image data based on the physical decomposition model, and the first feature primitive image is constructed based on the total radiation intensity information; linear polarization degree parameters are extracted from the linear polarization state component information of the infrared polarized image data based on the physical decomposition model, and the second feature primitive image is constructed based on the linear polarization degree parameters; polarization angle parameters are extracted from the linear polarization state component information based on the physical decomposition model, and the third feature primitive image is constructed based on the polarization angle parameters.

[0124] Optional, task processing module 303, specifically used for:

[0125] When the preset task mode is target detection mode, a mapping relationship is established between the original grayscale values ​​of pixels in the first feature primitive image and the grayscale values ​​of the target. The grayscale values ​​of the target region are adjusted according to the mapping relationship to obtain the first feature primitive image. The second feature primitive image is decomposed into multiple feature layers with different spatial resolutions according to preset scale parameters. Edge features and texture features are extracted from each feature layer. The edge features and texture features are superimposed and combined according to preset weight ratios to obtain the second feature primitive image. Noise pixels in the third feature primitive image are identified through preset noise judgment conditions. All non-noise pixels within a preset range around the noise pixel are selected. The statistical values ​​of all non-noise pixels are calculated, and the pixel values ​​of the noise pixel are replaced with the statistical values ​​to obtain the third feature primitive image.

[0126] Optional, the information abundance measurement calculation module 304 is specifically used for:

[0127] The first feature primitive image, the second feature primitive image, and the third feature primitive image are divided into multiple image blocks according to a preset size, resulting in a first number of first image blocks of the first feature primitive image, a second number of second image blocks of the second feature primitive image, and a third number of third image blocks of the third feature primitive image. A first weighted sum is obtained by weighting and summing the gray-level entropy value and the average gradient magnitude of the first image blocks. The first number of first weighted sums are then arranged according to the spatial coordinates of the first image blocks in the first feature primitive image to obtain the first region information abundance measure. The gray-level entropy value characterizes the information abundance within the image block. The complexity of the grayscale distribution, wherein the mean gradient magnitude characterizes the intensity of edge features within an image patch; a second weighted sum is obtained by weighting and summing the grayscale entropy value and the mean gradient magnitude of the second image patch, and the second number of second weighted sums are arranged according to the spatial coordinate position of the second image patch in the second feature primitive image to obtain the second region information abundance measure; a third weighted sum is obtained by weighting and summing the grayscale entropy value and the mean gradient magnitude of the third target image patch, and the third number of third weighted sums are arranged according to the spatial coordinate position of the third image patch in the third feature primitive image to obtain the third region information abundance measure.

[0128] Optional, the complementarity measurement calculation module 305 is specifically used for:

[0129] Calculate the first grayscale difference entropy between the first target image block of the first feature primitive image and the second target image block of the second feature primitive image, and use the first grayscale difference entropy as the first region complementarity measure. The first target image block is any one of the first number of first image blocks, and the second target image block is an image block in the second number of second image blocks that is in the same spatial coordinates as the first target image block. Calculate the second grayscale difference entropy between the first target image block and the third target image block of the third feature primitive image, and use the second grayscale difference entropy as the second region complementarity measure. The third target image block is an image block in the third number of third image blocks that is in the same spatial coordinate range as the first target image block. Calculate the third grayscale difference entropy between the second target image block and the third target image block, and use the third grayscale difference entropy as the third region complementarity measure.

[0130] Optional, the complementarity measurement calculation module 305 is specifically used for:

[0131] The difference in grayscale values ​​between each pair of pixels in the first target image block and the second target image block is calculated and its absolute value is taken to obtain a grayscale difference data group. Each pair of pixels includes a first pixel in the first target image block and a second pixel in the second target image block. The position of the first pixel in the first target image block is the same as the position of the second pixel in the second target image block. The frequency of each difference value in the grayscale difference data group is counted, and the frequency of each difference value is divided by the total number of pixels in the first image block to obtain the probability of each difference value. The probability of each difference value is multiplied by the logarithm of the probability to obtain an intermediate value. All intermediate values ​​are then added together and the result is negative to obtain the first grayscale difference entropy. The first grayscale difference entropy is determined as the first region complementarity measure between the first target image block and the second target image block.

[0132] Optionally, the fusion weight graph generation module 306 is specifically used for:

[0133] The first region information abundance measure, the first region complementarity measure, and the second region complementarity measure are weighted and fused to obtain a first comprehensive feature weight; the second region information abundance measure, the first region complementarity measure, and the third region complementarity measure are weighted and fused to obtain a second comprehensive feature weight; the third region information abundance measure, the second region complementarity measure, and the third region complementarity measure are weighted and fused to obtain a third comprehensive feature weight; the first comprehensive feature weight, the second comprehensive feature weight, and the third comprehensive feature weight are input into a normalized exponential function to calculate a first weight value matrix corresponding to the first comprehensive feature weight, a second weight value matrix corresponding to the second comprehensive feature weight, and a third weight value matrix corresponding to the third comprehensive feature weight; a fused weight map is generated based on the first weight value matrix, the second weight value matrix, and the third weight value matrix, and the fused weight map has the same spatial size as the first feature primitive image, the second feature primitive image, and the third feature primitive image.

[0134] This embodiment also discloses an electronic device, as shown in the reference. Figure 4 The electronic device may include: at least one processor 401, at least one communication bus 402, a user interface 403, a network interface 404, and at least one memory 405. The communication bus 402 is used to enable communication between these components. The user interface 403 may include a display screen or a camera; optionally, the user interface 403 may also include a standard wired interface or a wireless interface. The network interface 404 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The processor 401 may include one or more processing cores. The processor 401 connects to various parts of the server using various interfaces and lines, and performs various functions of the server and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory 405, and by calling data stored in the memory 405. Optionally, the processor 401 may be implemented using at least one hardware form selected from Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), and Programmable Logic Array (PLA). Processor 401 may integrate one or a combination of several of the following: a central processing unit (CPU), a graphics processing unit (GPU), and a modem.

[0135] exist Figure 4 In the electronic device shown, the user interface 403 is mainly used to provide an input interface for the user and to obtain the user input data; while the processor 401 can be used to call an application program stored in the memory 405 that is an image fusion method based on infrared polarization imaging. When executed by one or more processors 401, the electronic device executes one or more methods as described in the above embodiments.

Claims

1. An image fusion method based on infrared polarization imaging, characterized in that, Applied to a server, the method includes: Acquire infrared polarization image data of the target scene; Based on the bidirectional polarization reflection distribution function, the infrared polarization image data is decomposed into a first feature primitive image, a second feature primitive image, and a third feature primitive image. According to the preset task mode, feature optimization processing is performed on the first feature original primitive image, the second feature original primitive image and the third feature original primitive image respectively to obtain the first feature primitive image, the second feature primitive image and the third feature primitive image; Calculate the first region information abundance metric of the first feature primitive image, the second region information abundance metric of the second feature primitive image, and the third region information abundance metric of the third feature primitive image, respectively. Calculate a first region complementarity measure between the first feature primitive image and the second feature primitive image, a second region complementarity measure between the first feature primitive image and the third feature primitive image, and a third region complementarity measure between the second feature primitive image and the third feature primitive image; A fusion weight map is generated based on the first region information abundance measure, the second region information abundance measure, and the third region information abundance measure, along with the first region complementarity measure, the second region complementarity measure, and the third region complementarity measure, using a normalized exponential function. The first feature primitive image, the second feature primitive image, and the third feature primitive image are weighted and fused using the fusion weight map to output a fused image.

2. The method according to claim 1, characterized in that, The process of decomposing the infrared polarization image data into a first feature primitive image, a second feature primitive image, and a third feature primitive image based on the polarization bidirectional reflectance distribution function includes: A physical decomposition model is established using the polarization bidirectional reflection distribution function; The total radiation intensity information is extracted from the infrared polarization image data according to the physical decomposition model, and the first feature original primitive image is constructed according to the total radiation intensity information; Based on the physical decomposition model, linear polarization parameters are extracted from the linear polarization state component information of the infrared polarization image data, and the second feature original primitive image is constructed based on the linear polarization parameters. The polarization angle parameter is extracted from the linear polarization state component information according to the physical decomposition model, and the original primitive image of the third feature is constructed according to the polarization angle parameter.

3. The method according to claim 1, characterized in that, The step of performing feature optimization processing on the first feature primitive image, the second feature primitive image, and the third feature primitive image according to a preset task mode to obtain the first feature primitive image, the second feature primitive image, and the third feature primitive image includes: According to the preset task mode, a mapping relationship between the original gray values ​​of pixels and the target gray values ​​in the first feature primitive image is established, and the gray values ​​of the target region are adjusted according to the mapping relationship to obtain the first feature primitive image; The original primitive image of the second feature is decomposed into multiple feature layers with different spatial resolutions according to a preset scale parameter. Edge features and texture features are extracted from each feature layer. The edge features and texture features are superimposed and combined according to a preset weight ratio to obtain the second feature primitive image. The noise pixels in the original primitive image of the third feature are identified by a preset noise judgment condition. All non-noise pixels within a preset range around the noise pixel are selected, the statistical values ​​of all the non-noise pixels are calculated, and the pixel value of the noise pixel is replaced with the statistical values ​​to obtain the third feature primitive image.

4. The method according to claim 1, characterized in that, The calculation of the first region information abundance metric of the first feature primitive image, the second region information abundance metric of the second feature primitive image, and the third region information abundance metric of the third feature primitive image includes: The first feature primitive image, the second feature primitive image, and the third feature primitive image are divided into multiple image blocks according to a preset size, resulting in a first number of first image blocks of the first feature primitive image, a second number of second image blocks of the second feature primitive image, and a third number of third image blocks of the third feature primitive image. The gray-level entropy value and the average gradient magnitude of the first image block are weighted and summed to obtain a first weighted sum. The first number of first weighted sums are arranged according to the spatial coordinate position of the first image block in the first feature primitive image to obtain the first region information abundance measure. The gray-level entropy value represents the complexity of the gray-level distribution within the image block, and the average gradient magnitude represents the intensity of the edge features within the image block. The gray-level entropy value and the average gradient magnitude of the second image block are weighted and summed to obtain a second weighted sum. The second number of second weighted sums are arranged according to the spatial coordinate position of the second image block in the second feature primitive image to obtain the second region information abundance measure. The gray-level entropy value and the average gradient magnitude of the third image block are weighted and summed to obtain a third weighted sum. The third number of third weighted sums are arranged according to the spatial coordinate position of the third image block in the third feature primitive image to obtain the information abundance measure of the third region.

5. The method according to claim 4, characterized in that, The calculation of the first region complementarity measure between the first feature primitive image and the second feature primitive image, the second region complementarity measure between the first feature primitive image and the third feature primitive image, and the third region complementarity measure between the second feature primitive image and the third feature primitive image includes: Calculate the first gray-level difference entropy between the first target image block of the first feature primitive image and the second target image block of the second feature primitive image, and use the first gray-level difference entropy as the first region complementarity measure. The first target image block is any one of the first number of first image blocks, and the second target image block is the image block in the second number of second image blocks that is in the same spatial coordinate as the first target image block. Calculate the second gray-level difference entropy between the first target image block and the third target image block of the third feature primitive image, and use the second gray-level difference entropy as the second region complementarity measure. The third target image block is the image block among the third number of third image blocks that is in the same spatial coordinate range as the first target image block. Calculate the third grayscale difference entropy between the second target image block and the third target image block, and use the third grayscale difference entropy as the complementarity measure of the third region.

6. The method according to claim 5, characterized in that, The step of calculating the first gray-level difference entropy between the first target image block of the first feature primitive image and the second target image block of the second feature primitive image, and using the first gray-level difference entropy as the complementarity measure of the first region, includes: Calculate the difference in grayscale values ​​between each pair of pixels in the first target image block and the second target image block and take the absolute value to obtain a grayscale difference data group. Each pair of pixels includes a first pixel in the first target image block and a second pixel in the second target image block. The position of the first pixel in the first target image block is the same as the position of the second pixel in the second target image block. The frequency of each difference value in the grayscale difference data group is counted, and the frequency of each difference value is divided by the total number of pixels in the first image block to obtain the probability of each difference value. The first grayscale difference entropy is obtained by multiplying the probability of each difference value by the logarithm of the probability, and then summing all the intermediate values ​​and taking the negative value. The first grayscale difference entropy is determined as the first region complementarity measure between the first target image block and the second target image block.

7. The method according to claim 1, characterized in that, The step of processing the information abundance measures of the first region, the second region, the third region, the first region complementarity measure, the second region complementarity measure, and the third region complementarity measure using a normalized exponential function to generate a fusion weight map further includes: The first region information abundance measure, the first region complementarity measure, and the second region complementarity measure are weighted and fused to obtain the first comprehensive feature weight. The second region information abundance measure, the first region complementarity measure, and the third region complementarity measure are weighted and fused to obtain the second comprehensive feature weight. The third region information abundance measure, the second region complementarity measure, and the third region complementarity measure are weighted and fused to obtain the third comprehensive feature weight; Input the first comprehensive feature weight, the second comprehensive feature weight and the third comprehensive feature weight into the normalized exponential function to calculate the first weight value matrix corresponding to the first comprehensive feature weight, the second weight value matrix corresponding to the second comprehensive feature weight and the third weight value matrix corresponding to the third comprehensive feature weight. A fused weight map is generated based on the first weight value matrix, the second weight value matrix, and the third weight value matrix. The fused weight map has the same spatial size as the first feature primitive image, the second feature primitive image, and the third feature primitive image.

8. An image fusion system based on infrared polarization imaging, characterized in that, Specifically, it includes: The image data acquisition module is used to acquire infrared polarization image data of the target scene; The image decomposition module is used to decompose the infrared polarized image data into a first feature original primitive image, a second feature original primitive image, and a third feature original primitive image based on the polarization bidirectional reflectance distribution function. The task processing module is used to perform feature optimization processing on the first feature original primitive image, the second feature original primitive image and the third feature original primitive image respectively according to a preset task mode, so as to obtain the first feature primitive image, the second feature primitive image and the third feature primitive image. The information abundance measurement calculation module is used to calculate the information abundance measurement of the first region of the first feature primitive image, the information abundance measurement of the second region of the second feature primitive image, and the information abundance measurement of the third region of the third feature primitive image, respectively. The complementarity measurement calculation module is used to calculate the first region complementarity measurement between the first feature primitive image and the second feature primitive image, the second region complementarity measurement between the first feature primitive image and the third feature primitive image, and the third region complementarity measurement between the second feature primitive image and the third feature primitive image. The fusion weight graph generation module is used to process the information abundance measure of the first region, the information abundance measure of the second region, the information abundance measure of the third region, the complementarity measure of the first region, the complementarity measure of the second region, and the complementarity measure of the third region through a normalized exponential function to generate a fusion weight graph. The fused image output module is used to perform weighted fusion of the first feature primitive image, the second feature primitive image, and the third feature primitive image using the fusion weight map, and output the fused image.

9. An electronic device, characterized in that, The device includes a processor, a memory, a user interface, and a network interface. The memory is used to store instructions. The user interface and the network interface are both used to communicate with other devices. The processor is used to execute the instructions stored in the memory to cause the electronic device to perform the method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed, perform the method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Infrared polarization image super-resolution reconstruction and polarization information fusion method

    CN118747712A

  • Multi-scale polarization data fusion method based on 4K polarization imaging detector

    CN120339096A