Spatial non-cooperative target positioning method based on multi-source image fusion

By using the methods of registration of infrared and visible light images, HSI transformation, wavelet transformation and iterative update, combined with ResNet50 feature extraction and template matching, efficient fusion of multi-source images and accurate positioning of spatial non-cooperative targets are achieved, solving the problem of poor fusion effect in existing methods and improving detection accuracy and task stability.

CN120765736APending Publication Date: 2025-10-10HARBIN INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510854214.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

Existing methods have poor fusion effects on multi-source images, resulting in insufficient detection accuracy, tracking stability and recognition capabilities of non-cooperative targets in space, making it difficult to process large amounts of data in real time in complex spatial environments.

Method used

A spatial non-cooperative target localization method based on multi-source image fusion is adopted. Through the registration of infrared and visible light images, HSI transform decomposition, weighted fusion results and weighted fusion of high-frequency components, combined with the structural similarity index for iterative update, ResNet50 is used for feature extraction and template matching to determine the target's region of interest and point cloud in the three-dimensional coordinate system.

Benefits of technology

It improves the image fusion effect, enhances the target clarity and positioning accuracy, ensures mission stability and safety in complex space environments, and has rapid response capabilities and data consistency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120765736A_ABST
    Figure CN120765736A_ABST
Patent Text Reader

Abstract

The invention discloses a spatial non-cooperative target positioning method based on multi-source image fusion, and belongs to the technical field of multi-source image fusion. According to the invention, the problem of poor multi-source image fusion effect of the existing method is solved. According to the method, respective advantages of the visible light image and the infrared image are fully utilized, the images are fused through HSI transformation and wavelet transformation methods, the fusion effect of the visible light image and the infrared image can be improved while color information is reserved, iterative optimization is carried out on the obtained fused image based on the structural similarity index, and the fusion effect of the visible light image and the infrared image is improved. Therefore, the definition of the spatial non-cooperative target in the fused image is improved. Based on the fused image, the position of the non-cooperative target in the space can be accurately positioned, and the stability and safety of task execution in a complex space environment are ensured. According to the method, the visible light image and the infrared image can be efficiently and synchronously processed, and the method has a quick response capability. The method can be applied to spatial non-cooperative target positioning.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of multi-source image fusion, and particularly relates to a space non-cooperative target positioning method based on multi-source image fusion. BACKGROUND

[0002] With the continuous development of space exploration and application, the detection, tracking and identification of space non-cooperative targets have become an important technical challenge. Space non-cooperative targets generally refer to satellites, debris, rocket wreckage and other objects that do not have active signals or markings, cannot directly communicate and cooperate with observation platforms. These targets have high complexity in orbit, and their state and behavior often cannot be controlled or predicted through traditional instructions and protocols. Therefore, how to effectively perceive, track and identify these non-cooperative targets in a complex space environment has become a key problem in space missions.

[0003] Under this background, multi-source data fusion technology plays a crucial role in the monitoring and identification of space non-cooperative targets. Due to the characteristics of space non-cooperative targets, a single sensor cannot provide sufficient information to accurately describe the state and trajectory of the target. Traditional space monitoring methods usually rely on optical, radar or infrared sensors, etc., but these sensors each have different advantages and limitations. Multi-source data fusion technology can provide more comprehensive and accurate state information for space non-cooperative targets by integrating data from different sensors, such as optical images, radar echoes, infrared thermal imaging, lidar, etc. However, the existing method still has poor fusion effect on multi-source images, therefore, it is an urgent problem to be solved to propose a new multi-source image fusion method to compensate for the shortcomings of a single sensor, to process a large amount of data in real time in a dynamic environment, and to improve the detection accuracy, tracking stability and identification ability of the target. SUMMARY

[0004] The purpose of the present application is to solve the problem of poor fusion effect of existing methods on multi-source images, and a space non-cooperative target positioning method based on multi-source image fusion is proposed.

[0005] The technical solution adopted by the present application to solve the above technical problems is: a space non-cooperative target positioning method based on multi-source image fusion, which specifically comprises the following steps:

[0006] Step one, register the infrared image and the visible light image containing the same space non-cooperative target to obtain the registered infrared image and the visible light image;

[0007] Step two, convert the registered visible light image from the RGB color space to the HSI color space to obtain I channel information, H channel information and S channel information, and then perform wavelet transform decomposition on the I channel information;

[0008] And perform wavelet transform decomposition on the registered infrared image;

[0009] Step 3: After wavelet transform decomposition, the low-frequency components of the I channel information and the low-frequency components of the infrared image are weighted fused to obtain the weighted fusion result of the low-frequency components F. L , perform weighted fusion on the high-frequency components of the I channel information and the high-frequency components of the infrared image to obtain the weighted fusion result F of the high-frequency components H ;

[0010] Step 4: Perform inverse wavelet transform on the low-frequency component fusion result and the high-frequency component fusion result, and then perform inverse HSI transform on the inverse wavelet transform result to obtain the initial fusion image, and iteratively update based on the initial fusion image to obtain the final fusion image;

[0011] Step 5: Position the non-cooperative target in space based on the final fused image.

[0012] Furthermore, the infrared image and the visible light image containing the same spatial non-cooperative target are registered to obtain the registered infrared image and visible light image; the specific process is:

[0013] Step 11: preprocessing the infrared image and visible light image containing the same spatial non-cooperative target to obtain preprocessed infrared image and visible light image;

[0014] Step 1 and 2: extract key points and key point descriptors from the preprocessed infrared image and visible light image based on ORB features;

[0015] Step 13: Match the key points in the preprocessed infrared image and visible light image according to the descriptors of the key points to obtain matching point pairs;

[0016] Step 14: Calculate the homography matrix based on the matching point pairs obtained in step 13;

[0017] Step 15: Use the calculated homography matrix to transform the preprocessed infrared image so that the preprocessed infrared image is aligned with the preprocessed visible light image, thereby achieving registration of the infrared image and the visible light image.

[0018] Furthermore, the wavelet transform decomposition obtains a low-frequency component and three high-frequency components.

[0019] Furthermore, the low-frequency components of the I channel information and the low-frequency components of the infrared image are weightedly fused to obtain a weighted fusion result F of the low-frequency components. L ; The specific process is:

[0020] F L =wL ·cA base +(1-w L )·cA target

[0021] Among them, w L is the low frequency weight;

[0022] cA base It is the low-frequency component obtained by decomposing the I channel information of the visible light image through wavelet transform;

[0023] cA target It is the low-frequency component of the infrared image obtained by wavelet transform decomposition.

[0024] Furthermore, the high frequency components of the I channel information and the high frequency components of the infrared image are weightedly fused to obtain a weighted fusion result F of the high frequency components. H ; The specific process is:

[0025] Take the horizontal component in the high-frequency component as an example:

[0026] F H =max(w H ·cH base ,(1-w H )·cH target )

[0027] Among them, cH base It is the horizontal component obtained by decomposing the I channel information of the visible light image through wavelet transform;

[0028] cH target It is the horizontal component of the infrared image obtained by wavelet transform decomposition;

[0029] w H is the horizontal component weight;

[0030] F H Represents the horizontal component fusion result.

[0031] Furthermore, the iterative update is performed based on the initial fused image to obtain the final fused image; the specific process is:

[0032] Step 4.1. Record the initial fused image as Initialize the number of iterations i = 0;

[0033] Step 4.2: Calculate the fused image Compared with the original visible light image I base The structural similarity index value between:

[0034]

[0035] in, Represents the fused image Compared with the original visible light image I base The structural similarity index between them;

[0036] μ1 represents the fused image The average brightness of all pixels in ;

[0037] σ1 represents the fused image The variance of the brightness of all pixels in ;

[0038] μ2 represents the original visible light image I base The average brightness of all pixels in ;

[0039] σ2 represents the original visible light image I base The variance of the brightness of all pixels in ;

[0040] σ 12 Represents the fused image The brightness of all pixels in the original visible light image I base The covariance of the brightness of all pixels in ;

[0041] C1 and C2 are constants;

[0042] Step 43: Determine the structural similarity index calculated in step 42 Is it greater than or equal to the threshold?

[0043] If the structural similarity index If it is greater than or equal to the threshold, the image will be fused As the final fused image;

[0044] Otherwise, execute step 44;

[0045] Step 4. Subtract the step size Δw from the low-frequency component fusion weight of the previous iteration L , add the high-frequency component fusion weight of the previous iteration to the step size Δw H , use the updated fusion weight to fuse the I channel information with the infrared image again, and record the obtained fused image as

[0046] Step 45: Set i=i+1 and return to step 42.

[0047] Furthermore, the specific process of step five is as follows:

[0048] Step 5.1: Use ResNet50 to extract features from the final fused image to obtain the extracted feature map;

[0049] Step five two, determine the target's region of interest position in the image by template matching;

[0050] Step five three, obtain the target's corresponding point cloud in the three-dimensional coordinate system according to the target's region of interest position;

[0051] Step five four, estimate the target's 3D bounding box according to the target's corresponding point cloud in the three-dimensional coordinate system.

[0052] Further, the specific process of step five two is as follows:

[0053] A two-dimensional cross-correlation operation is performed on the extracted feature map by template matching, that is, by sliding the template image, the similarity of the template image and the extracted feature map at each offset position is calculated:

[0054]

[0055] Wherein: I(x, y) is the extracted feature map;

[0056] T(x, y) is the template image;

[0057] u is the horizontal displacement of the template image relative to the extracted feature map;

[0058] v is the vertical displacement of the template image relative to the extracted feature map;

[0059] R xy (u, v) is the correlation response value of the template image and the extracted feature map at the offset position (u, v);

[0060] After traversing each offset position, the offset position corresponding to the maximum correlation response value is taken as the peak position (u * ,v * ), and the peak position (u * ,v * ) represents the best matching position of the template image and the extracted feature map;

[0061] A rectangular region with the same size as the template image is demarcated in the extracted feature map with the peak position (u * ,v * ) as the center, and the demarcated rectangular region is taken as the target's region of interest.

[0062] Further, the specific process of step five three is as follows:

[0063] The rotation matrix R and the translation vector t of the visible light camera and the laser radar are obtained by offline demarcation, and the mapping relationship between the pixel coordinate system and the three-dimensional coordinates of the laser radar point cloud is established;

[0064] A frustum is formed in the three-dimensional space with the light center of the visible light camera as the top point and the region of interest of the target as the bottom surface, and the point cloud located inside the frustum is screened according to the spatial boundary of the frustum, and the screened point cloud is the corresponding point cloud of the target in the three-dimensional coordinate system.

[0065] Further, the specific process of step five four is:

[0066] Step five four one, the point cloud obtained in step five three is randomly down-sampled, and the down-sampled point cloud is normalized to obtain a normalized point cloud;

[0067] Step five four two, the coordinates of each normalized point cloud are taken as the input of the first MLP, and the local geometric features of each point cloud are extracted;

[0068] The local geometric features of each point cloud are processed through the softmax mechanism to obtain the weight of each normalized point cloud, and the local geometric features of each point cloud are weighted and summed using the weight to obtain a global feature vector;

[0069] And the global feature vector obtained is input into the first fully connected network, and a three-dimensional vector [x c ,y c ,z c ] is output through the first fully connected network, [x c ,y c ,z c ] is the predicted target center point coordinates;

[0070] Step five four three, the three-dimensional coordinates of each point cloud obtained in step five three are subtracted from the target center point coordinates, and the difference is taken as the new coordinates of the corresponding point cloud;

[0071] Step five four four, the new coordinates of each point cloud are taken as the input of the second MLP, and the new local geometric features of each point cloud are extracted;

[0072] The new local geometric features of each point cloud are spliced with the global feature vector of step five four two, and the spliced result is taken as the input of the second fully connected network, and a predicted target size vector [dx, dy, dz] is output through the second fully connected network, dx, dy, and dz represent the length, width, and height of the target, respectively;

[0073] Then a target 3D bounding box with [x c ,y c ,z c ] as the center, dx as the length, dy as the width, and dz as the height is obtained.

[0074] The beneficial effects of the present application are:

[0075] The present invention makes full use of the respective advantages of visible light images and infrared images, and fuses images through the methods of HSI transformation and wavelet transformation. It can improve the fusion effect of visible light images and infrared images while retaining color information, and iteratively optimizes the obtained fused image based on the structural similarity index to improve the clarity of spatial non-cooperative targets in the fused image. Based on the fused image, the position of non-cooperative targets in space can be accurately located, ensuring the stability and safety of executing tasks in complex spatial environments. The method of the present invention can efficiently and synchronously process visible light images and infrared images, and has a rapid response capability, ensuring the consistency and accuracy of data, and providing a reliable foundation for subsequent data analysis and task execution. BRIEF DESCRIPTION OF THE DRAWINGS

[0076] Figure 1 It is a flow chart of a spatial non-cooperative target multi-source image fusion method based on iterative wavelet transform of the present invention;

[0077] Figure 2 It is a flowchart of image registration;

[0078] Figure 3 It is a flow chart of the non-cooperative target localization method based on fused images. DETAILED DESCRIPTION

[0079] Specific implementation method 1: Combination Figure 1 This embodiment describes a spatial non-cooperative target positioning method based on multi-source image fusion, which specifically includes the following steps:

[0080] Step 1: registering the infrared image and the visible light image containing the same spatial non-cooperative target to obtain the registered infrared image and visible light image;

[0081] Step 2: Convert the registered visible light image from RGB color space to HSI color space to obtain I channel information, H channel information, and S channel information respectively, and then perform wavelet transform decomposition on the I channel information;

[0082] By converting the visible light image from RGB color space to HSI color space, the brightness information (I channel) and color information (H, S channels) can be separated. The purpose of this conversion is to process only the brightness component during the fusion process while keeping the color information unchanged to ensure the natural color of the fused image.

[0083] And perform wavelet transform decomposition on the registered infrared image;

[0084] Step 3: After wavelet transform decomposition, the low-frequency components of the I channel information and the low-frequency components of the infrared image are weighted fused to obtain the weighted fusion result of the low-frequency components F. L , perform weighted fusion on the high-frequency components of the I channel information and the high-frequency components of the infrared image to obtain the weighted fusion result F of the high-frequency components H ;

[0085] Step 4: Perform inverse wavelet transform on the low-frequency component fusion result and the high-frequency component fusion result, and then perform inverse HSI transform on the inverse wavelet transform result to obtain the initial fusion image, and iteratively update based on the initial fusion image to obtain the final fusion image;

[0086] Step 5: Position the non-cooperative target in space based on the final fused image.

[0087] Specific implementation method 2: Combination Figure 2 This embodiment further limits the first embodiment. Due to the differences in the shooting angles and resolutions of the two cameras, it is necessary to align the infrared image and the visible light image. The infrared image and the visible light image containing the same spatial non-cooperative target are aligned to obtain the aligned infrared image and visible light image. The specific process is as follows:

[0088] Step 11: Preprocessing (including denoising and enhancing) the infrared image and visible light image containing the same spatial non-cooperative target to obtain preprocessed infrared image and visible light image;

[0089] Step 1 and 2: extract key points and key point descriptors from the preprocessed infrared image and visible light image based on ORB features;

[0090] Step 13: Match the key points in the preprocessed infrared image and visible light image according to the descriptors of the key points to obtain matching point pairs;

[0091] Step 14: Calculate the homography matrix based on the matching point pairs obtained in step 13;

[0092] Step 15: Use the calculated homography matrix to transform the preprocessed infrared image so that the preprocessed infrared image is aligned with the preprocessed visible light image, thereby achieving registration of the infrared image and the visible light image.

[0093] Other steps and parameters are the same as those in the first embodiment.

[0094] ORB feature is a combination of the detection method of FAST feature points and BRIEF feature descriptor, and it is improved and optimized on the basis of them. The FAST corner detection is responsible for extracting feature points and calculating the main direction of the feature points to enhance the robustness to rotation changes. The BRIEF descriptor generates a compact and efficient binary feature descriptor by calculating the direction information. Compared with the traditional method such as SIFT, BRIEF provides a shortcut to calculate the binary string without constructing a complex feature descriptor. Its workflow includes: first, smoothing the image; second, selecting a patch around the feature point; third, selecting N groups of point pairs in the patch according to a specific way; fourth, calculating the intensity relationship between the point pairs to finally form a binary string with a length of N.

[0095] Specific implementation three: this implementation is a further limitation of specific implementation one, and the wavelet transform decomposition obtains a low-frequency component and three high-frequency components.

[0096] The other steps and parameters are the same as those in specific implementation one.

[0097] The low-frequency component (cA, approximation coefficient) obtained by wavelet transform decomposition represents the overall information of the image, and the high-frequency components (i.e. high-frequency components cH, cV, and cD representing horizontal, vertical, and diagonal detail coefficients) represent the detail information of the image.

[0098] Specific implementation four: this implementation is a further limitation of specific implementation three, and the low-frequency component of the I channel information and the low-frequency component of the infrared image are weighted and fused to obtain the weighted fusion result F of the low-frequency component. L The specific process is as follows:

[0099] F L = w L · cA base + (1-w L )· cA target

[0100] Wherein, w L is the low-frequency weight;

[0101] cA base is the low-frequency component obtained by wavelet transform decomposition of the I channel information of the visible light image;

[0102] cA target is the low-frequency component obtained by wavelet transform decomposition of the infrared image.

[0103] The other steps and parameters are the same as those in specific implementation three.

[0104] Specific embodiment 5: This embodiment is a further limitation of specific embodiment 4, wherein the high frequency component of the I channel information and the high frequency component of the infrared image are weightedly fused to obtain a weighted fusion result F of the high frequency component. H ; The specific process is:

[0105] Take the horizontal component in the high-frequency component as an example:

[0106] F H =max(w H ·cH base ,(1-w H )·cH target )

[0107] Among them, cH base It is the horizontal component obtained by decomposing the I channel information of the visible light image through wavelet transform;

[0108] cH target It is the horizontal component of the infrared image obtained by wavelet transform decomposition;

[0109] w H is the horizontal component weight;

[0110] F H Represents the horizontal component fusion result.

[0111] The fusion method for vertical components and diagonal detail components is the same as that for horizontal components.

[0112] Other steps and parameters are the same as those in the fourth embodiment.

[0113] Specific embodiment 6: This embodiment further limits specific embodiment 5. It iteratively updates the initial fused image to obtain the final fused image. The specific process is as follows:

[0114] Step 4.1. Record the initial fused image as Initialize the number of iterations i = 0;

[0115] Step 4.2: Calculate the fused image Compared with the original visible light image I base The structural similarity index (SSIM) value between:

[0116]

[0117] in, Represents the fused image Compared with the original visible light image I base The structural similarity index between them;

[0118] μ1 represents the fused image the average brightness of all the pixel points in the image;

[0119] σ1 represents the variance of the brightness of all the pixel points in the fusion image

[0120] μ2 represents the average brightness of all the pixel points in the original visible light image I base the average brightness of all the pixel points in the image;

[0121] σ2 represents the variance of the brightness of all the pixel points in the original visible light image I base the average brightness of all the pixel points in the image;

[0122] σ 12 represents the variance of the brightness of all the pixel points in the fusion image the brightness of all the pixel points in the fusion image and the original visible light image I base the brightness of all the pixel points in the fusion image and the original visible light image I

[0123] C1 and C2 are constants;

[0124] Step four three, determine whether the structural similarity index calculated in step four two is greater than or equal to a threshold value, the greater the SSIM value, the better the fusion effect:

[0125] If the structural similarity index is greater than or equal to the threshold value, the fusion image is taken as the final fusion image;

[0126] Otherwise, step four four is executed.

[0127] Step four four, subtract the step size Δw from the low-frequency component fusion weight of the last iteration L , add the step size Δw to the high-frequency component fusion weight of the last iteration H , and re-fuse the I channel information and the infrared image using the updated fusion weight, and the obtained fusion image is denoted as

[0128] Step four five, let i = i + 1, and return to step four two.

[0129] The other steps and parameters are the same as in the fifth embodiment.

[0130] Embodiment seven: in combination with Figure 3 to illustrate the embodiment. The embodiment is a further limitation of the sixth embodiment, and the specific process of step five is as follows:

[0131] Step five one, using ResNet50 to extract features from the final fusion image to obtain the extracted feature map;

[0132] Step five two, determining the position of the region of interest of the target in the image by template matching.​​​​

[0133] Step five three, according to the target region of interest position, obtaining the corresponding point cloud of the target in the three-dimensional coordinate system;

[0134] Step five four, estimating the 3D bounding box of the target according to the corresponding point cloud of the target in the three-dimensional coordinate system.

[0135] Other steps and parameters are the same as embodiment six.

[0136] Embodiment eight: this embodiment is a further limitation of embodiment seven, and the specific process of step five two is:

[0137] A two-dimensional cross correlation operation is performed on the extracted feature map in a template matching manner, that is, by sliding the template image, the similarity of the template image and the extracted feature map at each offset position is calculated:

[0138]

[0139] Wherein: I(x, y) is the extracted feature map;

[0140] T(x, y) is the template image;

[0141] u is the horizontal displacement of the template image relative to the extracted feature map;

[0142] v is the vertical displacement of the template image relative to the extracted feature map;

[0143] R xy (u, v) is the correlation response value of the template image and the extracted feature map at the offset position (u, v);

[0144] After traversing each offset position, the offset position corresponding to the maximum correlation response value is taken as the peak position (u * ,v * ), and the peak position (u * ,v * ) represents the best matching position of the template image and the extracted feature map;

[0145] A rectangular region with the same size as the template image is demarcated in the extracted feature map with the peak position (u * ,v * ) as the center, and the demarcated rectangular region is taken as the region of interest (ROI) of the target.

[0146] Other steps and parameters are the same as embodiment seven.

[0147] Specific embodiment nine: This embodiment further limits specific embodiment eight. The specific process of step five-three is as follows:

[0148] Based on the lidar's extrinsic parameter matrix and the camera's intrinsic parameter matrix K, the rotation matrix R and translation vector t of the visible light camera and lidar are obtained through offline calibration, and the mapping relationship between the pixel coordinate system and the lidar point cloud's 3D coordinates is established.

[0149] In three-dimensional space, a viewing cone is formed with the optical center of the visible light camera as the vertex and the target's area of ​​interest as the base. The point cloud inside the viewing cone is filtered out according to the spatial boundary of the viewing cone. The filtered point cloud is the point cloud corresponding to the target in the three-dimensional coordinate system.

[0150] Other steps and parameters are the same as those in the eighth embodiment.

[0151] Specific embodiment ten: This embodiment further limits specific embodiment nine. The specific process of step five-four is as follows:

[0152] Step 541: Randomly downsample the point cloud obtained in step 53 (npoints, usually 1024 points), and then normalize the downsampled point cloud to obtain a normalized point cloud;

[0153] Downsampling can reduce the amount of computation and improve generalization;

[0154] Step 542: The normalized coordinates of each point cloud are used as the input of the first MLP (Multi-layer Perceptron) to extract the local geometric features of each point cloud;

[0155] The local geometric features of each point cloud are processed through the softmax mechanism to obtain the weight of each point cloud after normalization. The weights are used to perform weighted summation on the local geometric features of each point cloud to obtain a global feature vector used to represent the semantic information of the entire target.

[0156] The obtained global feature vector is input into the first fully connected network (composed of at least two fully connected layers), and a three-dimensional vector [x c ,y c ,z c ],[x c ,y c ,z c ] is the predicted target center coordinate;

[0157] Step 543: Subtract the three-dimensional coordinates of each point cloud obtained in step 53 from the coordinates of the target center point, and use the subtraction results as the new coordinates of the corresponding point cloud;

[0158] By performing subtraction, each point cloud can be realigned with the target center as the origin. This operation can improve the accuracy and stability of subsequent size estimation.

[0159] Step 544: Use the new coordinates of each point cloud as the input of the second MLP (Multi-layer Perceptron) to extract the new local geometric features of each point cloud;

[0160] The new local geometric features of each point cloud are spliced ​​with the global feature vector from step 542. The splicing can enhance the semantic information. The splicing result is used as the input of the second fully connected network, and the second fully connected network outputs the predicted object size vector [dx, dy, dz], where dx, dy, and dz represent the length, width, and height of the object respectively.

[0161] Then we get [x c ,y c ,z c ] as the center, with dx as the length, dy as the width, and dz as the height of the target 3D bounding box.

[0162] Other steps and parameters are the same as those in the ninth embodiment.

[0163] The present invention verifies the effectiveness of the method of the present invention by the following method:

[0164] The coordinates of the eight vertices of the 3D bounding box are transformed through the camera intrinsic parameter matrix to obtain the corresponding coordinates of these vertices on the 2D image plane. Then, the minimum enclosing rectangle is calculated based on these projected points to obtain the projected 2D bounding box. Finally, the bounding box is compared with the 2D tracking result output in step 52. The consistency between the 3D estimation result and the 2D tracking result is evaluated by calculating the intersection over union (IoU) between the two. The IoU calculation formula is as follows:

[0165]

[0166] Where A represents the projected 2D bounding box area;

[0167] B represents the 2D ROI area obtained by tracking;

[0168] |A∩B| represents the intersection area of ​​two regions;

[0169] |A∪B| represents the union area of ​​two regions.

[0170] This demonstrates the robustness and accuracy of the method of the present invention.

[0171] The above examples are merely illustrative of the calculation model and process of the present invention and are not intended to limit the embodiments of the present invention. Persons skilled in the art will readily appreciate that other variations or modifications based on the above description are possible. This list of embodiments is not exhaustive; however, any obvious variations or modifications derived from the technical solution of the present invention remain within the scope of protection of the present invention.

Claims

1. A spatial non-cooperative target localization method based on multi-source image fusion, characterized in that: The method specifically comprises the following steps: Step 1: registering the infrared image and the visible light image containing the same spatial non-cooperative target to obtain the registered infrared image and visible light image; Step 2: Convert the registered visible light image from RGB color space to HSI color space to obtain I channel information, H channel information, and S channel information respectively, and then perform wavelet transform decomposition on the I channel information; And perform wavelet transform decomposition on the registered infrared image; Step 3: After wavelet transform decomposition, the low-frequency components of the I channel information and the low-frequency components of the infrared image are weighted fused to obtain the weighted fusion result of the low-frequency components F. L , perform weighted fusion on the high-frequency components of the I channel information and the high-frequency components of the infrared image to obtain the weighted fusion result F of the high-frequency components H ; Step 4: Perform inverse wavelet transform on the low-frequency component fusion result and the high-frequency component fusion result, and then perform inverse HSI transform on the inverse wavelet transform result to obtain the initial fusion image, and iteratively update based on the initial fusion image to obtain the final fusion image; Step 5: Position the non-cooperative target in space based on the final fused image.

2. The spatial non-cooperative target positioning method based on multi-source image fusion according to claim 1, characterized in that: The infrared image and the visible light image containing the same spatial non-cooperative target are registered to obtain the registered infrared image and visible light image; the specific process is: Step 11: preprocessing the infrared image and visible light image containing the same spatial non-cooperative target to obtain preprocessed infrared image and visible light image; Step 1 and 2: extract key points and key point descriptors from the preprocessed infrared image and visible light image based on ORB features; Step 13: Match the key points in the preprocessed infrared image and visible light image according to the descriptors of the key points to obtain matching point pairs; Step 14: Calculate the homography matrix based on the matching point pairs obtained in step 13; Step 15: Use the calculated homography matrix to transform the preprocessed infrared image so that the preprocessed infrared image is aligned with the preprocessed visible light image, thereby achieving registration of the infrared image and the visible light image.

3. The spatial non-cooperative target positioning method based on multi-source image fusion according to claim 1, characterized in that: The wavelet transform decomposition obtains a low-frequency component and three high-frequency components.

4. The spatial non-cooperative target positioning method based on multi-source image fusion according to claim 3, characterized in that: The low-frequency components of the I channel information and the low-frequency components of the infrared image are weightedly fused to obtain a weighted fusion result F of the low-frequency components. L ; The specific process is: F L =w L · cA base +(1-w L )·cA target Among them, w L is the low frequency weight; cA base It is the low-frequency component obtained by decomposing the I channel information of the visible light image through wavelet transform; cA target It is the low-frequency component of the infrared image obtained by wavelet transform decomposition.

5. The spatial non-cooperative target positioning method based on multi-source image fusion according to claim 4, characterized in that: The high-frequency components of the I channel information and the high-frequency components of the infrared image are weightedly fused to obtain a weighted fusion result F of the high-frequency components. H ; The specific process is: Take the horizontal component in the high-frequency component as an example: F H =max(w H ·cH base ,(1-w H )·cH target ) Among them, cH base It is the horizontal component obtained by decomposing the I channel information of the visible light image through wavelet transform; cH target It is the horizontal component of the infrared image obtained by wavelet transform decomposition; w H is the horizontal component weight; F H Represents the horizontal component fusion result.

6. The spatial non-cooperative target positioning method based on multi-source image fusion according to claim 5, characterized in that: The iterative update is performed based on the initial fused image to obtain the final fused image; the specific process is: Step 4.

1. Record the initial fused image as Initialize the number of iterations i = 0; Step 4.2: Calculate the fused image Compared with the original visible light image I base The structural similarity index value between: in, Represents the fused image Compared with the original visible light image I base The structural similarity index between them; μ1 represents the fused image The average brightness of all pixels in ; σ1 represents the fused image The variance of the brightness of all pixels in ; μ2 represents the original visible light image I base The average brightness of all pixels in ; σ2 represents the original visible light image I base The variance of the brightness of all pixels in ; σ 12 Represents the fused image The brightness of all pixels in the original visible light image I base The covariance of the brightness of all pixels in ; C1 and C2 are constants; Step 43: Determine the structural similarity index calculated in step 42 Is it greater than or equal to the threshold? If the structural similarity index If it is greater than or equal to the threshold, the image will be fused As the final fused image; Otherwise, execute step 44; Step 4. Subtract the step size Δw from the low-frequency component fusion weight of the previous iteration L , add the high-frequency component fusion weight of the previous iteration to the step size Δw H , use the updated fusion weight to fuse the I channel information with the infrared image again, and record the obtained fused image as Step 45: Set i=i+1 and return to step 42.

7. The spatial non-cooperative target positioning method based on multi-source image fusion according to claim 6, characterized in that: The specific process of step five is as follows: Step 5.1: Use ResNet50 to extract features from the final fused image to obtain the extracted feature map; Step 52: Determine the location of the region of interest of the target in the image by template matching; Step 53: According to the location of the target's region of interest, obtain the point cloud corresponding to the target in the three-dimensional coordinate system; Step 54: Estimate the 3D bounding box of the target based on the point cloud corresponding to the target in the three-dimensional coordinate system.

8. The spatial non-cooperative target positioning method based on multi-source image fusion according to claim 7, characterized in that: The specific process of step 52 is as follows: The extracted feature map is subjected to a two-dimensional cross-correlation operation using template matching. That is, by sliding the template image, the similarity between the template image and the extracted feature map at each offset position is calculated: Where: I(x,y) is the extracted feature map; T(x,y) template image; u is the horizontal displacement of the template image relative to the extracted feature map; v is the vertical displacement of the template image relative to the extracted feature map; R xy (u, v) is the correlation response value between the template image and the extracted feature map at the offset position (u, v); After traversing each offset position, the offset position corresponding to the maximum correlation response value is taken as the peak position (u * ,v * ), peak position (u * ,v * ) represents the position where the template image best matches the extracted feature map; In the extracted feature map, a peak position (u * ,v * ) as the center and a rectangular area with the same size as the template image, and the calibrated rectangular area is used as the target's region of interest.

9. The spatial non-cooperative target positioning method based on multi-source image fusion according to claim 8, characterized in that: The specific process of step 53 is as follows: Obtain the rotation matrix R and translation vector t of the visible light camera and lidar through offline calibration, and establish the mapping relationship between the pixel coordinate system and the three-dimensional coordinates of the lidar point cloud; In three-dimensional space, a viewing cone is formed with the optical center of the visible light camera as the vertex and the target's area of ​​interest as the base. The point cloud inside the viewing cone is filtered out according to the spatial boundary of the viewing cone. The filtered point cloud is the point cloud corresponding to the target in the three-dimensional coordinate system.

10. The spatial non-cooperative target positioning method based on multi-source image fusion according to claim 9, characterized in that: The specific process of step 54 is as follows: Step 541: randomly downsample the point cloud obtained in step 53, and then normalize the downsampled point cloud to obtain a normalized point cloud; Step 542: The normalized coordinates of each point cloud are used as the input of the first MLP to extract the local geometric features of each point cloud; The local geometric features of each point cloud are processed through the softmax mechanism to obtain the weight of each point cloud after normalization. The weights are used to perform weighted summation on the local geometric features of each point cloud to obtain the global feature vector. The obtained global feature vector is input into the first fully connected network, and a three-dimensional vector [x c ,y c ,z c ],[x c ,y c ,z c ] is the predicted target center coordinate; Step 543: Subtract the three-dimensional coordinates of each point cloud obtained in step 53 from the coordinates of the target center point, and use the subtraction results as the new coordinates of the corresponding point cloud; Step 544: Use the new coordinates of each point cloud as the input of the second MLP to extract the new local geometric features of each point cloud; Concatenate the new local geometric features of each point cloud with the global feature vector from step 542, and use the concatenated result as the input of the second fully connected network. The second fully connected network outputs the predicted object size vector [dx, dy, dz], where dx, dy, and dz represent the length, width, and height of the object, respectively. Then we get [x c ,y c ,z c ] as the center, dx as the length, dy as the width, and dz as the height of the target 3D bounding box.