Heterogeneous image target matching method and system for front downward-looking fixed target

Through projection transformation of remote sensing template diagrams and heterologous adaptive neural network feature extraction, combined with the feature matching of convolutional neural networks, the problem of matching the top-view remote sensing image and the front-down view infrared image in aircraft guidance is solved, and accurate target recognition and matching under low-safety conditions is achieved.

CN120451603APending Publication Date: 2025-08-08HUAZHONG UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510500775.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

During the aircraft guidance process, there are significant posture and scale differences in the matching of heterologous image between the top-view remote sensing image and the front-down infrared image. The difference in imaging principles and time phases make it difficult to match, and the long-distance targets are small in real-time images and the details are blurred, resulting in false alarms and algorithm failure.

Method used

The remote sensing template diagram is projected and transformed using the aircraft inertial navigation information and camera parameters. A heterologous adaptive neural network is used to extract feature points and combine the convolutional neural network for feature matching. Through the strategy of combining coarse matching and fine matching, a template library is built and confidence judgment is performed, and the template is updated to improve matching accuracy.

Benefits of technology

Under the low-income conditions, the precise target matching between the lateral view remote sensing image and the front-down view infrared image is achieved, which improves the accuracy and accuracy of target recognition, reduces false alarms, and improves the robustness and accuracy of matching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120451603A_ABST
    Figure CN120451603A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of image processing, and discloses a heterogenous image target matching method and system for front downward-looking fixed targets, and the method comprises the steps: carrying out the attitude transformation of a remote sensing template image according to inertial navigation information with a certain accumulative error; then extracting feature points from the front down-looking real-time image and the remote sensing template image after attitude change by using a neural network method, and carrying out image registration based on matched feature point pairs to realize target coarse positioning; when the aircraft is close to the target, enlarging the coarse positioning target frame by a certain multiple to intercept an image block as a new to-be-matched image; taking an area only containing the established target in the remote sensing template image as an initial target template, and performing template matching identification on the initial target template and a new image to be matched to realize fine positioning of the target; in the flight process of the aircraft, a new target template is selected to replace the initial target template according to the confidence degree of the matching result, and the fine positioning precision and the success rate are improved by updating the target template.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image processing technology, and more specifically, relates to a heterogeneous image target matching method and system for a front-view downward fixed target. Background Art

[0002] Image matching is a key technology in computer vision. Its purpose is to find the subregion within a single image or a batch of images that is most similar to a given target image of interest. Typically, the input to an image matching algorithm is an image pair (a template image and an image to be matched), with the target image serving as the template and the image to be matched serving as the template's matching target.

[0003] Deep learning methods possess powerful learning and generalization capabilities, capable of learning complex features and patterns in data. Consequently, deep learning-based matching methods have achieved very high performance and accuracy in many tasks. Convolutional neural networks (CNNs) excel in tasks such as image matching, with their convolutional layers effectively capturing local patterns within images. While Transformers can capture global features through their self-attention mechanism, CNNs typically infer faster data in real-time or resource-limited tasks, making them more suitable for hardware implementation. Therefore, using CNNs for matching homologous images can achieve relatively accurate matching results.

[0004] In some applications, due to the influence of imaging time, environment, angle, etc., the target undergoes significant changes, resulting in significant differences between the template image and the image to be matched, which increases the difficulty of matching. For example, in the guidance process of an aircraft, the target template is a downward-looking remote sensing image (visible light image) provided by a satellite, but the real-time image to be matched provided by the camera on board the aircraft is a small-angle (the angle between the camera on board the aircraft and the horizon is small) forward-looking infrared image (the view taken while the aircraft is flying forward while tilting downward). For the downward-looking remote sensing image as a template and the small-angle heterogeneous forward-looking real-time image, it is very difficult to directly match and identify the two. Specifically, it is manifested in the following aspects: (1) There are significant differences in posture and scale between the large-angle remote sensing image with a normal downward view and the small-angle forward-looking infrared image, which makes direct image matching challenging. (2) There are differences in imaging principles, shooting time, radiation detection system and time phase between the template image and the real-time image, resulting in larger image detail features. In addition, unlike conventional 2D to 2D or 3D to 3D (the 2D image also contains the height information of the target) image matching, which has mature data sets, there is no matching data set for matching the 2D remote sensing template image with the 3D real-time image. In addition, the requirements for self-made data sets are high and the difficulty is relatively high. (3) Target detection starts from a relatively long distance, and the actual target area in the real-time image detected by the aircraft is relatively small. For example, if the aircraft starts from 10 km away and the length, width and height of a specific building target are 20 meters respectively, its area in the real-time image is approximately 16*16 pixels. In addition, due to factors such as long-distance atmospheric disturbances and atmospheric scattering, the target details in the real-time image are blurred and deformed. Directly using the target template for matching and recognition will cause a large number of false alarms or even algorithm failure. Summary of the Invention

[0005] In response to the above defects or improvement needs of the prior art, the present invention provides a method and system for matching heterogeneous images of fixed targets of the front and downward view, which aims to improve the target matching accuracy of heterogeneous images of fixed targets of the front and downward view.

[0006] To achieve the above objectives, the present invention provides a heterogeneous image target matching method for a forward-looking, downward-looking, fixed target, comprising:

[0007] S1, using the real-time inertial navigation information of the aircraft and the camera parameters used by the aircraft to shoot the infrared front-down real-time image, the initial remote sensing template image P is projected and transformed to obtain the image R i A remote sensing template image P2 whose viewing angle, size, and spatial resolution are consistent within a preset error range; wherein i is the frame sequence number of the infrared front-down real-time image taken by the aircraft;

[0008] S2, using heterogeneous adaptive neural network to extract the remote sensing template image P2 and the real-time image R iThe features of the remote sensing template map P2 and the real-time map R are calculated. i Based on the coordinates of the feature point pairs, the remote sensing template map P2 to the real-time map R i The coordinate transformation matrix is used to transform the target coordinates in the remote sensing template map P2 to the real-time map R i In the real-time graph R i The target box in

[0009] S3, let i=i+1, repeat step S2, and get the real-time graph R i+1 The target box in

[0010] S4, real-time image R of two adjacent frames i and R i+1 Make the association. When the association is successful, R i and R i+1 The mean of the target box in R is used as the rough matching target box, and when the target is in R i+1 When the area in is lower than a preset value, the coarse matching target frame is used as the current matching result, and target matching of the next frame is performed.

[0011] Furthermore, when the target is in R i+1 When the area in exceeds a preset value, further comprising performing fine matching based on the coarse matching result, specifically comprising:

[0012] In real-time graph R i+1 In the first step of fine matching, the image block containing the target position intercepted from the remote sensing template image P2 is used as the template image; in the subsequent fine matching, the target frame obtained by fine matching in the real-time image of the previous fine matching is intercepted as the template image;

[0013] A convolutional neural network is used to perform fine matching on the current real-time image to be matched and the template image to obtain a fine matching target frame, and the fine matching target frame is used as the current matching result, and target matching of the next frame is performed.

[0014] Furthermore, after performing the fine matching on the multiple frames of real-time images, the method further includes:

[0015] Construct a fine matching template library, which includes templates t1, t2, and t3; wherein template t1 is an image block containing the target position intercepted from the remote sensing template image P2; template t3 is used in the fine matching process and has a Euclidean distance d between it and the template t1 for the first time. k Satisfy d kTemplate diagram of Q1 relationship, where Q1 is a set threshold; template t2 is used in the fine matching process and the Euclidean distance d between it and template t1 for the second time k satisfies d k Template diagram of <Q1 relationship;

[0016] Based on the current fine matching result, perform template confidence discrimination, specifically including:

[0017] Calculate the Euclidean distances d1, d2, and d3 between the template diagrams used in the current fine matching process and templates t1, t2, and t3 in the template library respectively, and obtain the average Euclidean distance If d < Q2, where Q2 is a set threshold, it is considered that the fine matching is successful; otherwise, it is considered that the fine matching fails, and the target matching for the next frame is performed;

[0018] When the fine matching is successful, calculate the Euclidean distances d4 and d5 between template t1 and templates t2 and t3 respectively. If (d1 + d2) < (d4 + d5), then update the template library: use template t2 as the current template t3, use the template diagram used in the current fine matching process as the current template t2, and perform the target matching for the next frame; otherwise, do not update the template library.

[0019] Furthermore, in S1, the inertial navigation information includes the height H of the aircraft relative to the ground, the latitude lat0 and longitude lon0 of the aircraft, the pitch angle anglex, roll angle angley, and yaw angle anglez of the aircraft; the camera parameters include the vertical viewable angle FOV zhi and the horizontal viewable angle FOV ping ;

[0020] S1 includes:

[0021] Calculate the horizontal distance r i between four boundary points of the aircraft to the real-time image R x and the projection distance r y :

[0022]

[0023] Calculate the distance r and angle θ between the aircraft and four boundary points of the real-time image R i :

[0024]

[0025] Calculate the latitudes lat i of four boundary points corresponding to four boundary points in the initial remote sensing template image P and the real-time image R j and longitudes lonj :

[0026]

[0027] Wherein, j=1, 2, 3, 4, respectively represent the four boundary points corresponding to the distance r;

[0028] Based on the image resolution res and longitude and latitude of the initial remote sensing template map P, the image resolution res and longitude and latitude of the initial remote sensing template map P are obtained. i The coordinates of the four boundary points corresponding to the four boundary points (Δx j ,Δy j ); and according to the coordinates (Δx j ,Δy j ) performing cropping on the initial remote sensing template image P to obtain a cropped remote sensing template image P1;

[0029] Use the real-time graph R i The coordinates of the four boundary points (x j ,y j ) and the coordinates (Δx j ,Δy j ) to transform the remote sensing template image P1 into the remote sensing template image P2.

[0030] Furthermore, in S2, the method for calculating the coordinates of the feature points includes:

[0031] The remote sensing template image P2 and the real-time image R i The features of the remote sensing template map P2 and the real-time map R are respectively obtained. i Coordinates of key feature points in ;

[0032] Based on the coordinates of the key feature points, the remote sensing template image P2 and the real-time image R are searched using a fast nearest neighbor search algorithm. i Perform rough matching on the key feature points in the image to obtain the coordinates of the feature point pairs after rough matching;

[0033] The dynamic adaptive distance constraint algorithm is used to dynamically and adaptively eliminate the feature point pairs after rough matching. The elimination conditions are:

[0034] dis j ≥dis′ j -avgdis;

[0035]

[0036] Among them, dis j 、dis′ jare the nearest neighbor and the next nearest neighbor of the jth pair of feature points, respectively, and N is the number of feature point pairs after the rough matching;

[0037] After the feature point pairs are eliminated, the remaining feature point pairs are accurately matched using a random sampling consensus algorithm to obtain the precise coordinates of the feature point pairs, which are used as the coordinates of the required feature point pairs.

[0038] Furthermore, in S2, the heterogeneous adaptive neural network is a VGG16 network, and the last convolutional layer is discarded, and the output of the fourth convolutional layer of the VGG16 network is directly selected as the remote sensing template image P2 and the real-time image R i characteristics.

[0039] Furthermore, in S4, two adjacent frames of real-time images R i and R i+1 When the length ratio and width ratio of are within the preset threshold range, the IOU of the two frames of real-time images is calculated:

[0040]

[0041] Among them, A and B are the target frames of two frames of real-time images respectively;

[0042] When IOU>Y, where Y is the set threshold, the two frames of real-time images are successfully associated.

[0043] The present invention also provides a heterogeneous image target matching system for a front-downward fixed target, comprising a computer-readable storage medium and a processor;

[0044] The computer-readable storage medium is used to store executable instructions;

[0045] The processor is configured to read the executable instructions stored in the computer-readable storage medium to execute any one of the above-mentioned heterogeneous image target matching methods for front-downward fixed targets.

[0046] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the heterogeneous image target matching method for a front-down fixed target as described in any one of the above items.

[0047] The present invention also provides a computer program product, comprising a computer program, which, when executed on a computer, enables the computer to execute any of the above-mentioned heterogeneous image target matching methods for forward-looking, downward-looking, fixed targets.

[0048] In general, the above technical solutions conceived by the present invention can achieve the following beneficial effects:

[0049] (1) The present invention performs real-time projection transformation on the remote sensing template image based on the inertial navigation information of the aircraft, transforming the top-down view into a front-down view. The feature distribution of the transformed remote sensing template image is closer to the real-time image to be matched. Feature point matching is performed based on the remote sensing template image after projection transformation and the implementation image, which improves the accuracy of target matching. In view of the fact that there are differences in detection mechanism and time phase between the remote sensing image and the real-time image, a neural network that supports the extraction of consistent features of heterogeneous images is used to extract features from the transformed remote sensing template image and the real-time image; considering that there is no matching data set for the matching of the two-dimensional remote sensing template image and the three-dimensional real-time image, the remote sensing template image P2 and the real-time image R are calculated based on the extracted features. i The coordinates of the feature point pairs are calculated based on the coordinates of the point pairs. The remote sensing template image P2 is converted to the real-time image R. i The coordinate transformation matrix is used to transform the target position in the template image to the current real-time image, thereby achieving point-matching positioning of the target and improving the accuracy of target point matching. This invention utilizes only the remote sensing template image and inertial navigation information with a certain cumulative error obtained by sensors onboard the aircraft. Under low security conditions, it achieves relatively accurate target matching and recognition between biased downward-looking, large-angle remote sensing images and small-angle, forward-looking, downward-looking infrared images.

[0050] (2) Furthermore, to address the problem of blurred local features of distant targets, when the area of the target in the real-time image increases to a certain threshold, a strategy of combining coarse and fine positioning of the target and gradually refining the identification area is adopted. The local area of the real-time image obtained on the basis of coarse matching is used as the new real-time image to be matched, and target template matching is performed for fine positioning, which further improves the target matching accuracy.

[0051] (3) Furthermore, after performing the fine matching on multiple frames of real-time images, the method further includes constructing a fine matching template library, performing template confidence judgment on the current fine matching result, and updating the template library according to the confidence to further improve the matching accuracy.

[0052] (4) As a preference, the initial remote sensing template image P is transformed in attitude according to the inertial navigation information and camera parameters of the aircraft, thereby reducing the significant pitch, yaw and roll angle differences and scale differences between the remote sensing image of the reference template and the real-time image.

[0053] (5) As a preferred method, the calculation method of the feature point coordinates in the present invention is to first use a fast nearest neighbor search package algorithm to calculate the remote sensing template image P2 and the real-time image R iThe key feature points in the image are roughly matched, and the dynamic adaptive distance constraint algorithm is used to dynamically and adaptively eliminate the feature point pairs after rough matching, thereby improving the heterogeneous adaptability between remote sensing images and infrared real-time images. After the feature point pairs are eliminated, the remaining feature point pairs are precisely matched using a random sampling consistency algorithm to obtain accurate feature point pair coordinates. The accuracy of the coordinate transformation matrix obtained based on the accurate feature point pair coordinates is also high, which improves the accuracy of point matching.

[0054] (6) As a preferred method, the VGG16 model without the last convolutional layer is used as the heterogeneous adaptive neural network, which can improve the accuracy of spatial positioning while integrating the advantages of shallow detail information and deep abstract features.

[0055] In summary, the present invention targets aircraft moving platforms and uses deep learning image matching and recognition technology to achieve low-security heterogeneous image target matching and recognition of fixed, low-profile targets in complex ground scenes using inertial navigation information, remote sensing template maps, and target longitude and latitude. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 Schematic diagram of a heterogeneous image target matching method for a forward-looking, downward-looking, fixed target according to an embodiment of the present invention;

[0057] Figure 2 Schematic diagram of a method for matching heterogeneous images of fixed forward-looking downward targets with combined coarse and fine coordinates in an embodiment of the present invention;

[0058] Figure 3 This is the overall strategy for target confidence determination and template updating in the embodiment of the present invention;

[0059] Figure 4 This is an example of a HW real-time graph in an embodiment of the present invention;

[0060] Figure 5 is an example of a visible light remote sensing image in an embodiment of the present invention;

[0061] Figure 6 Schematic diagram of the aircraft posture in an embodiment of the present invention;

[0062] Figure 7 The cropped remote sensing template image P1 in the embodiment of the present invention;

[0063] Figure 8 The remote sensing image P2 after attitude transformation in the embodiment of the present invention;

[0064] Figure 9 is a point pair matching coarse positioning network in an embodiment of the present invention, Figure 9(a)-(c) are the overall framework diagram, the backbone structure of the network feature map extraction, and the process of extracting key points based on the feature map;

[0065] Figure 10 This is an example of point pair matching coarse positioning in an embodiment of the present invention. The left side is a real-time image, and the box is the coarse positioning result obtained by transformation; the right side is a template image after posture transformation, and the box is the specified target box;

[0066] Figure 11 This is an example of template matching fine positioning in an embodiment of the present invention. The left side is the template image, and the right side is the clipped real-time image matching result. DETAILED DESCRIPTION

[0067] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only intended to illustrate the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below may be combined with each other as long as they do not conflict with each other.

[0068] Example 1

[0069] like Figure 1 As shown, an embodiment of the present invention provides a heterogeneous image target matching method for a forward-looking, downward-looking, fixed target, including:

[0070] S1, using the sensors on board the aircraft to obtain the aircraft's real-time inertial navigation information, using the inertial navigation information and the aircraft's camera parameters used to shoot infrared real-time images, to perform a projection transformation on the initial remote sensing template image P, so that the transformed remote sensing template image P2 is aligned with the real-time image R i The viewing angle, size, and spatial resolution of the image are consistent within the preset error range; where i is the frame sequence number of the infrared real-time image taken by the aircraft;

[0071] S2, using heterogeneous adaptive neural networks to extract remote sensing template images P2 and real-time images R i The characteristics of the remote sensing template map P2 and the real-time map R i The coordinates of the feature point pairs are converted to the remote sensing template map P2 and the real-time map R i Transformation matrix for coordinate transformation; based on the transformation matrix, the target coordinates in the remote sensing template map P2 are transformed to the real-time map R i , in order to achieve the target coarse positioning, get the real-time map R i The target box in i ,w i ,h i ), track the target in a larger area. i is the center point coordinate of the target frame, wi is the width of the target box, h i is the height of the target box, and the units are all pixels;

[0072] S3. Let i = i + 1, and repeat step S2 to obtain the real-time image R i+1 of the target box (cen i+1 , w i+1 , h i+1 ); where cen i+1 is the center point coordinate of the target box, w i+1 is the width of the target box, h i+1 is the height of the target box, and the units are all pixels;

[0073] S4. When the length ratio and width ratio of two adjacent frames R i and R i+1 are within the preset threshold range, perform the two-frame associated IOU calculation A and B are the target boxes of the two frames respectively; if IOU > Y, where Y is the set threshold, then the two frames are successfully associated, and the mean value of the two-frame target boxes is used as the matched target box (cen' i+1 , w' i+1 , h' i+1 ); where, cen' i+1 is the center point coordinate of the target box, w' i+1 is the width of the target box, h' i+1 is the height of the target box, and the units are all pixels, and If IOU < Y or the length ratio and width ratio of two adjacent frames R i and R i+1 are not within the preset threshold range, it is considered that the matching fails and returns to step S1 to re-match the next frame of the real-time image.

[0074] When the aircraft is far from the target, the target imaging in the real-time image lacks sufficient pixel count at this time, that is, w' i+1 < M and h' i+1 < M, where M is the set threshold, and the rough positioning result after two-frame association is used as the final target box (cen' i+1 , w' i+1 , h' i+1 ).

[0075] As Figure 2 , Figure 3 shown, as a further design of the present invention, as the distance between the aircraft and the target decreases, when the area of the target in the real-time image increases to a certain threshold (in the embodiment of the present invention, when the distance between the aircraft and the target is within 4 km), that is, w' i+1 > M or h' i+1When it is >M, it also includes a fine positioning process:

[0076] S5. In the current frame real-time image R i+1 截取以当前所得粗定位目标框扩大预设倍数的图像块作为新的待匹配图;在上一次细匹配的实时图中截取上一次细匹配所得目标框作为新的模板图s i+1 , use a convolutional neural network for fine template matching to obtain the target box (cen” i+1 ,w” i+1 ,h” i+1 ), where cen” i+1 is the center point coordinate of the target box, w” i+1 is the width of the target box, and h” i+1 is the height of the target box, and the units are all pixels. Among them, when initially performing fine positioning, the new template image is an image block containing the target position intercepted from the remote sensing template image P2.

[0077] S6. After performing fine positioning for a period of time, it also includes constructing a fine matching template library. The template library includes templates t1, t2, t3, that is Figure 3 the visible light remote sensing image t1, the short-term HW template image t2, and the long-term HW template image t3 in k ; among them, template t1 is an image block containing the target position intercepted from the remote sensing template image P2; template t3 is the Euclidean distance d k between the template t1 and the template t1 for the first time during the fine positioning process k satisfies d k <Q1 relationship template image, Q1 is a set threshold; template t2 is the Euclidean distance d

[0078] S7. During the flight of the aircraft, perform template confidence discrimination according to the fine positioning matching result, specifically including:

[0079] Calculate the Euclidean distances d1, d2, d3 between the template images used in the current fine matching process and the three templates t1, t2, t3 in the template library respectively, and obtain the average Euclidean distance If d < Q2, Q2 is a set threshold, it is considered that the fine matching is successful;

[0080] When d < Q2, that is, when the fine matching is successful, calculate the Euclidean distance d4 between template t1 and template t2, and calculate the Euclidean distance d5 between template t1 and template t3. If (d1 + d2) < (d4 + d5), then update the template library: discard the original template t3, use template t2 as the current template t3, use the template image used in the current fine matching process as the current template t2, and enter the matching or tracking of the next frame of real-time image; if (d1 + d2) ≥ (d4 + d5), then do not update the template library.

[0081] If d ≥ Q2, it is considered that the matching fails, and enter the target matching of the next frame of real-time image.

[0082] In the embodiment of the present invention, first capture the target through target matching, then enter the tracking mode. If the tracking fails, enter the target matching mode again.

[0083] The relevant thresholds in the embodiment of the present invention are empirical values and are set according to the actual situation.

[0084] In the embodiment of the present invention, the actual size of the target is 20 * 20 m, the distance between the aircraft and the target is 10 km, the flight altitude of the aircraft is 3 km, and the infrared detector generates 640 * 512 real-time HW data (infrared real-time image) at a frame frequency of 50 Hz, as Figure 4 shown, where HW represents infrared. The initial template image is a large-range visible light remote sensing image covering the target area, as Figure 5 shown.

[0085] Preferably, in S1, the inertial navigation information includes the current position of the aircraft (the height H of the aircraft relative to the ground and the longitude and latitude of the aircraft) and the three-axis attitude of the aircraft (pitch angle anglex, roll angle angley, yaw angle anglez). The camera parameters include: the vertical viewing angle FOV of the camera zhi and the horizontal viewing angle FOV ping .

[0086] S1 includes:

[0087] [[ID=二十九]]S11. Calculate the horizontal distance r i between the four boundary points of the aircraft to the real-time image R x in the aircraft coordinate system and the vertical (projection) distance r y :

[0088]

[0089] In the formula, the plus sign “+” in r x and r y corresponds to the real-time image R iThere are two upper boundary points, and the minus sign "-" corresponds to two lower boundaries; the upper boundary and the lower boundary are the upper boundary and the lower boundary of the imaging field of view, respectively.

[0090] S12, after correcting the direction with the yaw angle anglez, the aircraft goes to the real-time map R i The distance r and angle θ between the four boundary points are:

[0091]

[0092] Where, Corresponding to the angle between the two left boundary points in the imaging field of view, Corresponding to the angle between the two boundary points on the right side of the imaging field of view.

[0093] Based on the above distance r and angle θ, according to the radius of the earth R = 6371.0 km, the latitude lat0 and longitude lon0 of the aircraft, the initial remote sensing template map P and the real-time map R are calculated. i The latitudes of the four boundary points corresponding to the four boundary points j and longitude lon j (j=1,2,3,4), such as Figure 6 As shown:

[0094]

[0095] S13, through the image resolution res (pixel / km) of the initial remote sensing template map P and the longitude lon and latitude lat of the initial remote sensing template map P, obtain the initial remote sensing template map P and the real-time map R i The coordinates of the four boundary points corresponding to the four boundary points (Δx j ,Δy j ); According to the coordinates (Δx j ,Δy j ) is cropped on the initial remote sensing template image P to obtain the cropped remote sensing template image P1, as shown in Figure 7 As shown. In the image coordinate system, the coordinates of the upper left corner boundary point of the initial remote sensing template image P are (0,0). Therefore, in the embodiment of the present invention, the longitude lon and latitude lat of the upper left corner boundary point of the initial remote sensing template image P are used. Correspondingly, the coordinates (Δx j ,Δy j ) is calculated as:

[0096]

[0097] Δy j =(lat j -lat)×111×res

[0098] S14, through the real-time graph boundary four point coordinates (x j ,y j ) and the corresponding template Figure 4 Point coordinates (Δx j ,Δy j ), list the equations to obtain the transformation matrix H, and transform the cropped remote sensing template image P1 into a remote sensing template image P2 that is similar to the real-time image in terms of viewing angle, size, and spatial resolution, as shown in the following example: Figure 8 As shown, the viewing angle is close to real-time. The formula is as follows: each pair of points corresponds to two equations, and four pairs of points correspond to eight equations, which are used to obtain the eight parameters in the transformation matrix H.

[0099]

[0100] Preferably, in S2, the heterogeneous adaptive neural network is an improved VGG16 network. The original VGG16 model, as a typical deep convolutional neural network, includes five stacked convolutional layers and is mainly designed for image classification tasks. A notable feature of this network architecture is that the shallow convolutional layers can capture local, low-level features such as edges and corners, which have high spatial positioning accuracy. As the depth of the network increases, the extracted features gradually transform into more abstract and high-level global information, which enhances the robustness of the model to interference from images from different sources, but correspondingly sacrifices the accuracy of spatial positioning. In order to balance the high-level abstractness of features and the ability to accurately locate them, in an embodiment of the present invention, the last convolutional layer of VGG16 is discarded, and the output of the last convolutional layer of the fourth convolutional layer (i.e., the Conv4_3 layer) is directly selected as the feature map for key point extraction. This choice aims to combine the advantages of shallow detail information and deep abstract features, thereby optimizing the detection performance of feature points. The overall structure of the algorithm is as follows: Figure 9 As shown in (a), w and h represent the width and height of the image respectively, and the network backbone is as follows Figure 9 In other embodiments, heterogeneous adaptive neural networks such as ResNet may also be used.

[0101] In S2, the remote sensing template map P2 and the real-time map R i The calculation method of the coordinates of the feature point pairs includes:

[0102] Remote sensing template map P2 and real-time map R i The features are screened to obtain the remote sensing template map P2 and the real-time map R i The coordinates of the key feature points in the feature map are the locations with the largest feature values in the channel. The overall process is as follows Figure 9 As shown in (c) in .

[0103] According to the remote sensing template map P2 and real-time map R iThe key feature point coordinates in the remote sensing template map P2 and the real-time map R i Match the key feature points in to obtain feature point pairs and their coordinates, including:

[0104] According to the remote sensing template map P2 and real-time map R i The key feature point coordinates in the remote sensing template map P2 and the real-time map R are obtained by using the Fast Library for Approximate Nearest Neighbors (FLANN) algorithm. i Perform rough matching on the key feature points in the image to obtain the feature point pairs after rough matching;

[0105] The dynamic adaptive distance constraint algorithm is used to dynamically and adaptively eliminate the feature point pairs after rough matching; the elimination conditions are:

[0106] dis j ≥dis′ j -avgdis;

[0107]

[0108] Among them, dis j sid′ j are the nearest neighbor and the next nearest neighbor of the jth pair of feature points, respectively, and N is the number of feature point pairs after rough matching.

[0109] After removing the feature point pairs, the remaining feature point pairs are precisely matched using a random sampling consensus algorithm to obtain precise feature point pairs and their coordinates. In this embodiment of the present invention, the MAGSAC++ algorithm is used for precise matching to obtain precise feature point pairs and their coordinates. Then, based on the coordinates of the precise feature point pairs, the remote sensing template image P2 to the real-time image R is obtained. i The transformation matrix of coordinate transformation is used to determine the rough positioning coordinates of the target, such as Figure 10 shown.

[0110] As a preference, in S4, two adjacent frames R i and R i+1 The length ratio and width ratio of When , the two-frame associated IOU calculation is performed.

[0111] In S5, the current frame real-time graph R i+1 The size of the image block intercepted is L*L, L>w' i+1 , L>h' i+1 In the embodiment of the present invention, s is the distance between the aircraft and the target, ensuring that the L*L image block can cover the target at different distances.

[0112] As a preference, in S5, the neural network used for fine positioning is the classic residual neural network ResNet-18. In other embodiments, other convolutional neural networks can also be used for template matching fine positioning. In this embodiment of the present invention, the result of fine matching is as follows: Figure 11 shown.

[0113] In this embodiment of the present invention, the template image is first projected based on the position and attitude parameters provided by the moving platform (aircraft) to minimize perspective differences. Because the template image and the real-time image are derived from different sources and there are errors in the attitude and position parameters, a deep learning method with excellent generalization capabilities for scale, perspective, and spectral range is used for image matching. Furthermore, to improve the accuracy of close-range matching, a hierarchical matching strategy from coarse to fine is proposed.

[0114] Example 2

[0115] An embodiment of the present invention provides a heterogeneous image target matching system for a front-downward fixed target, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps of the heterogeneous image target matching method for a front-downward fixed target in the above-mentioned embodiment 1.

[0116] The relevant technical solutions are the same as above and will not be repeated here.

[0117] Example 3

[0118] An embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the heterogeneous image target matching method for forward-looking and downward-looking fixed targets in the above-mentioned embodiment 1 are implemented.

[0119] Specifically, the memory may include a high-speed random access memory, and may also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage device.

[0120] The relevant technical solutions are the same as above and will not be repeated here.

[0121] Example 4

[0122] An embodiment of the present application provides a computer program product, including a computer program. When the computer program is run on a computer, it enables the computer to execute the steps of the heterogeneous image target matching method for forward-looking and downward-looking fixed targets in the above-mentioned embodiment 1.

[0123] The relevant technical solutions are the same as above and will not be repeated here.

[0124] It will be easily understood by those skilled in the art that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A heterogeneous image target matching method for a forward-looking, downward-looking, fixed target, characterized in that: include: S1, using the real-time inertial navigation information of the aircraft and the camera parameters used by the aircraft to shoot the infrared front-down real-time image, the initial remote sensing template image P is projected and transformed to obtain the image R i A remote sensing template image P2 whose viewing angle, size, and spatial resolution are consistent within a preset error range; wherein i is the frame sequence number of the infrared front-down real-time image taken by the aircraft; S2, using heterogeneous adaptive neural network to extract the remote sensing template image P2 and the real-time image R i The features of the remote sensing template map P2 and the real-time map R are calculated. i Based on the coordinates of the feature point pairs, the remote sensing template map P2 to the real-time map R i The coordinate transformation matrix is used to transform the target coordinates in the remote sensing template map P2 to the real-time map R i In the real-time graph R i The target box in S3, let i=i+1, repeat step S2, and get the real-time graph R i+1 The target box in S4, real-time image R of two adjacent frames i and R i+1 Make the association. When the association is successful, R i and R i+1 The mean of the target box in R is used as the rough matching target box, and when the target is in R i+1 When the area in is lower than a preset value, the coarse matching target frame is used as the current matching result, and target matching of the next frame is performed.

2. The heterogeneous image target matching method for a forward-looking, downward-looking, fixed target according to claim 1, characterized in that: When the target is in R i+1 When the area in exceeds a preset value, further comprising performing fine matching based on the coarse matching result, specifically comprising: In real-time graph R i+1 In the first step of fine matching, the image block enlarged by a preset multiple of the coarse matching target frame is intercepted as the current real-time image to be matched; when performing fine matching for the first time, the template image used is the image block containing the target position intercepted from the remote sensing template image P2; when performing fine matching subsequently, the target frame obtained by fine matching in the real-time image of the previous fine matching is intercepted as the template image; A convolutional neural network is used to perform fine matching on the current real-time image to be matched and the template image to obtain a fine matching target frame, and the fine matching target frame is used as the current matching result, and target matching of the next frame is performed.

3. The heterogeneous image target matching method for a forward-looking, downward-looking, fixed target according to claim 2, characterized in that: After performing the fine matching on multiple frames of real-time images, the method further includes: Construct a fine matching template library, which includes templates t1, t2, and t3; among them, template t1 is an image block containing the target position intercepted from the remote sensing template map P2; template t3 is used in the fine matching process and is the first Euclidean distance d between it and template t1 k Satisfy d k The template map with <Q1 relationship, where Q1 is a set threshold; template t2 is used in the fine matching process and is the second Euclidean distance d between it and template t1 k Satisfy d k The template map with <Q1 relationship; The template confidence level is determined based on the current fine matching result, including: Calculate the Euclidean distances d1, d2, and d3 between the template image used in the current fine matching process and the templates t1, t2, and t3 in the template library respectively, and obtain the average Euclidean distance If d < Q2, where Q2 is a set threshold, it is regarded as a successful fine match; otherwise, it is regarded as a failed fine match, and the target matching for the next frame is performed; When the fine matching is successful, the Euclidean distances d4 and d5 between the template t1 and the template t2 and the template t3 are calculated respectively. If (d1+d2)<(d4+d5), the template library is updated: the template t2 is used as the current template t3, the template image used in the current fine matching process is used as the current template t2, and the target matching of the next frame is performed; otherwise, the template library is not updated.

4. The heterogeneous image target matching method for a forward-looking, downward-looking, fixed target according to any one of claims 1 to 3, characterized in that: In S1, the inertial navigation information includes the aircraft's height relative to the ground H, the aircraft's latitude lat0 and longitude lon0, the aircraft's pitch angle anglex, roll angle angley and yaw angle anglez; the camera parameters include the vertical viewing angle FOV zhi and horizontal field of view (FOV) ping ; S1 includes: Calculate the distance between the aircraft and the real-time map R in the aircraft coordinate system i The horizontal distance r between the four boundary points x and projection distance r y : Calculate the aircraft to the real-time map R i The distance r and angle θ between the four boundary points: Calculate the initial remote sensing template map P and the real-time map R i The latitudes of the four boundary points corresponding to the four boundary points j and longitude lon j : Wherein, j=1, 2, 3, 4, respectively represent the four boundary points corresponding to the distance r; Based on the image resolution res and longitude and latitude of the initial remote sensing template map P, the image resolution res and longitude and latitude of the initial remote sensing template map P are obtained. i The coordinates of the four boundary points corresponding to the four boundary points (Δx j ,Δy j ); and according to the coordinates (Δx j ,Δy j ) performing cropping on the initial remote sensing template image P to obtain a cropped remote sensing template image P1; Use the real-time graph R i The coordinates of the four boundary points (x j ,y j ) and the coordinates (Δx j ,Δy j ) to transform the remote sensing template image P1 into the remote sensing template image P2.

5. The heterogeneous image target matching method for a front-downward fixed target according to claim 4 is characterized in that: In S2, the method for calculating the coordinates of the feature points includes: The remote sensing template image P2 and the real-time image R i The features of the remote sensing template map P2 and the real-time map R are respectively obtained. i Coordinates of key feature points in ; Based on the coordinates of the key feature points, the remote sensing template image P2 and the real-time image R are searched using a fast nearest neighbor search algorithm. i Perform rough matching on the key feature points in the image to obtain the coordinates of the feature point pairs after rough matching; The dynamic adaptive distance constraint algorithm is used to dynamically and adaptively eliminate the feature point pairs after rough matching. The elimination conditions are: dis j ≥dis′ j -avgdis; Among them, dis j 、dis′ j are the nearest neighbor and the next nearest neighbor of the jth pair of feature points, respectively, and N is the number of feature point pairs after the rough matching; After removing the feature point pairs, the remaining feature point pairs are accurately matched using a random sampling consensus algorithm to obtain the precise feature point pair coordinates, which are used as the coordinates of the required feature point pairs.

6. The heterogeneous image target matching method for a forward-looking, downward-looking, fixed target according to claim 1, characterized in that: In S2, the heterogeneous adaptive neural network is a VGG16 network, and the last convolutional layer is discarded, and the output of the fourth convolutional layer of the VGG16 network is directly selected as the remote sensing template image P2 and the real-time image R i characteristics.

7. The heterogeneous image target matching method for a forward-looking, downward-looking, fixed target according to claim 1, characterized in that: In S4, two adjacent frames of real-time image R i and R i+1 When the length ratio and width ratio of are within the preset threshold range, the IOU of the two frames of real-time images is calculated: Among them, A and B are the target frames of two frames of real-time images respectively; When IOU>Y, where Y is the set threshold, the two frames of real-time images are successfully associated.

8. A heterogeneous image target matching system for forward-looking, downward-looking, fixed targets, characterized by: comprising a computer-readable storage medium and a processor; The computer-readable storage medium is used to store executable instructions; The processor is used to read the executable instructions stored in the computer-readable storage medium to execute the heterogeneous image target matching method for forward-looking fixed targets according to any one of claims 1 to 7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the heterogeneous image target matching method for a forward-looking, downward-looking, fixed target as described in any one of claims 1 to 7 is implemented.

10. A computer program product, characterized in that The invention comprises a computer program, which, when running on a computer, enables the computer to execute the heterogeneous image target matching method for a front-downward fixed target according to any one of claims 1 to 7.