A multi-modal real-time target localization and re-indication method
By installing a multi-spectral zoom camera on the drone, dynamically adjusting the focal length and generating a multi-layer positioning reference template, the problem of slow target re-indication speed, easy loss, and confusing targets in vision target detection and positioning is solved, and the accuracy and robustness of fast target re-indication and recognition of targets is achieved.
Patent Information
- Application Number
- CN202411267961.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-11
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2044-09-11
AI Technical Summary
In the prior art, in the detection and positioning of vision targets, the target is slow, easy to lose, and confusing the target.
Using the multimodal real-time target positioning and re-indication method, the multi-spectral zoom camera equipped with a drone is used to acquire images in real time, dynamically adjust the focal length, generate a multi-layer positioning reference template, and quickly indicate the target again.
It improves the accuracy and robustness of target recognition, achieves the fast target re-indication function, fast response speed and strong stability, and is suitable for fire monitoring, intelligent transportation, autonomous driving and other fields.
Smart Images

Figure CN119131633B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of long-range target detection and positioning, and relates to a multi-modal real-time target positioning and re-indication method. Background Art
[0002] At present, the zoom multi-spectral camera carried by an unmanned aerial vehicle (UAV) has become one of the mainstream technologies for target positioning and recognition. In the field of long-range target detection and positioning, during the search process for an indefinite number of targets, it is often necessary to quickly re-position the targets that have been searched in the subsequent search process, that is, to quickly re-indicate the targets. Target re-indication is the re-positioning of the same target during the target positioning search process, which is a common operation in multi-target recognition and positioning. However, the current target re-indication often involves re-searching for the target, resulting in problems such as slow target re-indication speed, easy loss, and confusion of targets. Summary of the Invention
[0003] The technical problem to be solved by the present invention is to provide a multi-modal real-time target positioning and re-indication method, which can complete the function of rapid target re-indication and improve the accuracy and robustness of target recognition.
[0004] To solve the above technical problem, the multi-modal real-time target positioning and re-indication method of the present invention is as follows:
[0005] Step 1: Perform target recognition and positioning on the original images collected in real time by the multi-spectral zoom camera carried by the UAV, assign an ID value to the recognized target, and record its GPS position;
[0006] Step 2: Dynamically adjust the focal length of the multi-spectral zoom camera, and use the images of the target area captured within each segmented adjustable focal length range as the positioning reference templates for the corresponding focal length levels, so as to generate multi-layer positioning reference templates containing focal length information;
[0007] Step 3: When receiving a target re-indication request, first obtain the GPS position and multi-layer positioning reference templates corresponding to the re-indicated target; adjust the viewing position of the multi-spectral zoom camera so that the GPS value corresponding to the central pixel of the image is the GPS position of the re-indicated target, and use the image at this time as a marked image; if the marked image matches all the positioning reference templates of the participating focal length levels, the target re-indication is successful.
[0008] In the second step, first adjust the attitude of the gimbal so that the target in the image is at or near the center position, and then capture the images of the target area within each segmented adjustable focal length range as the positioning reference templates for the corresponding focal length levels.
[0009] In the second step, the image of the target area is centered on the target coverage box and enlarged by a factor of 1.2 - 1.5 times.
[0010] In the first step, when multiple targets are recognized, one of the targets is located first, an ID value is assigned to it, and its GPS position is recorded. Then, a multi-layer positioning reference template corresponding to the target is generated through the second step. Similarly, the other recognized targets are located in turn, an ID value is assigned to each of them, its GPS position is recorded, and a corresponding multi-layer positioning reference template is generated.
[0011] In the third step, when a target re-indication request is received, if the re-indication target ID value is received at the same time, the corresponding GPS information and multi-layer positioning reference template are obtained according to the re-indication target ID value.
[0012] In the third step, when a target re-indication request is received, if the re-indication target ID value is not received, the targets corresponding to each ID value are re-indicated in the order from the latest to the earliest in time.
[0013] In the third step, the method for matching the marked image with the positioning reference template is as follows:
[0014] 1. Adjust the perspective position of the multi-spectral zoom camera so that the GPS value corresponding to the central pixel of the image is the GPS position of the re-indicated target, and use the image at this time as the marked image.
[0015] 2. Select the positioning reference template at the focal length level corresponding to the re-indicated target according to the current focal length of the multi-spectral zoom camera and match it with the marked image; if the matching fails, adjust the focal length of the multi-spectral zoom camera to a smaller focal length level, return to step 1, and if the matching is successful, move the registration position in the marked image to the center of the image and go to step 3.
[0016] 3. Adjust the focal length of the multi-spectral zoom camera to a larger focal length level, obtain a new marked image. If the marked image fails to match the positioning reference template at the corresponding focal length level, adjust the focal length of the multi-spectral zoom camera to the previous focal length level and return to step 1; if the matching is successful, adjust the focal length of the multi-spectral zoom camera to an even larger focal length level, obtain a new marked image, and then match the marked image with the positioning reference template at the corresponding focal length level; if the marked image matches the positioning reference templates at all the focal length levels participating in the matching, the target re-indication is successful.
[0017] Furthermore, if the number of times the marked image fails to match the positioning reference template exceeds the set maximum number of matches, the target re-indication fails.
[0018] The marked image and the positioning reference template are matched by using a feature matching algorithm, a histogram matching algorithm, a template-based matching algorithm, a deep learning-based matching algorithm, or a coefficient normalization algorithm.
[0019] Advantageous effects: The multi-modal real-time target localization and re-indication method of the present invention combines various sensor data (RGB images, infrared images, inertial navigation, positioning modules, etc.), and uses technologies such as artificial intelligence and traditional target matching for target recognition and localization, improving the accuracy and robustness of target recognition. Thus, while continuing to search for other targets after identifying a certain target, the camera image can be quickly switched back to this target at any time to complete the function of rapid target re-indication. At the same time, in order to more effectively respond to multiple targets for re-indication, ID tags are assigned to each identified target, enabling the function of multi-target re-indication. The present invention has the advantages of fast response speed and strong stability in target re-indication, and can be widely applied in fields such as fire monitoring, intelligent transportation, and autonomous driving. Brief Description of the Drawings
[0020] Figure 1 is the overall flowchart of the multi-modal real-time target localization and re-indication method of the present invention.
[0021] Figure 2 are images of target A and target B being recognized.
[0022] Figure 3 , Figure 4 , Figure 5 is the three-layer localization reference template corresponding to target A.
[0023] Figure 6 are images of target B and target C being recognized.
[0024] Figure 7 , Figure 8 , Figure 9 is the three-layer localization reference template corresponding to target B.
[0025] Figure 10 is the image of target C being recognized.
[0026] Figure 11 , Figure 12 , Figure 13 is the three-layer localization reference template corresponding to target C. Detailed Embodiments
[0027] The present invention will be further described in detail below in conjunction with the drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention.
[0028] In the description of the embodiments, the orientation or positional relationships such as "upper", "lower", "left", "right", etc. are based on the orientation or positional relationships shown in the drawings. They are only for convenience of description and simplifying operations, rather than indicating or implying that the devices or elements referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present invention. In addition, the terms "first" and "second" are only used for distinction in description and have no special meaning.
[0029] Fast and accurate target positioning is the key to achieving intelligent perception. Traditional single positioning methods, such as GPS, visual odometry, etc., are greatly affected by environmental factors and have low positioning accuracy. To solve this problem, this method proposes a target re-indication method based on a multi-layer positioning template, which can comprehensively utilize multi-dimensional templates to quickly re-position the target.
[0030] As Figure 1 shown, the multi-modal real-time target positioning and re-indication method of the present invention specifically includes the following steps:
[0031] Step 1: Use the multi-spectral zoom camera carried by the drone to perform real-time image acquisition. After preprocessing the acquired original images, such as filtering, calibration, alignment, etc., use general computer vision techniques for detection and feature extraction, identify the target and locate it, and determine the specific positions and sizes of each target; as Figure 2 shown, the target detected in the image may be one or multiple; first perform a manual indication or automatic annotation operation on target A, assign it an ID value id1, and record its GPS position; the multi-spectral zoom camera sets multiple adjustable focal length ranges according to the imaging effect; perform image data acquisition on the area where target A is located. When acquiring, adjust the attitude of the pan-tilt head so that target A in the image is at the center position or close to the center position, dynamically adjust the exposure time, aperture, focal length, etc. of the multi-spectral zoom camera, and segment the adjustable focal length range of the camera using the segmented ranges set by the software as the acquisition levels of the image. Take the images of the area where target A is located taken within each segmented adjustable focal length range as the positioning reference templates corresponding to the focal length levels, and thus generate a multi-layer positioning reference template containing focal length information, as Figure 3 、 Figure 4 、 Figure 5 shown, the image 1 of the area where target A is located is centered on the target coverage frame 2 and enlarged by a factor of 1.2 - 1.5 in equal proportion;
[0032] Step 2: Use the multi-spectral zoom camera to perform real-time image acquisition again, and locate the identified target B, as Figure 6 shown, assign it an ID value id2, record its GPS position, and use the same method as generating the multi-layer positioning reference template corresponding to target A to obtain the multi-layer positioning reference template corresponding to target B, as Figure 7 、Figure 8 , Figure 9 as shown; similarly, locate the identified target C, such as Figure 10 as shown, assign it an ID value of id3, record its GPS position, and generate a corresponding multi-layer positioning reference template with focal length information, such as Figure 11 , Figure 12 , Figure 13 as shown;
[0033] Step 3: During the search process, when a target re-indication request is received, if a re-indication target ID value is received simultaneously, obtain the corresponding GPS position and multi-layer positioning reference template according to the re-indication target ID value; if there is no re-indication target ID value, re-indicate the targets corresponding to id3, id2, and id1 in the order from the latest to the earliest time; taking the target A corresponding to id1 as an example, the re-indication method is as follows:
[0034] 1. Calculate the GPS position of the current image center pixel based on the focal length, height from the ground, field of view range, GPS, and inertial navigation pose information of the multi-spectral zoom camera, and use the pan-tilt to adjust the viewing angle position of the multi-spectral zoom camera in real time so that the GPS value corresponding to the image center pixel is the GPS position of the re-indication target A, and use the image at this time as the marked image; for the abnormal situation where the pan-tilt cannot adjust the GPS value corresponding to the image center pixel to the GPS position of the re-indication target A, the re-indication fails, directly return, recalculate the GPS position corresponding to the current image center pixel until the GPS value corresponding to the image center pixel is the GPS position of the re-indication target A, and use the image at this time as the marked image;
[0035] 2. Select the positioning reference template corresponding to the focal length level of the re-indication target A according to the current focal length of the multi-spectral zoom camera and match it with the marked image. If the match fails, adjust the focal length of the multi-spectral zoom camera to a smaller focal length level, return to step 1. If the match is successful, go to step 3;
[0036] 3. Adjust the focal length of the multi-spectral zoom camera to a larger focal length level, obtain a new marked image. If the marked image fails to match the positioning reference template corresponding to the focal length level, adjust the focal length of the multi-spectral zoom camera to the previous focal length level, return to step 1; if the match is successful, adjust the focal length of the multi-spectral zoom camera to an even larger focal length level, obtain a new marked image, and then match the marked image with the positioning reference template corresponding to the focal length level; if the marked image matches the positioning reference templates of all participating focal length levels, the target re-indication is successful.
[0037] The system can also set the maximum number of matches. If the number of times the marked image fails to match the positioning reference template exceeds the set maximum number of matches, the target re-indication fails.
[0038] Taking the positioning reference template with a total of 5 layers, the current focal length corresponding to the 3rd focal length level, and the maximum number of failed matching attempts being 4 as an example, the re-indication method is as follows:
[0039] 1. Adjust the perspective position of the multispectral zoom camera so that the GPS value corresponding to the central pixel of the image is the GPS position of the re-indication target, and use the image at this time as the marked image.
[0040] 2. If the positioning reference template at the 3rd focal length level fails to match the marked image, adjust the focal length of the multispectral zoom camera to the 2nd focal length level, return to step 1, re-adjust the perspective position of the multispectral zoom camera, obtain a new marked image, and match the positioning reference template at the 2nd focal length level with the marked image; if the positioning reference template at the 3rd focal length level matches the marked image successfully, move the registration position in the marked image to the center of the image, and go to step 3.
[0041] 3. Adjust the focal length of the multispectral zoom camera to the 4th focal length level, obtain a new marked image, and then match the marked image with the positioning reference template at the 4th focal length level; if the match fails, adjust the focal length of the multispectral zoom camera to the 3rd focal length level, return to step 1, re-adjust the perspective position of the multispectral zoom camera, obtain a new marked image, and match the positioning reference template at the 3rd focal length level with the marked image; if the marked image matches the positioning reference template at the 4th focal length level successfully, move the registration position in the marked image to the center of the image, adjust the focal length of the multispectral zoom camera to the 5th focal length level, obtain a new marked image, and then match the marked image with the positioning reference template at the 5th focal length level; if the match fails, adjust the focal length of the multispectral zoom camera to the 4th focal length level, return to step 1, re-adjust the perspective position of the multispectral zoom camera, obtain a new marked image, and match the positioning reference template at the 4th focal length level with the marked image; return to step 1; if the marked image matches the positioning reference template at the 5th focal length level successfully, the target re-indication is successful. During the matching process, if the number of failed matching attempts reaches 4 times, return failure.
[0042] The GPS position of the target is obtained through the following method: Use the GPS and inertial navigation system carried by the UAV to obtain the position and attitude information of the UAV, and combine the position of the target in the image to calculate the accurate GPS coordinates of the target; the specific method is as follows:
[0043] (1) Determine the UAV flight attitude parameters: During the flight of the UAV, it will be affected by factors such as wind force and gravity, resulting in changes in its flight attitude. Therefore, first use the inertial measurement unit (IMU) carried by the UAV to measure the attitude angles of the UAV (yaw angle (Yaw), pitch angle (Pitch), roll angle (Roll)) in real time.
[0044] (2) Calculate the exterior orientation elements of the multispectral zoom camera: Based on the flight attitude angles of the UAV and the pre-calibrated interior orientation elements of the multispectral zoom camera, calculate the exterior orientation elements of the multispectral zoom camera in the world coordinate system;
[0045] X = Xc + (Hc / Z) ×(x - xp)
[0046] Y = Yc + (Hc / Z) × (y - yp)
[0047] Z = Hc.
[0048] Among them, (X, Y, Z) are the position coordinates of the pixel on the image in the world coordinate system, Xc, Yc, and Hc are the position coordinates of the center point of the multispectral zoom camera in the world coordinate system; x, y are the position coordinates of the pixel on the image in the image coordinate system; xp, yp are the center point coordinates of the image; Z is the elevation of the object point.
[0049] (3) Establish the conversion relationship between the GPS coordinates of the image pixels: Through the above steps, the position coordinates of each pixel in the world coordinate system can be determined. Further, this position coordinate can be converted into the standard GPS coordinate to obtain the GPS coordinate corresponding to each pixel. The conversion formula is as follows:
[0050] Latitude = arctan[(Z - Hc) / sqrt(X^2 + Y^2)]
[0051] Longitude = arctan(Y / X).
[0052] Among them, Latitude is the latitude and Longitude is the longitude.
[0053] The positioning reference template and the marked image can be matched using existing feature matching algorithms, histogram matching algorithms, template-based matching algorithms, deep learning-based matching algorithms, or coefficient normalization algorithms (COEFF NORMED), etc.
[0054] Feature matching algorithm: Image matching is achieved by detecting feature points in the image and calculating feature descriptors. Commonly used feature matching algorithms include SIFT (Scale-Invariant Feature Transform), SURF (Speeded-Up Robust Features), ORB (Oriented FAST and Rotated BRIEF), etc.
[0055] Histogram matching algorithm: Image matching is achieved by calculating the histogram features of the image. Commonly used algorithms include histogram similarity matching, histogram-based color matching, etc.
[0056] Template matching algorithm: By sliding the image to be matched on the template image and calculating the similarity between the two images to achieve matching. Commonly used template matching algorithms include SSD (Sum of Squared Differences), NCC (Normalized Cross Correlation), etc.
[0057] Deep learning-based matching algorithm: Using methods such as convolutional neural network (CNN) and recurrent neural network (RNN) in deep learning technology to extract and match image features.
[0058] Coefficient normalization algorithm: Matching the relative value of the template image to its mean with the relative value of the image to be matched to its mean. 1 represents a perfect match, -1 represents a poor match, and 0 represents no correlation.
[0059] The present invention uses a multi-sensor fusion target optical detection method, which can improve the accuracy and robustness of detection. For target recognition and positioning, a multi-spectral zoom camera is used to obtain multi-spectral image data of the target, and the collected data is processed such as denoising, correction, alignment, etc., which can improve the quality of the image data. Using a dynamic target recognition method, the pan-tilt and multi-spectral zoom camera are controlled in real time. The image sensor has high sensitivity and a wide dynamic range, and can obtain clear images in complex environments; preprocessing such as filtering, correction, alignment, etc. is performed on the collected original images, which can eliminate noise and distortion and prepare for subsequent feature extraction.
Claims
1. A multi-modal real-time target positioning and re-indication method, characterized in that The method is as follows: S1. Perform target recognition and positioning on the original images collected in real time by the multi-spectral zoom camera carried by the drone, assign an ID value to the recognized target and record its GPS location; S2, dynamically adjust the focal length of the multi-spectral zoom camera, and use the target area image captured within the adjustable focal length range of each segment as the positioning reference template of the corresponding focal length level, so as to generate a multi-layer positioning reference template containing focal length information; S3. When receiving a target re-instruction request, first obtain the GPS position and multi-layer positioning reference template corresponding to the re-instruction target; If the marked image successfully matches the positioning reference templates of all focal length levels involved in the matching, the target re-indication is successful; the marked image is an image when the GPS value corresponding to the central pixel of the image is the GPS position of the re-indicating target; In step S3, the method for matching the marked image with the positioning reference template is as follows: S31 adjusts the multi-spectral zoom camera viewing angle position so that the GPS value corresponding to the center pixel of the image is then indicative of the target GPS position, and the image at this time is used as a marked image; S32. According to the current focal length of the multi-spectral zoom camera, the positioning reference template corresponding to the focal length level of the indicated target is selected and matched with the marked image; if the match fails, the focal length of the multi-spectral zoom camera is adjusted to a smaller focal length level, and the process returns to step 1. If the match succeeds, the registration position in the marked image is moved to the center of the image and the process goes to step 3; S33. Adjust the focal length of the multi-spectral zoom camera to a larger focal length level, obtain a new marked image, if the marked image fails to match the positioning reference template of the corresponding focal length level, adjust the focal length of the multi-spectral zoom camera to the previous focal length level, and return to step 1; if the match is successful, adjust the focal length of the multi-spectral zoom camera to a larger focal length level, obtain a new marked image, and then match the marked image with the positioning reference template of the corresponding focal length level; if the marked image successfully matches the positioning reference templates of all focal length levels involved in the matching, the target is indicated successfully again.
2. The method for real-time target positioning and re-indication based on multi-modal according to claim 1, characterized in that: In the step 2, the gimbal posture is first adjusted so that the target in the image is at or close to the center, and then the image of the target area is captured within the adjustable focal length range of each segment as a positioning reference template for the corresponding focal length level.
3. The method for real-time target positioning and re-indication based on multi-modal according to claim 2, characterized in that: In the step 2, the image of the target area is centered on the target coverage frame and is proportionally enlarged by 1.2-1.5 times.
4. The method for real-time target positioning and re-indication based on multi-modal according to claim 1, characterized in that: In the step one, multiple targets are identified, one of which is located first, an ID value is assigned to it, and its GPS position is recorded, and then a multi-layer positioning reference template corresponding to the target is generated through step two; similarly, other identified targets are located in turn, ID values are assigned to them, their GPS positions are recorded, and corresponding multi-layer positioning reference templates are generated.
5. The method for real-time target positioning and re-indication based on multi-modal according to claim 1, characterized in that: In the step three, when a target re-indication request is received, if a re-indication target ID value is received at the same time, the corresponding GPS information and multi-layer positioning reference template are obtained according to the re-indication target ID value.
6. The method for real-time target positioning and re-indication based on multi-modal according to claim 1, characterized in that: In the step 3, when a target re-indication request is received, if a re-indication target ID value is not received, the targets corresponding to the ID values are re-indicated in order from the latest to the earliest in time.
7. The method for real-time target positioning and re-indication based on multi-modal according to claim 1, characterized in that: If the number of failed matches between the marker image and the positioning reference template exceeds the set maximum number of matches, the target will then indicate failure.
8. The method for real-time target positioning and re-indication based on multi-modal according to claim 1, characterized in that The marked image and the positioning reference template are matched using a feature matching algorithm, a histogram matching algorithm, a template-based matching algorithm, a deep learning-based matching algorithm or a coefficient normalization algorithm.
Citation Information
Patent Citations
Small-sized ground marker capturing and positioning method
CN101509782A