An unmanned aerial vehicle vision navigation and positioning method, system and related device

By performing joint correction and depth feature matching on UAV images, and combining benchmark databases and UAV attitude data, the problem of insufficient positioning accuracy of traditional visual navigation systems in low-altitude flight and undulating terrain is solved, and high-precision UAV visual navigation is achieved.

CN120043532BActive Publication Date: 2025-11-11CHENGDU SHENGWEI POWER TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510290759.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-11-11
Estimated Expiration
2045-03-12

AI Technical Summary

Technical Problem

Traditional visual navigation systems lack sufficient positioning accuracy in low-altitude flight and undulating terrain, failing to meet the precise positioning requirements of UAVs.

Method used

By acquiring monocular images in real time, performing joint correction and deep feature matching, and combining benchmark databases and UAV attitude data, we can achieve rapid matching of high-dimensional semantic features and cross-modal comparison to determine the target region and improve positioning accuracy.

Benefits of technology

It improves the positioning accuracy of UAV visual navigation, reduces error accumulation, and enhances navigation reliability in the event of GNSS interruption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120043532B_ABST
    Figure CN120043532B_ABST
Patent Text Reader

Abstract

This invention relates to the field of unmanned aerial vehicle (UAV) navigation and control technology, and discloses a UAV visual navigation and positioning method, system, and related equipment. The method includes: real-time image acquisition; joint correction of the current image to determine a corrected image; acquisition of high-dimensional semantic features in the corrected image, and rapid coarse matching of candidate regions from a benchmark database; cross-modal comparison of the depth local features of the corrected image with the depth local features of each candidate region to determine the target region; determination of the current positioning information of the UAV based on the target region and the UAV's attitude data, and visual navigation based on the current positioning information. This application narrows the matching range through joint correction and rapid coarse matching, and then performs a second matching on the narrowed range through cross-modal matching, thereby improving the positioning accuracy. Finally, the matching target region is combined with the UAV flight data (attitude data) to further improve the positioning accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of unmanned aerial vehicle (UAV) navigation and control technology, and in particular to a UAV visual navigation and positioning method, system and related equipment. Background Technology

[0002] The Global Navigation Satellite System (GNSS) satellite signal has a ground power of only -130 dBm and operates in a publicly available frequency band, making it highly susceptible to interference. Once interfered with, a drone may be unable to complete its intended mission or even return safely. Although dead reckoning can be performed using data from the Inertial Measurement Unit (IMU) when GNSS signals are interfered with, the accumulation of errors over time becomes significant. Therefore, obtaining the drone's precise location (i.e., accurate latitude and longitude coordinates) is crucial for subsequent missions when GNSS is interrupted or unavailable.

[0003] Visual-based navigation (VBN) has attracted widespread attention due to its numerous advantages, including strong anti-interference capabilities, low power consumption, low cost, small size, simple device structure, passive positioning, and high positioning accuracy. Meanwhile, with the continuous development of remote sensing and mapping technologies, high-resolution orthophoto satellite or mapping images now cover almost all areas of the Earth, with each pixel labeled with precise coordinates. Based on this, matching ground-level aerial views captured by UAVs with corresponding remote sensing or mapping images enables accurate and rapid UAV positioning without cumulative errors, serving as an important supplement to current UAV integrated navigation systems.

[0004] However, traditional visual navigation systems currently suffer from insufficient positioning accuracy. For example, traditional scene-matching navigation is less effective when faced with changes in perspective caused by low-altitude flight or undulating terrain. Summary of the Invention

[0005] To overcome the problem of insufficient positioning accuracy in traditional visual navigation systems, this invention provides a visual navigation and positioning method, system, and related equipment for unmanned aerial vehicles (UAVs).

[0006] In a first aspect, to solve the above-mentioned technical problems, the present invention provides a visual navigation and positioning method for unmanned aerial vehicles (UAVs), comprising:

[0007] Real-time image acquisition; wherein the image is a monocular image captured by a monospectral camera or a monocular image captured by a multispectral camera;

[0008] Perform joint correction on the current image to determine the corrected image;

[0009] High-dimensional semantic features in the corrected image are obtained, and at least one candidate region that is similar to the high-dimensional semantic features is quickly coarsely matched from the benchmark database.

[0010] The depth local features of the corrected image are obtained, and the depth local features of the corrected image are compared with the depth local features of each candidate region across modalities to determine the target region.

[0011] Based on the target area and the drone's attitude data, the drone's current location information is determined, and visual navigation is performed based on the current location information.

[0012] Secondly, the present invention provides a visual navigation and positioning system for unmanned aerial vehicles (UAVs), comprising:

[0013] The image acquisition module is used to acquire images in real time; the images are monocular images taken by a monospectral camera or monocular images taken by a multispectral camera.

[0014] The image correction module is used to perform joint correction on the current image and determine the corrected image;

[0015] The candidate region acquisition module is used to acquire high-dimensional semantic features in the corrected image and quickly coarsely match at least one candidate region that is similar to the high-dimensional semantic features from the benchmark database.

[0016] The target region acquisition module is used to acquire the depth local features of the corrected image, and to perform cross-modal comparison between the depth local features of the corrected image and the depth local features of each candidate region to determine the target region.

[0017] The positioning and navigation module is used to determine the current positioning information of the UAV based on the target area and the UAV's attitude data, and to perform visual navigation based on the current positioning information.

[0018] Thirdly, the present invention provides a computing device, including a memory, a processor, and a program stored in the memory and running on the processor, wherein the processor executes the program to implement the steps of the above-described UAV visual navigation and positioning method.

[0019] Fourthly, the present invention provides a computer-readable storage medium storing instructions that, when executed on a terminal device, cause the terminal device to perform the steps of the above-described UAV visual navigation and positioning method.

[0020] The beneficial effects of this invention are as follows: Joint correction is performed on the acquired images to obtain a corrected image. Then, high-dimensional semantic features from the corrected image are extracted, and coarse matching is performed from a benchmark database to obtain multiple candidate regions. Finally, a target region is obtained through cross-domain matching using a graph neural network. The current positioning information can then be obtained using the target region and the UAV's attitude data, enabling visual navigation of the UAV. This application narrows the matching range through joint correction and fast coarse matching, and then performs a secondary matching on the narrowed range through cross-modal matching, thereby improving positioning accuracy. Finally, the matched target region is combined with the UAV's flight data (attitude data) to further improve positioning accuracy. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0022] Figure 1 This is a flowchart illustrating a visual navigation and positioning method for unmanned aerial vehicles (UAVs) according to an embodiment of the present invention.

[0023] Figure 2 This is a schematic diagram of the structure of a UAV visual navigation and positioning system according to an embodiment of the present invention. Detailed Implementation

[0024] The following embodiments are further explanations and supplements to the present invention and do not constitute any limitation on the present invention.

[0025] The following describes, with reference to the accompanying drawings, an embodiment of the present invention of a UAV visual navigation and positioning method, system and related equipment.

[0026] like Figure 1 As shown, this embodiment of the invention provides a visual navigation and positioning method for unmanned aerial vehicles (UAVs), including:

[0027] S1. Real-time image acquisition; wherein the image is a monocular image captured by a single-spectral camera or a monocular image captured by a multispectral camera.

[0028] In this embodiment, a global shutter camera is selected as the single / multispectral camera, and HDR mode is supported to cope with changes in lighting.

[0029] S2. Perform joint correction on the current image to determine the corrected image.

[0030] S3. Obtain high-dimensional semantic features in the corrected image and quickly coarsely match at least one candidate region that is similar to the high-dimensional semantic features from the benchmark database.

[0031] In this embodiment, the baseline database generates a multi-scale feature pyramid offline, which is combined with hash coding to accelerate retrieval. The layered geographic baseline data, i.e., satellite images or survey images, is stored, and then the layered geographic baseline data is sliced ​​into multiple candidate regions.

[0032] S4. Obtain the depth local features of the corrected image, and perform cross-modal comparison between the depth local features of the corrected image and the depth local features of each candidate region to determine the target region.

[0033] S5. Based on the target area and the attitude data of the UAV, determine the current positioning information of the UAV, and perform visual navigation based on the current positioning information.

[0034] In this embodiment, the attitude data of the UAV is provided by an IMU and a barometer. The IMU provides short-term attitude prediction, and the barometer assists in altitude estimation.

[0035] In this embodiment, the acquired images are jointly corrected to obtain corrected images. High-dimensional semantic features from these corrected images are then extracted, and coarse matching is performed from a benchmark database to obtain multiple candidate regions. Finally, a cross-domain matching method using a graph neural network is used to obtain the target region. The current positioning information can then be obtained using the target region and the UAV's attitude data, enabling visual navigation of the UAV. This application narrows the matching range through joint correction and fast coarse matching, and then performs a secondary matching on the narrowed range through cross-modal matching, thereby improving positioning accuracy. Finally, the matched target region is combined with the UAV's flight data (attitude data) to further improve positioning accuracy.

[0036] In this embodiment, the hardware on the UAV can adopt the Jetson Orin NX embedded platform, which supports 100 TOPS computing power. TensorRT quantization models (FP16 / INT8) can be deployed on the inference engine. Therefore, various deep learning neural networks can be embedded in the platform to execute a UAV visual navigation and positioning method as described above, thereby improving computational efficiency. At the same time, the algorithm also supports acceleration solutions such as GPU, CPU and NPU.

[0037] Optionally, joint correction is performed on the current image to determine the corrected image, including:

[0038] The current image is preprocessed using adaptive histogram equalization and dark channel dehazing algorithms to determine the preprocessed image;

[0039] The three-axis acceleration of the IMU on the UAV is obtained, and the motion blur of the preprocessed image is determined based on the three-axis acceleration and the average acceleration within the exposure time window.

[0040] If the degree of motion blur meets the preset requirements, the preprocessed image will be used as the correction image;

[0041] If the degree of motion blur does not meet the preset requirements, the current rotation angle of the UAV is obtained through the IMU, and motion blur compensation is performed on the preprocessed image based on the current rotation angle to determine the corrected image.

[0042] In this embodiment, adaptive histogram equalization aims to improve image contrast, particularly in local regions. It is achieved by applying histogram equalization to local areas of the image (typically small windows). This process calculates and equalizes the local histogram for each small window, thereby enhancing local contrast and making image details more apparent. To avoid over-enhancement, smoothing is typically performed, resulting in CLAHE (Contrast Limited AHE) adaptive histogram equalization.

[0043] This embodiment presents a dark channel dehazing algorithm, an image processing method for dehazing, based on a simple but effective assumption: in hazy images, certain color channels (such as red, green, and blue) in most image regions will have very small values; this value is called the dark channel. Even in foggy conditions, the value of this dark channel is usually small. The core idea of ​​dark channel dehazing is to recover a fog-free image by estimating the atmospheric light and transmission map (i.e., the fog density distribution) in the image. It proceeds through the following steps: 1. Calculate the dark channel image. 2. Estimate atmospheric light based on the minimum value of the dark channel. 3. Use this estimation result to infer the transmission map of each pixel in the image. 4. Recover the original fog-free image based on the atmospheric light and transmission map.

[0044] In this embodiment, the motion blur level meets the preset requirement, which means that the motion blur level is less than the threshold.

[0045] Optionally, the degree of motion blur in the preprocessed image is determined based on the triaxial acceleration and the average acceleration within the exposure time window, as shown in the following formula:

[0046]

[0047] Where K represents the degree of motion blur, a x a y a z Represents triaxial acceleration, μ α σ represents the mean acceleration within the exposure time window, N represents the length of the exposure time window, and σ represents the mean acceleration within the exposure time window. motion The function representing the degree of motion blur is σ. noise The standard deviation of Gaussian noise is given.

[0048] In this embodiment, a motion blur evaluation function is constructed to calculate the motion blur level of the current image. The blur level of the current image is quantified, which helps to determine whether motion blur compensation is needed for the current image based on the motion blur level, thereby improving the positioning accuracy of the UAV.

[0049] Optionally, motion blur compensation is performed on the preprocessed image based on the rotation angle to determine the corrected image, including:

[0050] Based on the rotation angle, the frequency domain compensation direction weight matrix is ​​determined as follows:

[0051]

[0052] Where W(u,v) represents the frequency domain compensation direction weight matrix, u and v are frequency domain coordinates, representing two orthogonal direction vectors in the frequency domain, which are the horizontal and vertical frequency domain axes of the image, θ represents the rotation angle, and σ represents the preset value;

[0053] Motion blur compensation is performed on the preprocessed image based on the frequency domain compensation direction weight matrix to determine the corrected image.

[0054] In this embodiment, the degree of motion blur of the UAV depends on the tilt angle of the UAV during flight. Therefore, the current rotation angle of the UAV can be obtained by means of IMU, and the frequency domain compensation direction weight matrix can be calculated by means of the current rotation angle, so as to perform motion blur compensation on the current image and improve the positioning accuracy of the UAV.

[0055] Optionally, high-dimensional semantic features in the corrected image are obtained, and at least one candidate region that approximates the high-dimensional semantic features is quickly and coarsely matched from a benchmark database, including:

[0056] High-dimensional semantic features are extracted from the corrected image using a lightweight convolutional neural network;

[0057] The random hyperplane projection method is used to map high-dimensional semantic features into binary hash codes;

[0058] Using the locality-sensitive hashing algorithm, at least one candidate region that approximates the binary hash code corresponding to the high-dimensional semantic features is quickly coarsely matched from the benchmark database.

[0059] In this embodiment, the lightweight convolutional neural network can be a series of backbone networks such as EfficientNet, MobileNet, or ResNet. The original network's last classification layer is removed through lightweight reconstruction, and generalized average pooling is used to obtain the high-dimensional semantic features of the corrected image.

[0060] In this embodiment, the hash code of each candidate region in the benchmark database is pre-calculated and stored in buckets according to the hash value. At this time, the binary hash code obtained by mapping the high semantic features can be quickly compared with the hash code values ​​corresponding to all candidate regions. The Hamming distance is used to measure the similarity of hash codes, thereby filtering out at least one candidate region that is closest.

[0061] Optionally, the depth local features of the corrected image are compared with the depth local features of each candidate region across modalities to determine the target region, including:

[0062] The local depth features of the corrected image are matched with the local depth features of each candidate region, and feature points with incorrect matching relationships are eliminated to determine the target candidate region.

[0063] Use the candidate region as the target region.

[0064] In this embodiment, deep local features are extracted from the corrected image and candidate regions using techniques such as SIFT (Scale-Invariant Feature Transform), DISK (DIScrete Keypoints, a reinforcement learning-based deep feature detection and matching framework), or SuperPoint (a deep learning algorithm for feature detection and matching).

[0065] In this embodiment, neural network algorithms such as SuperGlue, LightGlue, or GIM-LightGlue are used to query the approximation of the depth local features of the corrected image and the depth local features of the candidate region, and obtain the most approximate candidate region.

[0066] In this embodiment, neural network algorithms such as SuperGlue, LightGlue, or GIM-LightGlue compare the similarity between each feature point in the deep local features and each feature point in the candidate region, thereby obtaining feature points with correct matching relationships (feature points with high similarity) and feature points with incorrect matching relationships (feature points with low similarity).

[0067] In this embodiment, algorithms such as RANSAC (Random Sample Consensus), PROSAC (Progressive Sampling Consensus), or GC-RANSAC (Graph-Cut RANSAC Robust Estimator) are used to eliminate mismatches, i.e. feature points with incorrect matching relationships. The homography matrix is ​​then obtained to calculate the perspective transformation, and the image-level pose estimation is completed to obtain the target candidate region.

[0068] Optionally, the current positioning information of the UAV can be determined based on the target area and the UAV's attitude data, using the following formula:

[0069] Lat = Laattl + Cy·(Latbr - Laattl)

[0070] Lon=Lontl+Cx·(Lonbr-Lontl);

[0071] Where Lat and Lon represent latitude coordinates and progress coordinates respectively, and represent the current location information of the UAV, C x and C y These represent the normalized position information of the UAV in the current image coordinate system, Lat tl Lon tl This indicates the latitude and longitude of the top left corner of the target area. br Lon br This indicates the latitude and longitude of the lower right corner of the target area.

[0072] In this embodiment, the normalized position information in the current image coordinate system is transformed into data in the corresponding three-dimensional coordinate system by using the corresponding metadata (target area) in the reference database. The UAV attitude and altitude data (attitude data) obtained by the IMU are used to limit the transformation result, thereby obtaining high-precision positioning information.

[0073] like Figure 2 As shown, the present invention provides a visual navigation and positioning system for unmanned aerial vehicles (UAVs), comprising:

[0074] The image acquisition module is used to acquire images in real time; the images are monocular images taken by a monospectral camera or monocular images taken by a multispectral camera.

[0075] The image correction module is used to perform joint correction on the current image and determine the corrected image;

[0076] The candidate region acquisition module is used to acquire high-dimensional semantic features in the corrected image and quickly coarsely match at least one candidate region that is similar to the high-dimensional semantic features from the benchmark database.

[0077] The target region acquisition module is used to acquire the depth local features of the corrected image, and to perform cross-modal comparison between the depth local features of the corrected image and the depth local features of each candidate region to determine the target region.

[0078] The positioning and navigation module is used to determine the current positioning information of the UAV based on the target area and the UAV's attitude data, and to perform visual navigation based on the current positioning information.

[0079] Optionally, the image correction module is specifically used for:

[0080] The current image is preprocessed using adaptive histogram equalization and dark channel dehazing algorithms to determine the preprocessed image;

[0081] The three-axis acceleration of the IMU on the UAV is obtained, and the motion blur of the preprocessed image is determined based on the three-axis acceleration and the average acceleration within the exposure time window.

[0082] If the degree of motion blur meets the preset requirements, the preprocessed image will be used as the correction image;

[0083] If the degree of motion blur does not meet the preset requirements, the current rotation angle of the UAV is obtained through the IMU, and motion blur compensation is performed on the preprocessed image based on the current rotation angle to determine the corrected image.

[0084] Optionally, the image correction module is specifically used for:

[0085] The degree of motion blur in the preprocessed image is determined based on the triaxial acceleration and the average acceleration within the exposure time window, using the following formula:

[0086]

[0087] Where K represents the degree of motion blur, a x a y a z Represents triaxial acceleration, μ α σ represents the mean acceleration within the exposure time window, N represents the length of the exposure time window, and σ represents the mean acceleration within the exposure time window. motion The function representing the degree of motion blur is σ. noise The standard deviation of Gaussian noise is given.

[0088] Optionally, the image correction module is specifically used for:

[0089] Based on the rotation angle, the frequency domain compensation direction weight matrix is ​​determined as follows:

[0090]

[0091] Where W(u,v) represents the frequency domain compensation direction weight matrix, u and v are frequency domain coordinates, representing two orthogonal direction vectors in the frequency domain, which are the horizontal and vertical frequency domain axes of the image, θ represents the rotation angle, and σ represents the preset value;

[0092] Motion blur compensation is performed on the preprocessed image based on the frequency domain compensation direction weight matrix to determine the corrected image.

[0093] Optionally, the candidate region acquisition module is specifically used for:

[0094] High-dimensional semantic features are extracted from the corrected image using a lightweight convolutional neural network;

[0095] The random hyperplane projection method is used to map high-dimensional semantic features into binary hash codes;

[0096] Using the locality-sensitive hashing algorithm, at least one candidate region that approximates the binary hash code corresponding to the high-dimensional semantic features is quickly coarsely matched from the benchmark database.

[0097] Optionally, the target area acquisition module is specifically used for:

[0098] The local depth features of the corrected image are matched with the local depth features of each candidate region, and feature points with incorrect matching relationships are eliminated to determine the target candidate region.

[0099] Use the candidate region as the target region.

[0100] Optionally, the positioning and navigation module is specifically used for:

[0101] Based on the target area and the UAV's attitude data, the UAV's current location information is determined using the following formula:

[0102] Lat = Laattl + Cy·(Latbr - Laattl)

[0103] Lon=Lontl+Cx·(Lonbr-Lontl);

[0104] Where Lat and Lon represent latitude coordinates and progress coordinates respectively, and represent the current location information of the UAV, C x and C y These represent the normalized position information of the UAV in the current image coordinate system, Lat tl Lon tl This indicates the latitude and longitude of the top left corner of the target area. br Lon br This indicates the latitude and longitude of the lower right corner of the target area.

[0105] The present invention also provides a computing device, including a memory, a manager, and a program stored in the memory and running on the manager. When the manager executes the program, it implements some or all of the steps of the above-described UAV visual navigation and positioning method.

[0106] The computing device can be a computer, and the corresponding program is computer software. The parameters and steps of the computing device of the present invention can be referred to the parameters and steps in the embodiment of the UAV visual navigation and positioning method above, and will not be repeated here.

[0107] Those skilled in the art will recognize that this invention can be implemented as a system, method, or computer program product. Therefore, this disclosure can be embodied in the following forms: it can be entirely hardware, entirely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software, generally referred to herein as a "circuit," "module," or "system." Furthermore, in some embodiments, the invention can also be implemented as a computer program product contained in one or more computer-readable media, which contains computer-readable program code. Computer-readable storage media can be, for example, but not limited to—electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof.

[0108] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0109] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.

Claims

1. A visual navigation and positioning method for unmanned aerial vehicles (UAVs), characterized in that, include: Real-time image acquisition; wherein the image is a monocular image captured by a single-spectral camera or a monocular image captured by a multispectral camera; Perform joint correction on the current image to determine the corrected image; High-dimensional semantic features in the corrected image are obtained, and at least one candidate region that is similar to the high-dimensional semantic features is quickly coarsely matched from the benchmark database. The depth local features of the corrected image are obtained, and the depth local features of the corrected image are compared with the depth local features of each candidate region across modalities to determine the target region. Based on the target area and the attitude data of the UAV, the current positioning information of the UAV is determined, and visual navigation is performed based on the current positioning information; The step of performing joint correction on the current image to determine the corrected image includes: The current image is preprocessed using adaptive histogram equalization and dark channel dehazing algorithms to determine the preprocessed image; The three-axis acceleration of the IMU on the UAV is obtained, and the motion blur of the preprocessed image is determined based on the three-axis acceleration and the average acceleration within the exposure time window. If the degree of motion blur meets the preset requirements, the preprocessed image is used as the corrected image; If the degree of motion blur does not meet the preset requirements, the current rotation angle of the UAV is obtained through the IMU, and motion blur compensation is performed on the preprocessed image based on the current rotation angle to determine the corrected image; Based on the triaxial acceleration and the average acceleration within the exposure time window, the motion blur degree of the preprocessed image is determined using the following formula: Where K represents the degree of motion blur, a x a y a z Represents triaxial acceleration, μ α σ represents the mean acceleration within the exposure time window, N represents the length of the exposure time window, and σ represents the mean acceleration within the exposure time window. motion The function representing the degree of motion blur is σ. noise The standard deviation of Gaussian noise is given.

2. The method according to claim 1, characterized in that, The step of performing motion blur compensation on the preprocessed image based on the rotation angle to determine the corrected image includes: Based on the rotation angle, the frequency domain compensation direction weight matrix is ​​determined as follows: Where W(u,v) represents the frequency domain compensation direction weight matrix, u and v are frequency domain coordinates, representing two orthogonal direction vectors in the frequency domain, which are the horizontal and vertical frequency domain axes of the image, θ represents the rotation angle, and σ represents the preset value; Motion blur compensation is performed on the preprocessed image based on the frequency domain compensation direction weight matrix to determine the corrected image.

3. The method according to claim 1, characterized in that, The step of acquiring high-dimensional semantic features in the corrected image and quickly coarsely matching at least one candidate region that approximates the high-dimensional semantic features from a benchmark database includes: High-dimensional semantic features are extracted from the corrected image using a lightweight convolutional neural network; The random hyperplane projection method is used to map high-dimensional semantic features into binary hash codes; Using the locality-sensitive hashing algorithm, at least one candidate region that approximates the binary hash code corresponding to the high-dimensional semantic features is quickly coarsely matched from the benchmark database.

4. The method according to claim 1, characterized in that, The step of performing a cross-modal comparison between the depth local features of the corrected image and the depth local features of each candidate region to determine the target region includes: The depth local features of the corrected image are matched with the depth local features of each candidate region, and feature points with incorrect matching relationships are eliminated to determine the target candidate region. The candidate region is taken as the target region.

5. The method according to claim 1, characterized in that, The current positioning information of the UAV is determined based on the target area and the UAV's attitude data, using the following formula: Years=Years tl +C y ·(Years br -Years tl ) Lon=Lon tl +C x ·(Lon br -Lon tl ); Where Lat and Lon represent latitude and longitude coordinates respectively, and represent the current location information of the UAV, C x and C y These represent the normalized position information of the UAV in the current image coordinate system, Lat tl Lon tl This indicates the latitude and longitude of the top left corner of the target area. br Lon br This indicates the latitude and longitude of the lower right corner of the target area.

6. A visual navigation and positioning system for unmanned aerial vehicles (UAVs), characterized in that, The method applied to any one of claims 1 to 5 includes: An image acquisition module is used to acquire images in real time; wherein the images are monocular images captured by a monospectral camera or monocular images captured by a multispectral camera. The image correction module is used to perform joint correction on the current image and determine the corrected image; The candidate region acquisition module is used to acquire high-dimensional semantic features in the corrected image and quickly coarsely match at least one candidate region that is similar to the high-dimensional semantic features from the benchmark database. The target region acquisition module is used to acquire the depth local features of the corrected image, and to perform cross-modal comparison between the depth local features of the corrected image and the depth local features of each candidate region to determine the target region. The positioning and navigation module is used to determine the current positioning information of the UAV based on the target area and the UAV's attitude data, and to perform visual navigation based on the current positioning information.

7. A computing device, comprising a memory, a processor, and a program stored in the memory and running on the processor, characterized in that, When the processor executes the program, it implements the steps of a UAV visual navigation and positioning method as described in any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed on a terminal device, cause the terminal device to perform the steps of a UAV visual navigation and positioning method as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Synchronous positioning and mapping method integrating vision, IMU (Inertial Measurement Unit) and sonar

    CN113744337A

  • Aircraft visual navigation method based on deep learning matching and Kalman filtering

    CN116518981A