Omnidirectional vision positioning system and method and unmanned aerial vehicle

By using an omnidirectional vision positioning system that combines multi-fisheye cameras and inertial measurement sensors, the problem of low positioning accuracy of unmanned aerial vehicles in complex environments has been solved, enabling high-precision and robust autonomous flight.

CN121089720APending Publication Date: 2025-12-09CHINA GENERAL NUCLEAR POWER OPERATION
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202511121306.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2025-12-09

AI Technical Summary

Technical Problem

Existing visual positioning technologies have low positioning accuracy and poor robustness in complex environments, making it difficult to meet the high-precision autonomous flight requirements of unmanned aerial vehicles.

Method used

An omnidirectional vision positioning system, including multiple fisheye cameras and inertial measurement sensors, is adopted. Through distortion correction, feature extraction and feature matching, combined with inertial measurement data, the current pose information of the unmanned aerial vehicle in the spatial coordinate system is determined.

Benefits of technology

It improves the positioning accuracy and robustness of unmanned aerial vehicles in complex environments, meeting the requirements for high-precision autonomous flight.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121089720A_ABST
    Figure CN121089720A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of unmanned aerial vehicles, and provides an omni-directional vision positioning system and method and an unmanned aerial vehicle, the omni-directional vision positioning system is installed on a fuselage of the unmanned aerial vehicle, and the omni-directional vision positioning system comprises an image acquisition device used for acquiring a real-time scene image in an omni-directional range in the flight process of the unmanned aerial vehicle; the inertial measurement sensor is connected with the image acquisition equipment and is used for synchronously acquiring motion state data of the unmanned aerial vehicle; the navigation system is connected with the image acquisition equipment and the inertial measurement sensor and is configured to execute at least one of distortion correction processing and feature extraction processing on the real-time scene image to obtain a processed image; and according to the motion state data and the processed image, determining current pose information of the unmanned aerial vehicle in a space coordinate system constructed by taking a take-off point of the unmanned aerial vehicle as an original point. The positioning precision and robustness of the unmanned aerial vehicle in a complex environment are improved, and the requirement of high-precision autonomous flight is met.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of unmanned aerial vehicles, and more particularly relates to an omnidirectional visual positioning system and method and unmanned aerial vehicle. BACKGROUND

[0002] With the wide application of unmanned aerial vehicles in the fields of industrial inspection, emergency rescue, surveying and mapping, etc., higher requirements are put forward for the autonomous positioning accuracy and stability of unmanned aerial vehicles in complex environments. Although the traditional global satellite navigation system (GNSS) can achieve high-precision positioning in open environments, the positioning accuracy will be greatly reduced or even invalid in scenes such as nuclear power plants, urban canyons, indoor environments, etc., where satellite signals are severely blocked or interfered. Although the positioning scheme based on laser radar has high accuracy, it has problems such as high cost and poor adaptability to complex texture environments.

[0003] The existing visual positioning technology, such as monocular visual positioning system, often lacks depth information, resulting in insufficient positioning accuracy, and is prone to feature mismatching in scenes with varying light and similar textures. Although the binocular visual positioning system can obtain depth information, the field of view is limited, and it is difficult to meet the demand for omnidirectional perception. At the same time, the positioning method relying solely on visual information has the problems of high positioning delay and poor robustness in dynamic motion or rapid light changes, and cannot meet the demand for high-precision autonomous flight of unmanned aerial vehicles. Therefore, there is an urgent need for a positioning scheme that can adapt to complex environments to improve the positioning accuracy and robustness of unmanned aerial vehicles in complex environments. SUMMARY

[0004] The purpose of the embodiments of the present application is to provide an omnidirectional visual positioning method, device, electronic equipment and program product, aiming to solve the technical problems that related visual positioning technologies are difficult to adapt to complex environments and have low positioning accuracy in complex environments.

[0005] To achieve the above-mentioned purpose, according to the first aspect of the present application, an omnidirectional visual positioning system is provided, which is installed on the fuselage of an unmanned aerial vehicle, and the omnidirectional visual positioning system comprises:

[0006] An image acquisition device is configured to acquire real-time scene images in an omnidirectional range during the flight of the unmanned aerial vehicle;

[0007] An inertial measurement sensor is connected to the image acquisition device and configured to synchronously acquire motion state data of the unmanned aerial vehicle;

[0008] A navigation system is connected to the image acquisition device and the inertial measurement sensor and is configured to perform at least one of:

[0009] distortion correction processing and feature extraction processing on the real-time scene images to obtain processed images;

[0010] According to the motion state data and the processed image, current pose information of the unmanned aerial vehicle in a space coordinate system is determined, wherein the space coordinate system is constructed with a takeoff point of the unmanned aerial vehicle as an origin.

[0011] In an optional implementation, the image acquisition device further includes:

[0012] a plurality of fisheye cameras, the plurality of fisheye cameras are distributed on the fuselage in a predetermined shape, an included angle between optical axes of the fisheye cameras is a first angle, and a field of view angle of a single fisheye camera is not less than a second angle;

[0013] a hardware trigger synchronization circuit, configured to drive each fisheye camera to synchronously acquire images, wherein a time difference between the fisheye cameras is less than a predetermined time length.

[0014] In an optional implementation, the navigation system performs distortion correction processing on the real-time scene image, and is configured to perform:

[0015] obtain three-dimensional coordinates of each feature point in the real-time scene image on a normalized sphere;

[0016] map the three-dimensional coordinates to planar coordinates in a cylindrical projection image by using a cylindrical projection model, and compensate for planar coordinates that are not mapped by using a linear interpolation method, wherein the cylindrical projection model calibrates intrinsic matrixes and distortion coefficients of the fisheye cameras of the image acquisition device by using a sensor calibration algorithm.

[0017] In an optional implementation, the navigation system performs feature extraction processing on the real-time scene image, and is configured to perform:

[0018] extract a plurality of feature points of the real-time scene image and generate a descriptor of each feature point based on a deep learning algorithm;

[0019] construct a matching relationship of the plurality of feature points by using a feature matching network based on the descriptor of each feature point.

[0020] In an optional implementation, the navigation system determines current pose information of the unmanned aerial vehicle in a space coordinate system according to the motion state data and the processed image, and is configured to perform:

[0021] obtain an installation transformation relationship between the image acquisition device and the inertial measurement sensor, and an installation position relationship between the plurality of fisheye cameras included in the image acquisition device;

[0022] obtain a matching relationship between the feature points by performing feature extraction processing on the real-time scene image;

[0023] obtain depth information of a common view feature point of the plurality of fisheye cameras according to the installation transformation relationship, the installation position relationship, and the matching relationship between the feature points.

[0024] The bundle adjustment method is adopted to solve the pose change of the current frame scene image relative to the last frame scene image according to the depth information of the feature points in the common view area, with the feature point re-projection error as a constraint condition.

[0025] If the current scene is identified as matching the historical scene in the global map, the current pose information of the unmanned aerial vehicle in the spatial coordinate system is determined based on the global map and the pose change.

[0026] According to a third aspect of the present application, an unmanned aerial vehicle is provided for a global navigation satellite system (GNSS) denial scenario, characterized in that the unmanned aerial vehicle comprises: a fuselage and an omnidirectional visual positioning system according to any one of the preceding aspects; the omnidirectional visual positioning system is installed on the fuselage of the unmanned aerial vehicle, and the omnidirectional visual positioning system comprises:

[0027] an image acquisition device configured to acquire real-time scene images in an omnidirectional range during flight of the unmanned aerial vehicle;

[0028] an inertial measurement sensor connected to the image acquisition device and configured to synchronously acquire motion state data of the unmanned aerial vehicle;

[0029] a navigation system connected to the image acquisition device and the inertial measurement sensor and configured to perform at least one of:

[0030] distortion correction processing and feature extraction processing on the real-time scene images to obtain processed images; and determine current pose information of the unmanned aerial vehicle in a spatial coordinate system based on the motion state data and the processed images, wherein the spatial coordinate system is constructed with a takeoff point of the unmanned aerial vehicle as an origin.

[0031] According to a third aspect of the present application, an omnidirectional visual positioning method is provided, comprising:

[0032] acquiring real-time scene images in an omnidirectional range acquired by an image acquisition device of an unmanned aerial vehicle during flight of the unmanned aerial vehicle, and acquiring motion state data of the unmanned aerial vehicle synchronously acquired by an inertial measurement sensor of the unmanned aerial vehicle;

[0033] performing at least one of distortion correction processing and feature extraction processing on the real-time scene images to obtain processed images;

[0034] determining current pose information of the unmanned aerial vehicle in a spatial coordinate system based on the motion state data and the processed images, wherein the spatial coordinate system is constructed with a takeoff point of the unmanned aerial vehicle as an origin.

[0035] According to a third aspect of the present application, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to enable the electronic device to implement the method according to any one of the first aspect.

[0036] According to a fourth aspect of the present application, a computer readable storage medium is provided, which stores a computer program, wherein the computer program is executable by a processor to implement the method according to any one of the first aspect.

[0037] According to a fifth aspect of the present application, a computer program product is provided, which, when executed on an electronic device, enables the electronic device to implement the method according to any one of the first aspect.

[0038] It can be understood that the beneficial effects of the second aspect to the fifth aspect described above can be referred to the related description of the first aspect, which will not be repeated here.

[0039] The embodiment of the present application provides an omnidirectional visual positioning system, which is installed on a fuselage of an unmanned aerial vehicle. The omnidirectional visual positioning system comprises: an image acquisition device, configured to acquire real-time scene images in an omnidirectional range during flight of the unmanned aerial vehicle; an inertial measurement sensor, connected to the image acquisition device, configured to synchronously acquire motion state data of the unmanned aerial vehicle; and a navigation system, connected to the image acquisition device and the inertial measurement sensor, configured to perform at least one of distortion correction processing and feature extraction processing on the real-time scene images to obtain processed images, and determine current pose information of the unmanned aerial vehicle in a space coordinate system according to the motion state data and the processed images, wherein the space coordinate system is constructed with a takeoff point of the unmanned aerial vehicle as an origin. The positioning accuracy and robustness of the unmanned aerial vehicle in a complex environment are improved, and the demand for high-precision autonomous flight is met. BRIEF DESCRIPTION OF DRAWINGS

[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0041] Figure 1 is a structural schematic diagram of an omnidirectional visual positioning system provided by the embodiment of the present application;

[0042] Figure 2a is a schematic diagram of an original fisheye image provided by the embodiment of the present application;

[0043] Figure 2b is a schematic diagram of an optional corrected image provided by an embodiment of the present application;

[0044] Figure 2c is a schematic diagram of an optional image feature point provided by an embodiment of the present application;

[0045] Figure 3 is a schematic diagram of an optional optimization factor graph of visual-inertial perception information provided by an embodiment of the present application;

[0046] Figure 4 is a schematic diagram of an optional component of Compute Unified Device Architecture (CUDA) provided by an embodiment of the present application;

[0047] Figure 5 is a schematic diagram of an optional architecture of a pose estimation system provided by an embodiment of the present application;

[0048] Figure 6 is a schematic diagram of an optional pose estimation flow provided by an embodiment of the present application;

[0049] Figure 7 is a schematic diagram of an optional flow of an omnidirectional visual positioning method provided by an embodiment of the present application;

[0050] Figure 8 is a schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0051] In the following description, for the purposes of explanation and not limitation, specific details are set forth, such as particular sequences of acts, techniques, etc. in order to provide a thorough understanding of the present embodiments. However, it will be apparent to those skilled in the art that the present embodiments can be practiced in other embodiments that depart from these specific details. In other instances, detailed descriptions of well-known devices, circuits, and methods are omitted so as not to obscure the description of the present embodiments.

[0052] It is to be understood that the terminology "includes", "has", "holds", "contains" and / or other similar forms used in the present specification are used to describe the presence of an element, integer, step, operation, and / or component but do not exclude the presence or addition of one or more other elements, integers, steps, operations, components, and / or groups thereof.

[0053] It should also be understood that, in the description of the application, unless otherwise specified, the use of the term "or" in the description and the claims of the application means a "and / or" relationship, for example, A or B can mean A or B or A and B. In this application, "and / or" is merely a descriptive relationship of associated objects, which means that there can be three relationships, for example, A and / or B, which means that A exists alone, A and B exist together, and B exists alone, where A and B can be singular or plural. And in the description of the application, unless otherwise specified, "multiple" means two or more. "At least one of the following" or similar expressions means any combination of the items, including any combination of single or multiple items. For example, at least one of a, b, or c can mean a, b, c, a-b, a-c, b-c, or a-b-c, where a, b, and c can be single or multiple.

[0054] In addition, in order to facilitate the clear description of the technical solutions of the embodiments of the application, in the embodiments of the application, the terms "first", "second", etc. are used to distinguish the same items or similar items with basically the same function and role. Those skilled in the art can understand that the terms "first", "second", etc. do not limit the quantity and execution order, and are only used for distinction and description. The terms "first", "second", etc. do not necessarily mean different, and cannot be understood as indicating or implying relative importance.

[0055] As used in the specification and the appended claims of the application, the term "if" can be interpreted as "when" or "upon" or "in response to a determination" or "in response to detecting" depending on the context. Similarly, the phrase "if it is determined" or "if [a described condition or event] is detected" can be interpreted to mean "upon determining" or "in response to determining" or "upon detecting [a described condition or event]" or "in response to detecting [a described condition or event]" depending on the context.

[0056] In the description of the application, the reference "one embodiment" or "some embodiments" means that the specific features, structures or characteristics described in connection with the embodiment are included in one or more embodiments of the application. Therefore, the statements "in one embodiment", "in some embodiments", "in other some embodiments", "in additional some embodiments" and the like appearing in the specification are not necessarily all referring to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "include", "contain", "have" and their variants mean "include but not limited to", unless otherwise specifically emphasized.

[0057] An unmanned aerial vehicle, commonly known as a drone, is a remotely operated aircraft without a human pilot aboard. With the continuous advancement of microelectronics technology, small drones can carry high-performance lightweight computing platforms, have high mobility and high intelligence, and are particularly suitable for use in complex indoor environments such as nuclear power plants. One of the core technologies is how to use lightweight sensors to calculate the position of the drone relative to the takeoff point in the scene when GNSS positioning signals are blocked, which is the key application of simultaneous localization and mapping (SLAM) technology.

[0058] Simultaneous Localization and Mapping (SLAM) is a key robot technology that enables robots to autonomously localize and map unknown environments. SLAM is widely used in autonomous driving, robot navigation, augmented reality (AR), and other fields. Its core goal is to acquire environmental data through sensors such as lidar and vision sensors, and to simultaneously complete localization and mapping using these data. In the early days, SLAM algorithms needed to solve large-dimensional matrices and were usually run on the ground side and transmitted real-time scene information through a communication link to output the position of the drone. However, due to the existence of latency, the robustness of the system is affected. With the improvement of the performance of lightweight embedded processors and the continuous optimization of technology, a stable SLAM system can now be deployed on a drone platform with limited power consumption and payload, enabling high real-time autonomous positioning and meeting the needs of communication-limited environments.

[0059] In summary, there are many structural obstacles in complex environments such as nuclear power plants, which pose high requirements on the size and positioning method of small drones. Therefore, relative positioning algorithms that can be deployed on lightweight embedded computing units are needed to achieve real-time robust estimation of the drone's pose through lightweight environmental sensors. This will solve the positioning problem in the case of GNSS signal blocking, significantly improve the perception ability and efficiency of rotor drones, and have important application significance.

[0060] Limited by the carrying capacity of small drones, the current commonly used relative positioning technology is mainly based on vision sensors and small lidar sensors. Both have their own advantages and can achieve real-time relative pose calculation.

[0061] The UAV relative positioning technology based on laser radar sensor mainly obtains the accurate distance information of the environment through the laser radar sensor. This technology uses the measurement of laser beam reflection time to generate a high-precision environment map, and estimates the pose transformation matrix through the matching of front and back frame point clouds, and then constructs the map based on the point cloud and the calculated pose. This scheme provides high pose estimation accuracy, and since the environment point cloud is measured in real time, it can ensure better mapping quality. In the early stage, two-dimensional laser radar systems have been applied to the autonomous positioning of ground unmanned vehicles, but since the unmanned aerial vehicle is a three-dimensional motion in the air, the traditional two-dimensional scheme cannot effectively obtain height information, thereby limiting the movement range of the unmanned aerial vehicle and affecting its stable flight in complex structural environments. With the development of lightweight three-dimensional laser radars, small three-dimensional radar products can meet the demand of three-dimensional pose estimation. However, with the increasing complexity of the scene structure and the decline of the carrying capacity due to the size limitation of small unmanned aerial vehicles, existing products cannot meet the autonomous positioning needs of all small unmanned aerial vehicles.

[0062] The UAV relative positioning technology based on visual sensor mainly obtains the image information of the environment through visual sensor (such as RGB camera or depth camera), and uses computer vision technology to extract the environmental features in the image information for positioning and map construction. Through the matching of feature points or feature regions extracted from the image, the relative pose change of the UAV in the environment is estimated, and then positioning and mapping are completed.

[0063] Through the matching of feature points of front and back frame images, visual SLAM can be realized to construct an accurate three-dimensional map, and the attitude is estimated according to the pose change matrix. Since the visual sensor provides rich environmental information, and image data is easy to obtain and process, high-precision positioning and map construction can be realized at a low cost. In the early stage, the visual positioning technology based on monocular camera has been widely used in small unmanned aerial vehicles, which can realize pose estimation through motion estimation between consecutive image frames.

[0064] However, due to the lack of depth information in monocular vision, traditional monocular SLAM methods cannot directly obtain accurate positioning in global scale, thus limiting the application of UAVs in large-scale complex environments. With the advancement of binocular camera and depth camera technology, positioning technology based on stereo vision or RGB-D camera can overcome this problem by fusing depth information to improve the accuracy and robustness of pose estimation, supporting UAVs to fly stably in more complex three-dimensional environments. However, with the complexity of environmental structure and the diversification of UAV flight scenarios, relative positioning technology based on visual sensors also faces certain challenges. For example, light changes, dynamic objects, and occlusion problems may affect the extraction and matching of image features, leading to a decrease in positioning accuracy. In addition, the processing of visual sensors requires a large amount of computation and high real-time performance, which puts higher requirements on the processing power and battery life of UAVs.

[0065] In summary, due to the low load capacity requirement of visual image information sensors on UAVs and the higher information entropy contained in the information, the focus of research and development of autonomous positioning technology for small UAVs is the direction of meeting the use requirements of complex structure scenarios.

[0066] Currently, there are many software and hardware solutions for relative positioning technology based on visual sensors. At the software level, due to the limited field of view of traditional visual sensors, it is difficult to cover the environmental information around the UAV, thus lacking the perception of panoramic information; the extraction of image feature information by the visual front-end cannot balance the real-time and accuracy requirements, and false feature information reduces the robustness of the system; the image information processing and large-dimensional matrix information processing in the operation are less efficient on embedded CPUs, which cannot meet the real-time use requirements, making it difficult for relative positioning technology based on vision to meet the safe and robust flight requirements of small UAVs in complex structure scenarios.

[0067] To overcome the deficiencies of the prior art, the embodiments of the present application provide a relative positioning technology for a UAV based on a four-eye fisheye camera and an inertial sensor. The angle between the optical axes of the four-eye fisheye camera is 90 degrees, and the four-eye fisheye camera can cover panoramic information of the UAV. The technology adopts a cylindrical projection process to process real-time scene images, avoids the clipping of image edge information in a traditional rectification scheme, and thus obtains more extensive visual information. By combining a visual neural network model and a traditional SURF corner extraction algorithm, robust feature points are efficiently extracted, and the interference of false feature points is reduced. Based on a traditional VINS-Fusion visual SLAM algorithm, the embodiments of the present application normalize the spherical coordinate information of the feature points of the four-eye fisheye camera, and adopts a GPU-accelerated backend graph optimization algorithm for global positioning optimization, so as to reduce positioning errors and ensure efficient and stable relative positioning. In addition, the embodiments of the present application construct a positioner structure by using a ROS operating system, realize real-time communication with an open-source flight control system, and thus control the autonomous positioning and flight functions of the UAV in a GNSS denial scenario. The technology can effectively improve the positioning accuracy and robustness of the UAV in a complex environment, and meets the demand for high-precision autonomous flight.

[0068] The present application provides an example of an omnidirectional visual positioning system, please refer to Figure 1 The present application provides an example of an omnidirectional visual positioning system, please refer to Figure 1 The present application provides an example of an omnidirectional visual positioning system, please refer to

[0069] The image acquisition device 101 is used to acquire real-time scene images in an omnidirectional range during the flight of the UAV.

[0070] The inertial measurement sensor 102 is connected to the image acquisition device 101 and is used to synchronously acquire motion state data of the UAV.

[0071] The navigation system 103 is connected to the image acquisition device 101 and the inertial measurement sensor 102 and is configured to perform at least one of:

[0072] distortion rectification processing and feature extraction processing on the real-time scene images to obtain processed images; and determining current pose information of the UAV in a spatial coordinate system according to the motion state data and the processed images, wherein the spatial coordinate system is constructed with a takeoff point of the UAV as an origin.

[0073] In some embodiments, the omnidirectional vision positioning system is installed on the top or side of the body of the UAV, which adopts a lightweight design, for example, a small UAV with a total weight of not more than 500 g and a wheelbase of 200 mm. The omnidirectional vision positioning system at least includes the following components: an image acquisition device, an inertial measurement sensor, and a navigation system, each of which can be rigidly connected through a 3D printed structural member to ensure the stability of the mechanical structure.

[0074] In some embodiments, the image acquisition device includes a plurality of fisheye cameras distributed at a specific angle, for example, four fisheye cameras distributed in an X shape, each of which adopts a super-wide-angle optical lens and can cover a field of view angle of 210 degrees, and the included angle between the optical axes of adjacent cameras is 90 degrees. The four-eye image is synchronously acquired through a hardware trigger synchronization circuit, for example, with a time difference of less than 1 ms. The resolution of the raw image acquired by each fisheye camera is 1280x720 pixels.

[0075] In some embodiments, the image acquisition device is connected to the navigation system through a USB 3.0 interface, and the omnidirectional scene image (real-time scene image) acquired is transmitted to the navigation system in real time.

[0076] In some embodiments, the inertial measurement sensor can be a 10-axis inertial measurement unit (IMU) including a three-axis accelerometer, a three-axis gyroscope, and a four-axis magnetometer, which is rigidly connected with the fisheye camera array and installed close to the center of mass of the UAV to reduce motion noise interference. The inertial measurement unit (IMU, for example, with a sampling frequency of 100 Hz) outputs motion state data, for example, including three-axis acceleration, three-axis angular velocity, and magnetic field intensity, to the navigation system through an SPI interface, providing inertial motion constraints for pose solving.

[0077] In some embodiments, the navigation system first performs intrinsic calibration on the real-time scene image, adopts a cylindrical projection model to complete distortion correction processing of the real-time scene image, and retains large field of view angle image information; extracts a plurality of feature points in the real-time scene image and constructs a cross-camera matching relationship through a combination of a deep learning algorithm and a traditional feature detection algorithm, and eliminates false matching points to ensure feature reliability. A unified coordinate system is established with the inertial measurement sensor as the origin, the plurality of feature points are converted to the unified coordinate system, and a factor graph optimization model is constructed in combination with inertial data. Then, a heterogeneous computing architecture is used to accelerate the solving of the re-projection error equation, iteratively optimize the pose parameters of the UAV, and finally output the three-dimensional coordinates and attitude angle relative to the takeoff point.

[0078] By using the embodiment of the present application, based on the integrated image acquisition device, inertial measurement sensor and navigation system in the omnidirectional positioning system, real-time pose information is realized to support autonomous navigation. It is suitable for unmanned aerial vehicle positioning in GNSS denial scenarios, especially in complex environments such as nuclear power plants, and can replace manual remote visual inspection of high-altitude equipment.

[0079] The omnidirectional visual positioning system in the embodiment of the present application belongs to the autonomous positioning technology of the rotor unmanned aerial vehicle in the GNSS denial scenario, covers key technologies such as indoor autonomous positioning, real-time state estimation, multi-sensor fusion and computer vision, and is especially suitable for efficient and stable positioning in complex environments such as large nuclear power plants. It specifically relates to fisheye camera cylindrical imaging model processing technology, efficient feature point extraction method and sliding window based back-end graph optimization technology. The omnidirectional visual positioning system can specifically obtain panoramic image information of the unmanned aerial vehicle (unmanned aerial vehicle) during flight, and convert the panoramic image information into global sparse point cloud of the environment and pose information of the unmanned aerial vehicle relative to the takeoff origin, thereby supporting the multi-rotor unmanned aerial vehicle to perform autonomous navigation and scene perception tasks in a small and complex environment.

[0080] The embodiment of the present application specifically belongs to the field of simultaneous localization and mapping (SLAM) technology in unmanned aerial vehicle technology, relates to visual inertial positioning, fisheye image processing, feature point extraction and graph optimization technology, and realizes real-time pose estimation and panoramic information acquisition of the unmanned aerial vehicle in the case of GNSS signal denial. The system is particularly suitable for small unmanned aerial vehicle autonomous positioning and scene perception tasks in complex indoor environments such as large nuclear power plants, and meets the positioning requirements of high precision and high reliability.

[0081] In an alternative embodiment, the image acquisition device further comprises:

[0082] A plurality of fisheye cameras, the plurality of fisheye cameras are distributed in the fuselage in a predetermined shape, the included angle between the optical axes of each fisheye camera is a first angle, and the field of view angle of a single fisheye camera is not less than a second angle;

[0083] A hardware trigger synchronization circuit for driving each fisheye camera to synchronously acquire images, wherein the acquisition time difference between each fisheye camera is less than a predetermined time length.

[0084] In some embodiments, the plurality of fisheye cameras adopt an X-shaped or cross-shaped layout and are uniformly distributed on the top or periphery of the unmanned aerial vehicle fuselage, the included angle (i.e. the first angle) between the optical axes of adjacent fisheye cameras is 90°, so that the four-camera combination can cover a panoramic range of 360° in the horizontal direction.

[0085] In some embodiments, the optical axis of each fisheye camera forms a specific elevation angle (e.g. 45°) with the horizontal plane, and the vertical field of view angle after superposition can cover a pitch range of ±90°, forming a complete spherical field of view.

[0086] In some embodiments, the field of view (i.e., the second angle) of a single fisheye camera is not less than 210°, and an ultra-wide-angle fisheye lens (such as a 1 / 2.3-inch CMOS sensor with an f=1.8mm focal length lens) can capture panoramic images with extreme distortion.

[0087] In some embodiments, the camera resolution is not less than 1280×720 pixels and the frame rate is ≥30fps, which meets the real-time image processing requirements in dynamic scenes.

[0088] In some embodiments, the hardware trigger synchronization circuit can be built based on an FPGA or a dedicated clock chip, and send trigger pulses to four fisheye cameras simultaneously through differential signal lines to ensure that the exposure start time difference (i.e., the predetermined duration) of each fisheye camera is less than 1ms, thereby avoiding misalignment of multi-view image stitching due to time asynchrony.

[0089] In some embodiments, each fisheye camera can also connect to the navigation system via a USB 3.0 interface or a GigE interface, using time division multiplexing (TDM) technology to package and transmit the four-channel image data to the navigation system, ensuring timing alignment. At the navigation system end, a timestamp matching algorithm is used to synchronize the image frames and IMU data, achieving precise fusion of visual image data and inertial state data.

[0090] Through the aforementioned layout and synchronization mechanism, the image acquisition equipment can obtain seamlessly stitched panoramic image sequences, with an overlap area of ​​≥30% between adjacent cameras, providing sufficient redundant information for subsequent feature matching and 3D reconstruction. In nuclear power plant inspection scenarios, using this omnidirectional vision positioning system, unmanned aerial vehicles can effectively capture minute defects (such as cracks and deformations) on the surface of nuclear power plant equipment, and achieve millimeter-level precision 3D positioning through the triangulation principle of multi-view vision.

[0091] Because the imaging principle of a fisheye camera is different from that of a pinhole camera, it results in... Figure 2a The original fisheye image captured by the fisheye camera shown is significantly distorted compared to a normal image. Traditional distortion correction schemes for fisheye camera imaging only correct distortion near the image center, thus losing the large field of view characteristic of fisheye images. Therefore, this embodiment employs a distortion correction scheme based on a cylindrical projection model, which can preserve a larger field of view even in scenes with small distortions.

[0092] It should be understood that the cylindrical projection model is a method of converting the omnidirectional range of real-time scene images captured by the fisheye lens into visual planar images. Since the fisheye lens usually has a very wide viewing angle range, it can capture almost 210-degree panoramic images, and the cylindrical projection model converts these panoramic images from a spherical or annular region to a two-dimensional planar image with a standard aspect ratio. The corrected image as shown in Figure 2b Although there is still a small amount of distortion, it can be suitable for display and analysis on a standard display device.

[0093] In order to achieve accurate cylindrical projection effect, it is necessary to first correct the distortion of the original fisheye image (panoramic image, real-time scene image). In the image correction process, the OpenCV algorithm is used to effectively correct the distortion of the panoramic image captured by the fisheye lens, thereby obtaining the corrected image.

[0094] First, the standard fisheye model is used to solve the three-dimensional coordinates of each point in the image on the normalized sphere. Then, the three-dimensional coordinates are converted to the corresponding pixel positions in the cylindrical projection image by the cylindrical projection model. In order to ensure the smoothness of the mapping result, for the unmapped pixel positions, the linear interpolation method is used for compensation in the embodiments of the present application, so as to ensure that the image remains good coherence and accuracy in a large field of view angle range. In actual use, due to the installation deviation between cameras, the imaging position of the image is not uniform, and loss in the height direction is easy to occur, so the embodiments of the present application use a fixed window, and a fixed size window (890*770) is used according to the position information of the calibration to ensure that the imaging ranges of the four fisheye images are consistent, which is conducive to subsequent image stitching.

[0095] After obtaining the four images, the embodiments of the present application further perform feature point extraction processing. In order to improve processing efficiency and ensure high-quality feature point extraction, the self-supervised interest point detection and description algorithm (Super Point) is used as the main feature point extraction method. The image feature points detected based on the Super Point feature point detection method are as shown in Figure 2c The Super Point is a self-supervised learning method based on deep learning, which can automatically extract stable and highly recognizable feature points (key points) from images and generate descriptors for each feature point. The advantage of this algorithm is its efficiency and accuracy, especially suitable for resource-constrained embedded environments. In addition, by deploying and running on the GPU of the embedded device, Super Point can provide high processing speed and real-time performance, ensuring that the feature point extraction and matching tasks can be completed with low latency.

[0096] In an alternative embodiment, the navigation system performs distortion correction processing on the real-time scene image, which is configured to perform:

[0097] obtain the three-dimensional coordinates of each feature point in the real-time scene image on a normalized sphere.

[0098] map the three-dimensional coordinates to planar coordinates in a cylindrical projection image using a cylindrical projection model, and compensate for the planar coordinates that are not mapped using a linear interpolation method.

[0099] The cylindrical projection model calibrates the intrinsic matrix and distortion coefficient of each fisheye camera of the image acquisition device through a sensor calibration algorithm.

[0100] In some optional embodiments, the distortion correction processing of the navigation system on the real-time scene image relies on sensor calibration data and projection transformation algorithms. For example, a sensor calibration algorithm (such as the Kalibr algorithm) can be used to calibrate the intrinsic parameters of each fisheye camera in the image acquisition device. Specifically, by shooting multiple groups of calibration board images at different angles (for example, a calibration board containing a matrix of known size two-dimensional codes), the intrinsic matrix (including focal length, principal point coordinates, and other parameters) and distortion coefficient (used to describe the nonlinear distortion characteristics of the fisheye lens) of each camera are calculated.

[0101] During the calibration process, an optimization algorithm is used to minimize the positional deviation between the actual feature points and the projected feature points, ensuring that the accuracy of the intrinsic matrix and the distortion coefficient meets the subsequent processing requirements. The intrinsic matrix and the distortion coefficient obtained by calibration are stored in the memory of the navigation system as basic parameters for subsequent distortion correction, and are used for coordinate conversion of real-time images.

[0102] In some embodiments, for each feature point in the real-time scene image, the distortion coefficient obtained by calibration is first used for distortion correction to eliminate the radial distortion and tangential distortion caused by the fisheye lens, and the preliminary corrected image coordinates are obtained. Then, a geometric transformation algorithm is used to project each feature point from the image plane to the unit sphere, forming a point coordinate in three-dimensional space, which reflects the directional information of the feature point in the camera coordinate system.

[0103] In some embodiments, a cylindrical projection model is used to map the three-dimensional coordinates on the normalized sphere to planar coordinates in a cylindrical projection image. For example, the spherical point is projected radially onto a virtual cylindrical surface, and then the cylindrical surface is unfolded into a two-dimensional planar image, thereby converting the ultra-wide-angle image of the fisheye camera into a rectangular image suitable for standard display devices, while maximizing the preservation of the edge field of view information of the original image.

[0104] In some embodiments, for the plane coordinates that are not covered in the projection process (i.e., pixel holes that may occur in the original image edge area after projection), a linear interpolation method is used for compensation: by calculating the gray value or color value of the adjacent known pixels, the unmapped positions are filled in proportion to ensure that the projected image remains coherent and complete within a large field of view range, avoiding feature extraction failure due to pixel loss.

[0105] With the above embodiments, the camera intrinsic distortion coefficient is obtained by the Kalibr algorithm, the cylindrical projection model is used to retain the large field of view characteristics of the fisheye image, and the linear interpolation compensation is used to ensure the integrity of the image, thereby providing high-quality visual data for subsequent feature extraction and pose solving, and meeting the positioning needs of the unmanned aerial vehicle in complex environments (such as high-altitude equipment detection in nuclear power plants).

[0106] In order to construct the correlation of the feature points, the embodiment of the application uses a global feature matching network based on Mobile Net VLAD as a feature point matcher. Through this network, the matching relationship between the feature points can be effectively established. The result of feature point matching can be used for further depth information calculation. Combined with the extrinsic information of the camera, the depth information of the feature points in the camera common view range can be solved, and these information is used for camera pose estimation. In the pose estimation process, the BA (Bundle Adjustment) optimization algorithm can be used to solve the pose transformation matrix of the camera between the two frames, and then the pose information of the camera is roughly inferred. This process provides important support for further camera positioning and scene reconstruction.

[0107] Considering that in some environments there may be weak feature information, the embodiment of the application proposes to use the Speeded-Up Robust Features (SURF) algorithm as a supplementary feature point detection method. SURF is a classic and efficient feature point extraction algorithm that can provide additional stable points in scenes with insufficient feature points.

[0108] In order to improve the accuracy of feature point matching, the embodiment of the application further introduces the Random Sample Consensus (RANSAC) algorithm for matching result verification. The RANSAC algorithm can identify and remove inconsistent matching points through random sampling method, thereby improving the robustness and accuracy of feature point matching, so that the embodiment of the application can efficiently complete the feature point matching and subsequent image stitching task in complex environments.

[0109] The feature point extraction and tracking algorithm based on the fisheye image adopted by the embodiments of the present application can efficiently and accurately process the distortion correction, feature extraction and matching of the four-fisheye image, and can use GPU computing resources to accelerate image data processing and calculation, which is beneficial to improve the real-time calculation and edge device application of the omnidirectional vision positioning system in a dynamic environment.

[0110] To convert each feature point in the four-camera to a unified coordinate system, the embodiments of the present application take the installation position of the inertial measurement unit (IMU) as the origin of the unified coordinate system, use the extrinsic matrix of each camera and the IMU to convert the normalized spherical coordinates based on the camera into the normalized spherical coordinates based on the IMU, and filter out the matched points, thereby reducing the dimension of the subsequent optimization matrix.

[0111] To obtain stable and reliable pose estimation values, after obtaining the above feature points and IMU sensing data, the autonomous positioning algorithm is brought into the framework of the Visual-Inertial Navigation System (VINS) algorithm, and the four-fisheye autonomous positioning algorithm based on the VINS algorithm is constructed. In the initialization stage, the installation transformation relationship between the multi-camera and the IMU is established through a certain time of visual-IMU joint calibration, the depth information of the feature points in the common view area is solved by using the matching relationship between the feature points and the prior installation position relationship between the cameras in the front-end feature point processing, and the Bundle Adjustment (BA) is used to obtain the position and attitude change of the current frame image relative to the previous frame by iteratively optimizing and converging the feature point re-projection error as a constraint, and the inertial data constraint condition between the two frames is established by using the IMU pre-integration between the two image frames.

[0112] Due to the existence of cumulative error, a local joint optimization and repositioning module based on a sliding window is adopted to establish a constraint relationship between the current frame pose relationship and a certain number of historical key frames, and a more robust unmanned aerial vehicle pose relationship is obtained. In addition, the current key frame is connected with the global map, when the loop detector detects that the current scene matches the global map, the global pose optimizer is used to eliminate the long-time cumulative error, thereby ensuring the unbiased relative position of the unmanned aerial vehicle.

[0113] In an alternative embodiment, the navigation system performs feature extraction processing on real-time scene images, which is configured to perform:

[0114] extracting a plurality of feature points of the real-time scene image based on a deep learning algorithm and generating a descriptor of each feature point.

[0115] Based on the descriptor of each feature point, a feature matching network is used to construct a matching relationship of a plurality of feature points.

[0116] In some embodiments, the navigation system relies on deep learning and feature matching network to realize feature extraction processing of real-time scene images. For example, a lightweight deep learning model such as Super Point algorithm is deployed on the GPU of the navigation system. Through self-supervised learning, the model can detect feature points in real-time scene images and generate corresponding descriptor vectors.

[0117] In some embodiments, for the pre-processed real-time scene image, the deep learning model first extracts multi-scale features through the convolution layer, and then generates a feature point heat map through the Soft max classifier to locate the feature points in the image. Meanwhile, another branch network generates a 256-dimensional descriptor vector for each feature point. The vector has rotation and scale invariance, which is used for subsequent cross-image matching.

[0118] In some embodiments, a lightweight network is used to construct a feature matcher (for example, a lightweight image retrieval deep neural network, Mobile Net VLAD), which realizes efficient cross-view feature matching through local feature aggregation and global descriptor generation.

[0119] In some embodiments, after performing feature extraction on the real-time scene images collected by the four cameras in parallel, the initial matching relationship is established by calculating the cosine similarity of the feature descriptors between different cameras. Then, the bidirectional matching verification and epipolar constraint are used to filter the mismatched points, and finally a reliable cross-camera feature point correspondence relationship is constructed to provide visual constraints for subsequent pose solving.

[0120] By extracting image feature points and generating descriptors through deep learning algorithms, and constructing cross-camera feature association using a feature matching network, visual constraints are provided for the positioning of unmanned aerial vehicles in complex environments. In the scene of detecting high-altitude equipment in nuclear power plants, this scheme can effectively identify the subtle defect features on the surface of the equipment, and realize high-precision three-dimensional positioning through multi-view feature matching.

[0121] In an alternative implementation, the navigation system determines the current pose information of the unmanned aerial vehicle in the spatial coordinate system according to the motion state data and the processed images, and is configured to perform:

[0122] Obtain the installation transformation relationship between the image acquisition device and the inertial measurement sensor, and the installation position relationship between the multiple fisheye cameras included in the image acquisition device.

[0123] Obtain the feature extraction processing of the real-time scene images to obtain the matching relationship between the feature points.

[0124] According to the installation transformation relationship, the installation position relationship, and the matching relationship between the feature points, the depth information of the common view feature points of the multiple fisheye cameras is solved.

[0125] A bundle adjustment method is used to solve the pose change of the current frame scene image relative to the previous frame scene image according to the depth information of the feature points in the common view area, with the feature point re-projection error as a constraint condition.

[0126] If the current scene is identified as matching the historical scene in the global map, the current pose information of the unmanned aerial vehicle in the spatial coordinate system is determined based on the global map and the pose change.

[0127] In an alternative embodiment, the navigation system preforms a specific motion trajectory by controlling the unmanned aerial vehicle, synchronously collects motion state data and real-time scene images in an omnidirectional range, solves the installation transformation relationship between the image acquisition device and the inertial measurement sensor through iterative optimization of the equation, and simultaneously uses a high-precision calibration board to observe from multiple angles, solves the rotation matrix and translation vector between the cameras through a nonlinear optimization algorithm, and acquires the installation position relationship between the multiple fisheye cameras.

[0128] Next, the processed image after distortion correction is used to extract multi-scale feature points in parallel on a GPU using an improved Super Point algorithm, an initial match is constructed based on the cosine similarity of the feature descriptors, and the installation position relationship between the cameras is used to eliminate mismatched points through epipolar constraint to obtain the matching relationship between the feature points. Then, the initial depth value is calculated through triangulation for the feature point pairs that have successfully matched using the installation position relationship between the cameras, and the initial depth value is used as a priori to construct an optimization problem and iteratively solve to obtain the depth information of the feature points in the common view area. Then, a bundle adjustment method is used to construct a sliding window optimization problem, and the camera pose and the three-dimensional coordinates of the feature points are used as state variables, and the relative pose transformation is calculated through IMU pre-integration between adjacent two frames and added as a constraint term to the optimization to solve the pose change of the current frame relative to the previous frame, realize fast and stable inter-frame pose solving, and meet the real-time requirement.

[0129] Finally, when the current scene is identified as matching the historical scene in the global map through the bag-of-words model, a pose graph containing camera pose nodes and loop edges is constructed for global optimization to reduce the cumulative positioning error, and the current pose information of the unmanned aerial vehicle in the spatial coordinate system is determined based on the global map and the pose change. The unmanned aerial vehicle can also achieve high-precision positioning and stable navigation in complex environments, and meet the stringent requirements of nuclear power plant inspection and other scenes.

[0130] The core idea of the back-end optimization module is to use the estimated state to construct a nonlinear least squares optimization problem using the constraint relationship. Due to the large number of state variables to be solved, a large-dimensional matrix equation needs to be solved. The classic algorithm is limited by the performance of the CPU in the embedded system, and the solving speed will decrease, which will lead to the failure of the pose solution. Therefore, the embodiments of the present application parallelize and optimize the large-time-consuming module based on GPU, and use GPU to perform the large-dimensional matrix equation solving calculation in the module, which can improve the running efficiency of the algorithm on the embedded computing device with the GPU module, and provide stable pose information for the subsequent autonomous obstacle avoidance navigation system.

[0131] The factor graph constructed in the back-end graph optimization is as shown in Figure 3 As shown in the figure, each state is regarded as a node, and the mathematical relationship of different state vectors is regarded as an edge, so that the factor graph can be constructed, and the target function can be analyzed more intuitively with the help of the factor graph. Figure 3 The circular node in represents the observation, mainly the IMU observation, the visual front-end observation and the residual information, and the square node represents the constraint relationship between different types of observations at different times, so that the graph optimization method based on the factor graph can be used to construct the matrix equation of each iteration quantity.

[0132] Since the state variable dimension in the matrix equation usually reaches hundreds, the traditional matrix solving method has a low calculation speed, so the Schur complement solving scheme is used on the traditional Cere graph optimization solver, that is, the camera pose, IMU sensor bias, and state iteration of the camera and IMU extrinsic parameters are decoupled from the depth information state iteration in the solving process, and the calculation process is simplified.

[0133] In addition, since the lightweight embedded system cannot carry high-performance CPU computing resources, the traditional VINS algorithm cannot run on the platform with limited CPU computing resources to have real-time performance and stability, so improving the resource usage mode of the traditional algorithm and fully utilizing the embedded computing resources can effectively improve the solving speed of the algorithm. The Jetson series of embedded development platforms provide efficient GPU computing resources and the matching CUDA (Compute Unified Device Architecture) GPU computing architecture, which can construct kernel functions to implement the parallel steps in the algorithm. Compared with the traditional embedded heterogeneous platform, the CUDA function can be programmed and developed in a general way through a C-like language, and the GPU can be easily called for parallel operation. Main components and programming modes of CUDA.

[0134] In actual use, the present application puts large-scale matrix multiplication operation on the GPU for running, and places the matrix equation solving content on the CPU for running, and the running framework is as shown in Figure 4As shown. In addition, since the matrix needs to use a large amount of GPU memory resources, the application embodiment adopts the two-level separation adaptive memory allocation algorithm (TLSF) provided by the memory allocator (Vulkan Memory Al locator, VMA), which allows the existing GPU memory allocation to be allocated twice, improving the algorithm execution efficiency.

[0135] The application embodiment method is based on visual-inertial multi-sensor fusion, and constructs a pose estimation system as shown in the schematic diagram of the pose estimation system.As shown in the schematic diagram of the pose estimation system, the fisheye camera is used to acquire environmental visual information, the IMU inertial sensor is used to collect motion inertia data, through the links of feature extraction, matching, sliding window optimization and loop detection, precise pose estimation is realized, and the system process covers data input, processing, optimization and output stages. Figure 5

[0136] (I) Data input and preprocessing

[0137] Fisheye camera data: four fisheye cameras (fisheye cameras 1-4) are used to collect environmental images. Since the fisheye camera has distortion, the “graph distortion correction based on cylindrical projection model” operation needs to be performed on the images collected by each camera. Using the cylindrical projection model, according to the internal parameters of the fisheye camera (such as focal length, distortion coefficient, etc.), the distorted image pixel points are mapped to the cylindrical projection plane, the radial and tangential distortion is corrected, and the corrected orthographic image is obtained, providing high-quality visual data for subsequent feature extraction.

[0138] IMU data collection: the IMU inertial sensor collects real-time inertial data such as angular velocity and linear acceleration of the carrier as motion perception input for pose estimation.

[0139] (II) Visual feature extraction and matching

[0140] Feature point extraction: after the fisheye camera image is corrected for distortion, the Super Point feature point extraction algorithm and the SURF feature point extraction algorithm are used in parallel. Super Point uses a deep learning network to extract robust feature points and corresponding descriptors from images, suitable for different lighting and texture environments; the SURF algorithm constructs a scale space and uses the Hessian matrix to detect key points in the image, calculates their direction and descriptor, and obtains local features. The feature points extracted by the two algorithms are jointly input into the feature point descriptor generation link, and the pixel gradient, texture and other information around the feature points are encoded to form unique descriptors for subsequent matching.

[0141] Feature matching:

[0142] Co-view camera matching: using "RANSAC algorithm co-view camera feature matching", the feature points of different fisheye cameras at the same or adjacent time (there are co-view areas) are calculated to obtain the distance (such as Hamming distance, Euclidean distance) between the descriptors, and the initial matching pairs are screened; then, the RANSAC algorithm is used for random sample consensus test to remove the false matching and retain the correct matching pairs that meet the geometric constraints (such as epipolar constraint), so as to realize the alignment of the features between the co-view cameras.

[0143] Cross-camera matching: using "Mobile Net VLAD cross-camera feature matching", Mobile Net VLAD uses a lightweight neural network to extract global image features, and a vector aggregation descriptor (VLAD) is used to aggregate and encode the features to construct a cross-camera feature vector space. By calculating the similarity of the feature vectors of different cameras, the feature matching between non-co-view cameras is realized, and the range of environmental feature correlation is expanded.

[0144] (Three) IMU data processing

[0145] IMU pre-integration: the angular velocity and linear acceleration data collected by the IMU are subjected to "IMU pre-integration" operation. Based on the basic equation of inertial navigation, the continuous IMU measurement values are integrated in a short time window to obtain the relative attitude, position and velocity changes of the carrier. During the pre-integration process, the IMU noise model (such as Gaussian white noise, random walk) is considered to compensate for the errors of the integration results, and the relative motion constraints are generated to provide motion prior information for visual-inertial fusion.

[0146] IMU propagation and pose calculation: after the pre-integration processing, the motion information obtained by the pre-integration is used to recursively calculate the "IMU pose" of the carrier, including attitude angles (heading angle, pitch angle, roll angle), position coordinates, etc., by IMU propagation and combining the initial pose provided by the initialization link, as the inertial prediction result of pose estimation.

[0147] (Four) sliding window and key frame management

[0148] Initialization:

[0149] Sensor alignment: performing visual-inertial sensor alignment, by synchronizing the image acquisition time of the fisheye camera with the data acquisition time of the IMU, establishing a timestamp correspondence relationship, and ensuring that the multi-sensor data is matched in the time domain, laying a foundation for fusion calculation.

[0150] Motion pose recovery: Use visual-inertial information-based motion pose recovery. In the system startup or re-initialization phase, combine the visual feature matching results at the initial time (such as the motion parallax of adjacent frame feature points) and the IMU pre-integration data, use the Perspective-n-Point (PNP) algorithm or the Extended Kalman Filter (EKF) initialization method, and recover the initial pose (position and attitude) of the carrier, providing initial values for subsequent sliding window optimization.

[0151] Sliding window construction: Enter the sliding window phase, and the window contains multiple key frames (key frames 1-k). Key frame selection is based on image feature change rate, motion blur degree, etc. When the feature difference (such as feature point number change, descriptor distance average) between the current frame and the key frames in the window exceeds the set threshold, or the motion blur causes the feature extraction quality to decrease, the current frame is determined as a new key frame and added to the window, and the oldest key frame is removed, maintaining the number of key frames in the window within a reasonable range (such as 5-15 frames), ensuring calculation efficiency and accuracy.

[0152] Local visual-inertial odometry: In the sliding window, local visual-inertial odometry based on sliding window information is performed. Fuse the visual feature matching constraints of each key frame in the window, such as the three-dimensional spatial coordinates of feature points and the projection residual, and the IMU pre-integration motion constraints, use graph optimization or filtering algorithms (such as Bundle Adjustment (BA) algorithm based on sliding window, Extended Kalman Filter (EKF)), optimize the poses of all key frames in the window and the coordinates of map points (obtained by triangulation of visual features), and obtain local high-precision pose estimation results, updating the carrier motion state in real time.

[0153] (5) Loop detection and optimization

[0154] Loop detection process: sequentially perform loop detection → key frame determination → loop detection → loop feature pairing. Loop detection compares the feature descriptors (such as MobileNet VLAD global features, Super Point / SURF local features) of the current key frame and the historical key frames (from the key frame database), calculates the similarity index (such as cosine similarity, distance ratio), and when the similarity exceeds the set threshold, it is determined that there is a loop. Key frame determination further selects the key frame pair participating in the loop, ensuring that it has an overlapping area in space and a reasonable motion relationship; loop detection verifies the authenticity of the loop through geometry (such as calculating the essential matrix, the fundamental matrix); loop feature pairing accurately matches the feature points between the loop key frames, establishing long-term spatial constraints.

[0155] Global pose graph optimization: based on the constraints obtained by loop closure feature matching, perform "4-DoF global pose graph optimization" (4-DoF usually refers to position (x, y, z) and heading angle, or defined according to actual needs). Construct a global pose graph, with nodes as key frame poses and edges as visual feature matching constraints (including loop closure constraints) and IMU pre-integration constraints. Use a graph optimization algorithm (such as g2o, Ceres) to optimize the global pose graph, minimize the sum of squares of all constraint residuals, eliminate cumulative errors, and obtain globally consistent pose estimation results. At the same time, update the key frame database to store the optimized key frame poses and feature information for subsequent loop closure detection.

[0156] Through the above process, the final output global pose information, including the accurate position (such as three-dimensional coordinates) and attitude (such as quaternion, Euler angle representation) of the carrier in the global coordinate system, can be applied to robot autonomous navigation, unmanned aerial vehicle positioning, AR / VR space positioning, etc. Scene, to provide real-time, high-precision pose perception for the carrier, to support subsequent path planning, motion control, etc. Function implementation.

[0157] The method provided by the embodiments of the application is based on the fusion of fisheye cameras and IMU inertial sensors to construct a pose estimation process as shown in Figure 6 which covers image distortion correction, feature extraction and matching, IMU pre-integration, sliding window management, pose optimization and loop closure detection, etc. It realizes from local to global accurate estimation of pose, and adapts to robot navigation, unmanned aerial vehicle positioning, etc. Scene.

[0158] (I) Image acquisition and distortion correction

[0159] Image acquisition: deploy 4 fisheye cameras (fisheye cameras 1-4) to synchronously acquire environment images from different perspectives, covering the three-dimensional space around the carrier, and providing multi-view visual data for subsequent processing.

[0160] Distortion correction: for the image output by each fisheye camera, perform "cylindrical projection model-based image distortion correction". According to the fisheye camera calibration parameters (focal length, distortion coefficient, etc.), establish the mapping relationship between the distorted pixels and the cylindrical projection plane pixels, correct the radial and tangential distortion, and convert the fisheye image into an approximately orthographic image. Facilitate feature extraction, complete single camera image distortion correction, and 4 cameras enter the feature extraction link after processing.

[0161] (II) Visual feature extraction and matching

[0162] Feature point extraction: run "Super Point feature point extraction" and "SURF feature point extraction" in parallel on the distorted image.

[0163] Super Point: Using a pre-trained deep learning network, input the distortion-corrected image, extract the feature map through the convolution layer, filter the key points through non-maximum suppression, generate two-dimensional feature point coordinates and corresponding descriptors (such as 128-dimensional vectors), and adapt to complex texture and lighting environment.

[0164] SURF: Construct image scale space, calculate Hessian matrix to detect key points, assign key point direction (based on neighborhood pixel gradient), generate 64 / 128-dimensional descriptors, and quickly extract local significant features. The feature points extracted by the two algorithms are summarized and entered into "feature point descriptor generation", unified coding format, and enhanced matching compatibility.

[0165] Feature matching:

[0166] Co-view camera matching: "RANSAC algorithm co-view camera feature matching" is used. For feature points of different fisheye cameras (with overlapping field of view) at the same time, calculate the Euclidean distance of the descriptors, select the initial matching pairs; use RANSAC algorithm for random sampling, combined with epipolar constraint to remove false matches, retain matching pairs that meet geometric consistency, and establish feature association between co-view cameras.

[0167] Cross-camera matching: "Mobile Net VLAD cross-camera feature matching" is executed. Mobile Net extracts global image features, VLAD aggregates and encodes feature vectors, and constructs a cross-camera feature database; calculate the similarity (such as cosine similarity) of feature vectors of different cameras, match non-co-view camera features, expand the coverage of environmental features, and provide more constraints for pose estimation.

[0168] (Three) IMU data processing and sliding window management

[0169] IMU pre-integration: IMU inertial sensor collects angular velocity and linear acceleration data, and performs "IMU pre-integration". Based on the principle of inertial navigation, within a short time window, the relative motion of the carrier is integrated and calculated. Considering the IMU noise (Gaussian white noise, random walk), real-time compensation is performed through error state Kalman filtering, and pre-integrated motion constraints are generated for pose fusion.

[0170] Sliding window management: A "sliding window (record s historical key frames)" is constructed, and the window contains key frames (k-s) to k. Key frame selection is based on feature change rate (difference in the number of feature points between adjacent frames, average descriptor distance) and motion blur degree (image gradient variance). When the current frame meets the key frame conditions (such as feature change rate exceeding threshold, low blur degree), it is added to the window, and the oldest key frame k-s is removed. "Edge processing" is performed, and through Schur complement or sparse matrix decomposition, the constraint information between key frames (such as feature matching residuals, IMU pre-integration constraints) is retained, the calculation amount is compressed, and s key frames are maintained in the window (s is set to 5-15 according to the calculation resources and precision requirements).

[0171] (Four) Pose Optimization

[0172] Local pose optimization: Based on the key frames in the sliding window, "local information-based pose optimization" is performed. The visual feature matching constraints and the IMU pre-integration constraints are fused to construct a least squares optimization problem, and the Levenberg-Marquardt algorithm is used for iterative solution. The key frame poses and map point coordinates in the window are optimized, and "current pose information" is output to update the carrier motion state in real time.

[0173] Global pose optimization: "Loop closure detection" and "loop closure feature matching" are combined. Loop closure detection compares the feature descriptors (such as Mobile Net VLAD global features) of the current key frame and the historical key frames (stored in the global database) to calculate the similarity (such as distance ratio) and determine the loop closure (i.e., the carrier returns to the area it has passed through). Loop closure feature matching accurately associates the feature points of the loop closure key frame to establish long-term spatial constraints. Based on this, "global information-based pose graph optimization" is performed to construct a global pose graph (nodes are key frame poses, and edges are visual constraints, IMU constraints, and loop closure constraints). The graph optimization framework (such as g2o) is used to minimize the global constraint residuals to eliminate accumulated errors, and "global pose information" is output to achieve global consistent pose estimation.

[0174] (Five) Application Scenario Adaptation

[0175] Through the above method process, the output global pose information (including position coordinates, attitude angle / quaternion) can be directly applied to mobile robot autonomous navigation (path planning, obstacle avoidance), unmanned aerial vehicle positioning (flight path control), AR / VR space positioning (virtual scene registration), etc. scenarios, providing centimeter-level, real-time pose perception for the carrier, and supporting motion control and interaction functions in complex environments.

[0176] In the embodiments of the present application, in order to ensure the synchronization of fisheye camera images, the four camera images are merged into one image in the running, and finally the algorithm is run on Jetson Orin NX, and the pose information is output at a speed of 10Hz.

[0177] According to an embodiment of the present application, an unmanned aerial vehicle is provided for a global navigation satellite system (GNSS) denial scenario, and the unmanned aerial vehicle comprises: a fuselage and an omnidirectional visual positioning system according to any one of the preceding embodiments; the omnidirectional visual positioning system is mounted on the fuselage of the unmanned aerial vehicle, and the omnidirectional visual positioning system comprises:

[0178] an image acquisition device configured to acquire real-time scene images in an omnidirectional range during flight of the unmanned aerial vehicle.

[0179] an inertial measurement sensor connected to the image acquisition device and configured to synchronously acquire motion state data of the unmanned aerial vehicle.

[0180] a navigation system connected to the image acquisition device and the inertial measurement sensor and configured to perform at least one of:

[0181] distortion correction processing and feature extraction processing on the real-time scene images to obtain processed images; and determining current pose information of the unmanned aerial vehicle in a spatial coordinate system according to the motion state data and the processed images, wherein the spatial coordinate system is constructed with a takeoff point of the unmanned aerial vehicle as an origin.

[0182] In some embodiments, the omnidirectional visual positioning system is mounted on a top or side of the fuselage of the unmanned aerial vehicle, and the unmanned aerial vehicle adopts a lightweight design, for example, a small unmanned aerial vehicle with a total weight of not more than 500 g and a wheelbase of 200 mm. The omnidirectional visual positioning system comprises at least the following components: the image acquisition device, the inertial measurement sensor, and the navigation system, and each component can be rigidly connected through a 3D printed structural member to ensure mechanical stability.

[0183] In some embodiments, the image acquisition device comprises a plurality of fisheye cameras distributed at specific angles, for example, four fisheye cameras distributed in an X shape, each fisheye camera adopts an ultra-wide-angle optical lens and can cover a field of view angle of 210 degrees, and an optical axis angle between adjacent cameras is 90 degrees. The four fisheye cameras are driven by a hardware trigger synchronization circuit to realize synchronous image acquisition, for example, with a time difference of less than 1 ms. The resolution of a raw image acquired by each fisheye camera is 1280x720 pixels.

[0184] In some embodiments, the image acquisition device is connected to the navigation system through a USB 3.0 interface and transmits the acquired omnidirectional scene images (real-time scene images) to the navigation system in real time.

[0185] In some embodiments, the inertial measurement sensor can be a 10-axis inertial measurement unit (IMU) including a three-axis accelerometer, a three-axis gyroscope, and a four-axis magnetometer, rigidly connected with the fisheye camera array, and installed close to the center of mass of the UAV to reduce motion noise interference. The inertial measurement unit (IMU) outputs motion state data, such as three-axis acceleration, three-axis angular velocity, and magnetic field strength, to the navigation system through an SPI interface, providing inertial motion constraints for pose solving.

[0186] In some embodiments, the navigation system first performs intrinsic calibration on real-time scene images, uses a cylindrical projection model to complete distortion correction processing of the real-time scene images, and retains image information of a large field of view. Through a combination of deep learning algorithms and traditional feature detection algorithms, multiple feature points in the real-time scene images are extracted and a cross-camera matching relationship is constructed, and false matching points are removed to ensure feature reliability. A unified coordinate system is established with the inertial measurement sensor as the origin, the multiple feature points are converted to the unified coordinate system, and a factor graph optimization model is constructed in combination with inertial data. Then, a heterogeneous computing architecture is used to accelerate the solution of the re-projection error equation, and the pose parameters of the UAV are iteratively optimized, and finally the three-dimensional coordinates and attitude angles relative to the takeoff point are output.

[0187] By using the embodiments of the present application, based on the image acquisition device, the inertial measurement sensor, and the navigation system integrated in the omnidirectional positioning system, real-time pose information is realized to support autonomous navigation. It is suitable for UAV positioning in GNSS denial scenarios, especially in complex environments such as nuclear power plants, and can replace manual remote visual inspection tasks of high-altitude equipment.

[0188] The omnidirectional visual positioning system in the embodiments of the present application belongs to the autonomous positioning technology of the rotor UAV in the GNSS denial scenario, covers key technologies such as indoor autonomous positioning, real-time state estimation, multi-sensor fusion, and computer vision, and is especially suitable for efficient and stable positioning in complex environments such as large nuclear power plants. It specifically relates to fisheye camera cylindrical imaging model processing technology, efficient feature point extraction method, and back-end graph optimization technology based on a sliding window. The omnidirectional visual positioning system can specifically obtain panoramic image information of the UAV (unmanned aerial vehicle) during flight, and convert the panoramic image information into global sparse point cloud of the environment and pose information of the UAV relative to the takeoff point, thereby supporting the multi-rotor UAV to perform autonomous navigation and scene perception tasks in a small and complex environment.

[0189] This application specifically belongs to the field of Simultaneous Localization and Mapping (SLAM) technology in unmanned aerial vehicle (UAV) technology. It involves technologies such as visual inertial positioning, fisheye image processing, feature point extraction, and graph optimization. Even in the absence of GNSS signals, it enables real-time pose estimation and panoramic information acquisition for UAVs. This system is particularly suitable for autonomous positioning and scene perception tasks of small UAVs in complex indoor environments such as large nuclear power plants, meeting the requirements for high-precision and high-reliability positioning.

[0190] The unmanned aerial vehicle provided in this application embodiment can replace manual labor to remotely inspect high-altitude static mechanical equipment such as high-altitude air doors, high-altitude supports and hangers of conventional islands, expansion joints, internal fasteners of condensers, fire-fighting glass bulbs, and steel structures of pump station drum nets in nuclear power plants, without the need for personnel to climb onto the inner ring crane of the nuclear island or to erect scaffolding. This reduces safety risks and ensures the overhaul schedule.

[0191] For example, visual inspections of the dome fasteners are required before and after the reactor cover is opened and closed. This work requires workers to climb onto the inner hoist of the nuclear island and is classified as a Level 1 high-risk operation. In addition, there are inspections of the ventilation doors in the nuclear island that use non-standard suspended baskets, the high-altitude supports and hangers of the conventional island, expansion joints, internal fasteners of the condenser, fire-fighting glass bulbs, and the steel structure of the pump station's ductwork. Because the objects being inspected are located at high altitudes, these types of work are classified as Level 2 or higher high-risk operations. Furthermore, scaffolding needs to be erected for each object before performing these types of work, which involves a significant amount of scaffolding erection work.

[0192] According to the embodiments of this application, please refer to Figure 7 As shown, Figure 7 This is a flowchart illustrating an omnidirectional visual positioning method according to an embodiment of this application. An omnidirectional visual positioning method is provided, including:

[0193] S701 acquires real-time scene images of the omnidirectional range from the image acquisition device of the UAV during its flight, and acquires motion state data of the UAV synchronously collected by the inertial measurement sensor of the UAV.

[0194] S702 performs at least one of distortion correction processing and feature extraction processing on the real-time scene image to obtain the processed image.

[0195] S703 determines the current pose information of the unmanned aerial vehicle in the spatial coordinate system based on the motion state data and the processed image. The spatial coordinate system is constructed with the take-off point of the unmanned aerial vehicle as the origin.

[0196] In some embodiments, the omnidirectional vision positioning system is installed on the top or side of the body of the UAV, which adopts a lightweight design, for example, a small UAV with a total weight of not more than 500 g and a wheelbase of 200 mm. The omnidirectional vision positioning system at least includes the following components: an image acquisition device, an inertial measurement sensor, and a navigation system, each of which can be rigidly connected through a 3D printed structural member to ensure the stability of the mechanical structure.

[0197] In some embodiments, the image acquisition device includes a plurality of fisheye cameras distributed at a specific angle, for example, four fisheye cameras distributed in an X shape, each of which adopts a super-wide-angle optical lens and can cover a field of view angle of 210 degrees, and the included angle between the optical axes of adjacent cameras is 90 degrees. The four-eye image is synchronously acquired through a hardware trigger synchronization circuit, for example, with a time difference of less than 1 ms. The resolution of the raw image acquired by each fisheye camera is 1280x720 pixels.

[0198] In some embodiments, the image acquisition device is connected to the navigation system through a USB 3.0 interface, and the omnidirectional scene image (real-time scene image) acquired in real time is transmitted to the navigation system.

[0199] In some embodiments, the inertial measurement sensor can be a 10-axis inertial measurement unit (IMU) including a three-axis accelerometer, a three-axis gyroscope, and a four-axis magnetometer, which is rigidly connected with the fisheye camera array and installed close to the center of mass of the UAV to reduce motion noise interference. The inertial measurement unit (IMU, for example, with a sampling frequency of 100 Hz) outputs motion state data, for example, including three-axis acceleration, three-axis angular velocity, and magnetic field intensity, to the navigation system through an SPI interface, providing inertial motion constraints for pose solving.

[0200] In some embodiments, the navigation system first performs intrinsic calibration on the real-time scene image, adopts a cylindrical projection model to complete distortion correction processing of the real-time scene image, and retains large field of view angle image information; extracts a plurality of feature points in the real-time scene image and constructs a cross-camera matching relationship through a combination of a deep learning algorithm and a traditional feature detection algorithm, and eliminates false matching points to ensure feature reliability. A unified coordinate system is established with the inertial measurement sensor as the origin, the plurality of feature points are converted to the unified coordinate system, and a factor graph optimization model is constructed in combination with inertial data. Then, a heterogeneous computing architecture is used to accelerate the solving of the re-projection error equation, iteratively optimize the pose parameters of the UAV, and finally output the three-dimensional coordinates and attitude angle relative to the takeoff point.

[0201] The embodiment of the present application is based on the integrated image acquisition device, inertial measurement sensor and navigation system in the omnidirectional positioning system, and realizes real-time release of the pose information to support autonomous navigation. The embodiment of the present application is suitable for unmanned aerial vehicle positioning in a GNSS denial scenario, and can replace manual work to complete remote visual detection of high-altitude equipment in a complex environment such as a nuclear power plant.

[0202] The omnidirectional visual positioning system in the embodiment of the present application belongs to autonomous positioning technology of a rotor unmanned aerial vehicle in a GNSS denial scenario, covers key technologies such as indoor autonomous positioning, real-time state estimation, multi-sensor fusion and computer vision, and is especially suitable for efficient and stable positioning in a complex environment such as a large nuclear power plant. Specifically, the embodiment of the present application relates to fisheye camera cylindrical imaging model processing technology, efficient feature point extraction method and back-end graph optimization technology based on a sliding window. The omnidirectional visual positioning system can specifically acquire panoramic image information of an unmanned aerial vehicle (unmanned aerial vehicle) during flight, and convert the panoramic image information into global sparse point cloud of an environment and pose information of the unmanned aerial vehicle relative to a takeoff origin, thereby supporting autonomous navigation and scene perception tasks of a multi-rotor unmanned aerial vehicle in a small and complex environment.

[0203] The embodiment of the present application specifically belongs to the field of simultaneous localization and mapping (SLAM) technology in unmanned aerial vehicle technology, and relates to visual inertial positioning, fisheye image processing, feature point extraction and graph optimization technology. In the case of GNSS signal denial, the embodiment of the present application realizes real-time pose estimation and panoramic information acquisition of an unmanned aerial vehicle. The system is especially suitable for autonomous positioning and scene perception tasks of a small unmanned aerial vehicle in a complex indoor environment such as a large nuclear power plant, and meets the positioning requirements of high precision and high reliability.

[0204] It should be understood that the size of the serial number of each step in the above embodiment does not mean the order of execution, and the execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiment of the present application.

[0205] The embodiment of the present application also provides an electronic device, which comprises one or more processors and a memory;

[0206] The memory is coupled with the one or more processors, and the memory is used to store computer program code, the computer program code comprising computer instructions, and the one or more processors invoke the computer instructions to enable the electronic device to execute the omnidirectional visual positioning method shown in the foregoing.

[0207] Figure 8A structural schematic diagram of an electronic device is provided in the embodiments of the present application. The electronic device 800 can be a mobile phone, a smart screen, a tablet computer, a wearable electronic device, a vehicle-mounted electronic device, an augmented reality (AR) device, a virtual reality (VR) device, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), a projector, or a server, a memory, a base station, or a communication device, or a smart car. The embodiments of the present application do not make any limitation on the specific type of the electronic device.

[0208] The memory 801 can be used to store computer software programs 802 and modules, and the processor 803 can execute various function applications and data processing of the electronic device by running the software programs and modules stored in the memory 801. The memory 801 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application program required by a function (such as a sound playing function, an image playing function, etc.), and the like; and the data storage area can store data created according to the use of the electronic device (such as audio data, a phone book, etc.), and the like. In addition, the memory 801 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state memory device.

[0209] The processor 803 can include one or more of a central processor, an application processor (AP), a baseband processor, and the like. The processor can be the nerve center and command center of the wireless router. The processor 803 can generate operation control signals according to instruction operation codes and timing signals, and complete the control of fetching and executing instructions. The memory 801 can be used to store computer executable program codes, and the executable program codes include instructions. The processor 803 can execute various function applications and data processing of the network device by running the instructions stored in the memory. The memory 801 can include a program storage area and a data storage area, such as data of a sound signal to be played, and the like. For example, the memory can be a double data rate synchronous dynamic random access memory (DDR) or a flash memory (Flash), and the like.

[0210] The embodiments of the present application also provide a computer readable storage medium, wherein the computer readable storage medium stores computer instructions; and when the computer readable storage medium runs on an electronic device, the electronic device executes the aforementioned omnidirectional visual positioning method.

[0211] The computer instructions can be stored in or transferred from one computer-readable storage medium to another computer-readable storage medium, such as from one website, computer, server, or data center to another website, computer, server, or data center, through wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible by a computer or data storage device, such as one or more servers, data centers, etc., integrated with the medium. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium, or a semiconductor medium (such as a solid state disk (SSD)), etc.

[0212] The embodiments of the present application also provide a computer program product containing computer instructions, which, when the computer program product runs on an electronic device, enables the electronic device to execute the omnidirectional visual positioning method shown in the foregoing.

[0213] The computer storage medium and the computer program product provided by the embodiments of the present application are used to execute the method provided in the foregoing, and thus the beneficial effects that can be achieved by the computer storage medium and the computer program product can refer to the beneficial effects of the method provided in the foregoing, which will not be described herein again.

[0214] In the embodiments described above, all or part of the embodiments can be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present application are generated. The computer can be a general purpose computer, a special purpose computer, a computer network or other programmable apparatus. The computer instructions can be stored in a computer readable storage medium or transmitted from one computer readable storage medium to another computer readable storage medium, such as from a website site, computer, server or data center to another website site, computer, server or data center through wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) mode. The computer readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server, data center, etc. integrated with one or more available media. The storage medium can be a magnetic disk, an optical disk, a read-only memory (Rom), a random access memory (RAM), a flash memory, a hard disk drive (HDD) or a solid state drive (SSD), etc. The storage medium can also include a combination of the above types of memory.

[0215] In the above embodiments, the description of each embodiment has its own focus, and the parts not described or recorded in detail in a certain embodiment can be referred to the relevant description of other embodiments.

[0216] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments applied herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0217] In the embodiments provided by the present application, it should be understood that the disclosed apparatus / network device and method can be implemented in other manners. For example, the embodiments of the apparatus / network device described above are merely illustrative. For example, the division of the modules or units is merely logical function division, and there can be another division manner in actual implementation. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections between the units can be indirect couplings or communication connections through some interfaces, devices or units, and can be electrical, mechanical or in other forms.

[0218] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., can be located in one place, or can be distributed on a plurality of network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiments.

[0219] The above-described embodiments are merely used to illustrate the technical solutions of the present application, but not limit the present application; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: the technical solutions recorded in the foregoing embodiments can still be modified, or some technical features can be replaced by equivalent replacements; and these modifications or replacements do not make the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.

Claims

1. An omnidirectional visual positioning system, characterized in that, The omnidirectional visual positioning system, installed on the fuselage of the unmanned aerial vehicle, includes: Image acquisition equipment is used to acquire real-time scene images of the entire omnidirectional range during the flight of the unmanned aerial vehicle; An inertial measurement sensor, connected to the image acquisition device, is used to synchronously acquire motion state data of the unmanned aerial vehicle; The navigation system, connected to the image acquisition device and the inertial measurement sensor, is configured to perform: The real-time scene image is subjected to at least one of distortion correction processing and feature extraction processing to obtain the processed image; Based on the motion state data and the processed image, the current pose information of the unmanned aerial vehicle in the spatial coordinate system is determined, wherein the spatial coordinate system is constructed with the takeoff point of the unmanned aerial vehicle as the origin.

2. The omnidirectional visual positioning system according to claim 1, characterized in that, The image acquisition device also includes: Multiple fisheye cameras are distributed on the body in a predetermined shape. The optical axis angle between each fisheye camera is a first angle, and the field of view of a single fisheye camera is not less than a second angle. A hardware-triggered synchronization circuit is used to drive each of the fisheye cameras to acquire images synchronously, wherein the acquisition time difference between the various fisheye cameras is less than a predetermined duration.

3. The omnidirectional visual positioning system according to claim 1, characterized in that, The navigation system performs distortion correction processing on the real-time scene image and is configured to perform the following: Obtain the three-dimensional coordinates of each feature point in the real-time scene image on a normalized sphere; A cylindrical projection model is used to map the three-dimensional coordinates to planar coordinates in the cylindrical projection image, and a linear interpolation method is used to compensate for the unmapped planar coordinates. The cylindrical projection model calibrates the intrinsic parameter matrix and distortion coefficients of each fisheye camera of the image acquisition device through a sensor calibration algorithm.

4. The omnidirectional visual positioning system according to claim 1, characterized in that, The navigation system performs feature extraction processing on the real-time scene image and is configured to perform the following: Multiple feature points of the real-time scene image are extracted based on a deep learning algorithm, and a descriptor for each feature point is generated. Based on the descriptor of each feature point, a feature matching network is used to construct the matching relationship between feature points.

5. The omnidirectional visual positioning system according to claim 1, characterized in that, The navigation system, based on the motion state data and the processed image, determines the current pose information of the unmanned aerial vehicle in the spatial coordinate system and is configured to execute: The installation transformation relationship between the image acquisition device and the inertial measurement sensor, as well as the installation position relationship between the multiple fisheye cameras included in the image acquisition device, are obtained. The real-time scene image is processed by feature extraction to obtain the matching relationship between feature points; Based on the installation transformation relationship, the installation position relationship, and the matching relationship between the feature points, the depth information of the common field of view feature points of multiple fisheye cameras is obtained; Using the bundle adjustment method with feature point reprojection error as a constraint, the pose change of the current frame scene image relative to the previous frame scene image is calculated based on the depth information of the common view feature points. If the current scene matches a historical scene in the global map, the current pose information of the unmanned aerial vehicle in the spatial coordinate system is determined based on the global map and the pose change.

6. An unmanned aerial vehicle (UAV) for use in GNSS denial scenarios, characterized in that, The unmanned aerial vehicle (UAV) comprises: a fuselage and an omnidirectional visual positioning system as described in any one of claims 1 to 5; the omnidirectional visual positioning system is mounted on the fuselage of the UAV, and the omnidirectional visual positioning system comprises: Image acquisition equipment is used to acquire real-time scene images of the entire omnidirectional range during the flight of the unmanned aerial vehicle; An inertial measurement sensor, connected to the image acquisition device, is used to synchronously acquire motion state data of the unmanned aerial vehicle; The navigation system, connected to the image acquisition device and the inertial measurement sensor, is configured to perform: The real-time scene image is subjected to at least one of distortion correction processing and feature extraction processing to obtain a processed image; based on the motion state data and the processed image, the current pose information of the unmanned aerial vehicle in the spatial coordinate system is determined, wherein the spatial coordinate system is constructed with the take-off point of the unmanned aerial vehicle as the origin.

7. An omnidirectional visual positioning method, characterized in that, include: During the flight of the unmanned aerial vehicle (UAV), real-time scene images of the omnidirectional range are acquired by the image acquisition device of the UAV, and motion state data of the UAV are acquired synchronously by the inertial measurement sensor of the UAV. The real-time scene image is subjected to at least one of distortion correction processing and feature extraction processing to obtain the processed image; Based on the motion state data and the processed image, the current pose information of the unmanned aerial vehicle in the spatial coordinate system is determined, wherein the spatial coordinate system is constructed with the takeoff point of the unmanned aerial vehicle as the origin.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it causes the electronic device to perform the method as described in claim 7.

9. A computer program product, characterized in that, Includes a computer program, which, when run, causes the method of claim 7 to be performed.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in claim 7.

Citation Information

Cited By

  • Aircraft positioning method and device, aircraft, medium and program product

    CN121475158A

  • Geographic registration method, system and product for aerial image of unmanned aerial vehicle

    CN121746442A

  • Online self-calibration and SLAM method and system for multi-depth vision unmanned aerial vehicle

    CN121837385A

  • Passive geographic information superposition system and method based on cross-modal features

    CN122087022A