A Visual Fusion Localization and Navigation Method and Device
Through the visual fusion method of binocular camera and IMU sensor, feature points positions are extracted and updated, and a global map is built, which solves the accuracy and stability of mobile device positioning navigation, and achieves efficient and reliable positioning navigation effects.
Patent Information
- Application Number
- CN202510314818.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-03-18
AI Technical Summary
The prior art is difficult to provide more reliable data support for positioning and navigation of mobile devices without increasing system complexity and cost, especially because binocular cameras are affected by light and occlusion, resulting in inaccurate feature points extraction and matching, and the cumulative error of IMU sensors affects the positioning and navigation accuracy.
The binocular camera collects environmental images, extracts feature points and their descriptors, combines the acceleration and angular velocity information of the IMU sensor, predicts the position of the binocular camera, updates the position of the feature points, and fuses them with three-dimensional point cloud data to build a global map, determines the location of the mobile device and plans the path.
Improves the accuracy and stability of mobile device positioning navigation, makes full use of existing sensor resources, no additional hardware required, and reduces system costs and complexity.
Smart Images

Figure CN119860766B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of positioning and navigation, and particularly to a visual fusion positioning and navigation method and device. Background Art
[0002] In the field of positioning and navigation, binocular cameras and IMU sensors are two important data sources. Binocular cameras can collect three-dimensional information of the environment, providing a rich data source for positioning and navigation. The IMU sensor can obtain the attitude information of the device in real time, which helps to maintain the stability of the system when the image information is missing or unreliable.
[0003] However, binocular cameras are easily affected by factors such as light and occlusion when collecting images, resulting in inaccurate feature point extraction and matching. The IMU sensor will generate cumulative errors after long-term use, affecting the accuracy of positioning and navigation.
[0004] To solve these problems, the prior art corrects the cumulative error of the IMU by fusing data from other sensors. However, the prior art requires additional hardware support, increasing the complexity and cost of the system.
[0005] Therefore, how to provide more reliable data support for the positioning and navigation of mobile devices without increasing the complexity and cost of the system to improve the accuracy of positioning and navigation has become an urgent problem to be solved. Summary of the Invention
[0006] In the embodiments of the present application, by providing a visual fusion positioning and navigation method and device, the problem of how to provide more reliable data support for the positioning and navigation of mobile devices without increasing the complexity and cost of the system to improve the accuracy of positioning and navigation is solved.
[0007] In a first aspect, an embodiment of the present application provides a visual fusion positioning and navigation method, which includes: using a binocular camera mounted on a mobile device to collect environmental images, and preprocessing the collected images; extracting feature points and their descriptors from the preprocessed images; using the descriptors to perform feature point matching between the left and right images of the binocular camera, and restoring the three-dimensional point cloud data of the environment according to the matching result; reading in real time the IMU data obtained by the IMU sensor mounted on the mobile device, including acceleration and angular velocity information; using the IMU data to predict the pose of the binocular camera, and updating the feature point positions of the current frame according to the prediction result; fusing the updated feature point positions with the three-dimensional point cloud data to construct a global map; determining the current position of the mobile device according to the global map and the current pose of the binocular camera, planning a path according to the target position and the current position, and performing positioning and navigation of the mobile device according to the planned path.
[0008] In a possible implementation, the preprocessing of the acquired image includes: obtaining the grayscale histogram of the acquired image; calculating the mean value of the grayscale histogram, and if the mean value is less than the first preset threshold, determining that the image is too dark; calculating the standard deviation of the grayscale histogram, and if the standard deviation is less than the second preset threshold, determining that the image has insufficient contrast; if it is determined that the image is too dark or has insufficient contrast, adjust the brightness or contrast of the image.
[0009] In a possible implementation, the extraction of feature points and their descriptors from the preprocessed image includes: calculating the gradient magnitude and direction of the preprocessed image to obtain gradient information; calculating the corner response value of each pixel point of the image according to the gradient information; regarding the pixel points with corner response values greater than the third preset threshold as potential feature points; only retaining the feature point with the largest corner response value in the local area as the target feature point; for each target feature point, select a fixed-size image block around it as the neighborhood; randomly select the first preset number of pixel points in the neighborhood; sample each pixel point to obtain sampling points; record the index of the direction interval to which the gradient direction of each sampling point belongs and the relative magnitude of the gradient magnitude; construct a descriptor vector with a fixed length to extract the descriptor; where each element of the descriptor vector corresponds to the accumulation of the gradient magnitudes of the sampling points in a direction interval.
[0010] In a possible implementation, using the IMU data to predict the pose of the binocular camera and updating the position of the feature points in the current frame according to the prediction result includes: defining the pose of the binocular camera and constructing a dynamic model to describe the change of the pose of the binocular camera over time; discretizing the dynamic model using the fourth-order Runge-Kutta method to predict the pose of the binocular camera at the next moment and obtain the prediction result; using three-dimensional point back-projection to transform the feature points in the current frame back into points in three-dimensional space; using the prediction result and coordinate transformation to transform the points in three-dimensional space from the camera coordinate system at the current moment to the camera coordinate system at the next moment; using the reprojection equation to project the transformed points in three-dimensional space back onto the two-dimensional image plane to obtain the updated feature point coordinates to update the position of the feature points in the current frame.
[0011] In a possible implementation, the pose of the binocular camera is defined as: ; where is the pose of the binocular camera at the current moment, used to describe the rigid body transformation in three-dimensional space, is the set of all rigid body transformations, is 's three-dimensional rotation orthogonal matrix, is 's column vector, representing the position of the origin of the camera coordinate system in the world coordinate system, is The row vector, is the homogeneous term of homogeneous coordinates; the dynamic model is: ; where, is the time derivative of the pose of the binocular camera at the current moment on indicating the rate of change of the pose of the binocular camera with time, is the extended element, , is the angular velocity skew-symmetric matrix, , , is the three-dimensional acceleration vector, is the angular velocity on the axis component, is the angular velocity on the axis component; the expression for discretizing the dynamic model using the fourth-order Runge-Kutta method is: ; where, is the predicted pose of the binocular camera at the next moment, which is used as the prediction result, is the time interval from the current moment to the next moment, function is the matrix exponential, , , and are the slope terms, , , , , is the extended element of the binocular camera at the current moment, is the extended element of the binocular camera at the intermediate moment, is the extended element of the binocular camera before the next moment; the expression for three-dimensional point back-projection is: ; where, is the point in three-dimensional space, is the back-projection function, is the coordinate of the feature point in the current frame, is the depth value corresponding to the feature point, is the inverse matrix of the internal parameters of the binocular camera, is to expand the two-dimensional coordinates into homogeneous coordinate form; the expression for coordinate transformation is: ; where, is the point in the transformed three-dimensional space, is the pose of the binocular camera at the predicted next moment, is the inverse matrix of the pose of the binocular camera at the current moment; the reprojection equation is: ; where, are the coordinates of the feature points of the updated current frame, is the projection function, are the first three components of the point in the transformed three-dimensional space, is the homogeneous coordinate component of the point in the transformed three-dimensional space, is the internal parameter matrix of the binocular camera.
[0012] In a possible implementation manner, the fusing the updated feature point positions with the three-dimensional point cloud data to construct a global map includes: extracting descriptors for the feature points of each updated frame; generating a global point cloud map using the three-dimensional point cloud data; in the global point cloud map, searching for the second preset number of candidate points closest to each extracted descriptor; calculating the similarity between each extracted descriptor and the descriptors of the candidate points, and retaining the matching pairs with the similarity less than the preset similarity threshold; the calculation formula for the similarity is: ; where, is the similarity between the extracted descriptor and the descriptors of the candidate points, is the extracted descriptor, are the descriptors of the candidate points, is the Euclidean distance; the expression for projecting the candidate points from the three-dimensional space to the two-dimensional image plane of the current frame using the pose of the binocular camera at the predicted next moment is: ; where, are the candidate points, is the position after the candidate points are projected to the current frame through the pose of the binocular camera at the predicted next moment; calculating the reprojection error, and removing the abnormal matching pairs with the reprojection error greater than the pixel error threshold; the calculation formula for the reprojection error is: ; where, is the reprojection error, are the coordinates of the feature points of the updated current frame; fusing the remaining matching pairs with the candidate points in the corresponding three-dimensional point cloud data to construct a global map.
[0013] In a possible implementation, determining the current position of the mobile device based on the global map and the current pose of the binocular camera, planning a path based on the target position and the current position, and performing positioning and navigation of the mobile device according to the planned path includes: dividing the global map into multiple grid cells, where each grid cell represents a spatial position that the mobile device can occupy; marking the impassable grid cells according to the positions of the obstacles; the goal of path planning is to find a path that can bypass all impassable grid cells and minimizes the sum of the Euclidean distances between adjacent grid cells passed through.
[0014] In a second aspect, an embodiment of the present application provides a visual fusion positioning and navigation device, including: a preprocessing module for collecting environmental images using a binocular camera mounted on the mobile device and preprocessing the collected images; an extraction module for extracting feature points and their descriptors from the preprocessed images; a matching module for performing feature point matching between the left and right images of the binocular camera using the descriptors and restoring the three-dimensional point cloud data of the environment according to the matching results; a reading module for reading in real time the IMU data obtained by an IMU sensor mounted on the mobile device, including acceleration and angular velocity information; a prediction module for predicting the pose of the binocular camera using the IMU data and updating the positions of the feature points in the current frame according to the prediction results; a construction module for fusing the updated positions of the feature points with the three-dimensional point cloud data to construct a global map; a positioning and navigation module for determining the current position of the mobile device based on the global map and the current pose of the binocular camera, planning a path based on the target position and the current position, and performing positioning and navigation of the mobile device according to the planned path.
[0015] In a third aspect, an embodiment of the present application provides a visual fusion positioning and navigation server, including a memory and a processor; the memory is used for storing computer-executable instructions; the processor is used for executing the computer-executable instructions to implement the method described in the first aspect or any possible implementation manner of the first aspect.
[0016] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium storing executable instructions, and when a computer executes the executable instructions, it can implement the method described in the first aspect or any possible implementation manner of the first aspect.
[0017] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects:
[0018] The embodiment of the present application provides a visual fusion positioning and navigation method. By collecting environmental images through a binocular camera and restoring three-dimensional point cloud data, combined with the real-time acceleration and angular velocity information provided by the IMU sensor, the positioning of the mobile device can be achieved. Binocular vision provides rich environmental information, while IMU data helps to maintain the positioning accuracy and stability of the system when the image information is missing or unstable, thereby improving the overall navigation performance. The IMU data is used to predict the pose of the binocular camera, and the feature point positions of the current frame are updated according to the prediction results, so as to reduce the cumulative error caused by image matching errors. The present application makes full use of the existing binocular camera and IMU sensor resources on the mobile device, without the need for additional high-precision sensors or complex devices, thereby reducing the system cost and complexity. Whether in an indoor or outdoor environment, the present application can fuse the updated feature point positions with the three-dimensional point cloud data to construct a global map, determine the current position of the mobile device, and perform path planning based on the target position and the current position. It solves the problem of how to provide more reliable data support for the positioning and navigation of mobile devices without increasing the system complexity and cost, so as to improve the positioning and navigation accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments of the present application or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings.
[0020] Figure 1 It is a flowchart of a visual fusion positioning and navigation method provided by an embodiment of the present application;
[0021] Figure 2 It is a schematic diagram of a visual fusion positioning and navigation device provided by an embodiment of the present application;
[0022] Figure 3 It is a schematic diagram of a visual fusion positioning and navigation server provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] The following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0024] The following describes some of the technologies involved in the embodiments of the present application to facilitate understanding. It should be considered that they are merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present application. Similarly, for the sake of clarity and conciseness, some descriptions of well-known functions and structures are omitted in the following description.
[0025] The embodiments of the present application provide a visual fusion positioning and navigation method. As Figure 1 shown, the method includes steps S101 to S107. Among them, Figure 1 This is only an execution order shown in the embodiments of the present application and does not represent the only execution order of a visual fusion positioning and navigation method. In the case where the final result can be achieved, Figure 1 the steps shown can be executed in parallel or reversed.
[0026] S101: Use the binocular camera mounted on the mobile device to collect environmental images and preprocess the collected images.
[0027] It should be noted that the mobile device of the present application can be an unmanned robot and a navigation unmanned vehicle.
[0028] Preprocessing the collected images includes: obtaining the grayscale histogram of the collected images. Calculate the mean value of the grayscale histogram. If the mean value is less than the first preset threshold, it is judged that the image is too dark. Calculate the standard deviation of the grayscale histogram. If the standard deviation is less than the second preset threshold, it is judged that the image has insufficient contrast. If it is judged that the image is too dark or has insufficient contrast, adjust the brightness or contrast of the image.
[0029] Specifically, the grayscale histogram is an array of size 256, and the subscripts of the array range from 0 to 255, corresponding to each grayscale level from 0 to 255 in the grayscale image respectively. By traversing each pixel point of the image, the number of occurrences of each grayscale level in the image is counted to obtain the grayscale histogram. The mean value of the grayscale histogram reflects the overall brightness level of the image. The mean value calculation method is the sum of the products of the number of occurrences of all grayscale levels and their grayscale values divided by the total number of pixels. The standard deviation of the grayscale histogram reflects the distribution dispersion degree of the grayscale levels in the image, that is, the contrast. The standard deviation calculation method is the square root of the sum of the products of the number of occurrences of each grayscale level and the square of the difference between its grayscale value and the mean value divided by the square root of the total number of pixels. It should be noted that the first preset threshold can be 128, and the second preset threshold can be 30. In practical applications, the first preset threshold and the second preset threshold can also be flexibly adjusted according to specific situations.
[0030] Further, if it is determined that the image is too dark, it can be adjusted by increasing the brightness value of the image. The specific method is to add a brightness increment to the grayscale value of each pixel in the image. If it is determined that the image has insufficient contrast, it can be improved by adjusting the contrast of the image. The specific method is to multiply the grayscale value of each pixel in the image by a contrast coefficient. When the contrast coefficient is greater than 1, the contrast can be enhanced. It should be noted that the brightness increment can be 30 and the contrast coefficient can be 1.2. In practical applications, the brightness increment and the contrast coefficient can also be flexibly adjusted according to specific situations.
[0031] S102: Extract feature points and their descriptors from the preprocessed image.
[0032] Extract feature points and their descriptors from the preprocessed image, including: calculating the gradient magnitude and direction of the preprocessed image to obtain gradient information. Calculating the corner response value of each pixel point in the image according to the gradient information. Considering the pixel points with corner response values greater than the third preset threshold as potential feature points. Only retaining the feature point with the largest corner response value in the local area as the target feature point. For each target feature point, selecting an image patch of a fixed size around it as the neighborhood. Randomly selecting the first preset number of pixel points in the neighborhood. Sampling each pixel point to obtain sampling points. Recording the index of the direction interval to which the gradient direction of each sampling point belongs and the relative magnitude of the gradient magnitude. Constructing a descriptor vector with a fixed length to extract the descriptor. Each element of the descriptor vector corresponds to the accumulation of the gradient magnitudes of the sampling points in a direction interval.
[0033] Specifically, a differential operator can be used to perform a convolution operation on the preprocessed image to calculate the gradient magnitude and direction of the preprocessed image, thereby obtaining gradient information. The gradient magnitude reflects the speed of change of pixel values in the image, while the gradient direction indicates the direction of change. The differential operator can be the Sobel operator. The corner response value reflects the degree of change in brightness around a pixel point and is an important basis for identifying corners. Pixel points with a corner response value greater than the third preset threshold are regarded as potential feature points. This step helps to exclude those pixel points whose degree of change is not sufficient to be regarded as feature points. In a local area, only the feature point with the largest corner response value is retained as the target feature point. This step can ensure that each local area contains only one most representative feature point, thereby reducing the number of feature points and improving the accuracy of matching. Randomly selecting the first preset number of pixel points for sampling in the neighborhood helps to reduce the influence of noise and outliers and improve the robustness of the descriptor. Random sampling also introduces a certain degree of randomness, making the descriptor less sensitive to minor image changes to a certain extent. Since the descriptor is based on the accumulation of gradient magnitudes of the sampled points, even if the local brightness of the image changes, as long as the relative change between the gray values remains unchanged, the gradient magnitude will not be affected, and it can resist local brightness changes in the image to a certain extent. The relative magnitude of the gradient magnitude can be achieved by normalizing the gradient magnitude. In practical applications, appropriate parameters need to be selected according to the characteristics of the image and processing requirements, such as the third preset threshold, neighborhood size, and the first preset number, etc.
[0034] S103: Use the descriptor to perform feature point matching between the left and right images of the binocular camera, and restore the three-dimensional point cloud data of the environment according to the matching result.
[0035] Specifically, the matching process aims to find the best corresponding points in the right image for each feature point in the left image, and these corresponding points should reflect the projections of the same physical point under two different perspectives. Combining the three-dimensional coordinates of all the matching feature points can form the three-dimensional point cloud data for restoring the environment.
[0036] S104: Read in real-time the IMU data obtained by the IMU sensor mounted on the mobile device, including acceleration and angular velocity information.
[0037] Specifically, the IMU sensor is an inertial measurement unit sensor, and the IMU data is inertial measurement unit data.
[0038] S105: Use the IMU data to predict the pose of the binocular camera, and update the positions of the feature points in the current frame according to the prediction result.
[0039] Predict the pose of the binocular camera using IMU data, and update the position of the feature points in the current frame according to the prediction result, including: Define the pose of the binocular camera, and construct a dynamic model to describe the change of the pose of the binocular camera over time. Discretize the dynamic model using the fourth-order Runge-Kutta method to predict the pose of the binocular camera at the next moment and obtain the prediction result. Use three-dimensional point back-projection to transform the feature points in the current frame back into points in three-dimensional space. Utilize the prediction result and use coordinate transformation to transform the points in three-dimensional space from the camera coordinate system at the current moment to the camera coordinate system at the next moment. Use the reprojection equation to project the transformed points in three-dimensional space back onto the two-dimensional image plane to obtain the updated feature point coordinates to update the position of the feature points in the current frame.
[0040] Define the pose of the binocular camera as: . Among them, is the pose of the binocular camera at the current moment, used to describe the rigid body transformation in three-dimensional space, is the set of all rigid body transformations, is 's three-dimensional rotation orthogonal matrix, is 's column vector, representing the position of the origin of the camera coordinate system in the world coordinate system, is 's row vector, is the homogeneous term of the homogeneous coordinate.
[0041] The dynamic model is: . Among them, is the time derivative of the pose of the binocular camera at the current moment on , representing the rate of change of the pose of the binocular camera over time, is the extended element, , is the angular velocity 's skew-symmetric matrix, , , is the three-dimensional acceleration vector, is the angular velocity on the axis component, is the angular velocity on the axis component.
[0042] The expression for discretizing the dynamic model using the fourth-order Runge-Kutta method is: . Among them, is the predicted pose of the binocular camera at the next moment, which is used as the prediction result, is the time interval from the current moment to the next moment, The function is the matrix exponential, , , and are the slope terms, , , , , is the extended element of the binocular camera at the current moment, is the extended element of the binocular camera at the intermediate moment, is the extended element of the binocular camera before the next moment.
[0043] The expression for the back-projection of a 3D point is: . Where, is the point in 3D space, is the back-projection function, is the coordinate of the feature point in the current frame, is the depth value corresponding to the feature point, is the inverse matrix of the internal parameter matrix of the binocular camera, is to extend the 2D coordinate to the homogeneous coordinate form.
[0044] The expression for the coordinate transformation is: . Where, is the point in the transformed 3D space, is the pose of the binocular camera at the predicted next moment, is the inverse matrix of the pose of the binocular camera at the current moment.
[0045] The reprojection equation is: . Where, is the updated coordinate of the feature point in the current frame, is the projection function, are the first three components of the point in the transformed 3D space, is the homogeneous coordinate component of the point in the transformed 3D space, is the internal parameter matrix of the binocular camera.
[0046] Specifically, by combining IMU data and visual information from a binocular camera, the present application can more accurately predict the movement trajectory of the camera, thereby improving the positioning accuracy of the mobile device. In complex or dynamic environments, a single sensor may be affected by factors such as noise, occlusion, or light changes, resulting in positioning failures. Combining IMU data and a binocular camera can complement each other and enhance the robustness of the system. By using numerical methods such as the fourth-order Runge-Kutta method, the dynamic model equations can be efficiently solved, thereby optimizing the computational efficiency while ensuring accuracy.
[0047] S106: Fuse the updated feature point positions with the three-dimensional point cloud data to construct a global map.
[0048] Fusing the updated feature point positions with the three-dimensional point cloud data to construct a global map includes: extracting descriptors for each frame of the updated feature points. Generating a global point cloud map using the three-dimensional point cloud data. In the global point cloud map, for each extracted descriptor, search for the second preset number of candidate points closest to it. Calculate the similarity between each extracted descriptor and the descriptors of the candidate points, and retain the matching pairs with similarity less than the preset similarity threshold. The formula for calculating similarity is: . Where is the similarity between the extracted descriptor and the descriptor of the candidate point, is the extracted descriptor, is the descriptor of the candidate point, is the Euclidean distance. The expression for projecting the candidate points from the three-dimensional space to the two-dimensional image plane of the current frame using the predicted pose of the binocular camera at the next moment is: . Where is the candidate point, is the position of the candidate point after projection to the current frame through the predicted pose of the binocular camera at the next moment. Calculate the reprojection error and eliminate the abnormal matching pairs with reprojection error greater than the pixel error threshold. The formula for calculating the reprojection error is: . Where is the reprojection error, are the coordinates of the updated feature points in the current frame. Fuse the remaining matching pairs with the candidate points in the corresponding three-dimensional point cloud data to construct a global map.
[0049] Specifically, in the global point cloud map, for each extracted descriptor, searching for the second preset number of candidate points closest to it can be achieved through a nearest neighbor search algorithm (such as a KD tree) to efficiently find the points most similar to the descriptor. To retain reliable matching pairs, a preset similarity threshold is set. Only when the similarity between descriptors is less than this threshold are they considered a valid matching pair. This step helps to eliminate unreliable matches caused by noise, occlusion, or incorrect matching. To further verify the validity of the match, using the predicted pose of the binocular camera at the next moment, project the candidate points from the three-dimensional space onto the two-dimensional image plane of the current frame, and then calculate the reprojection error. If the reprojection error is greater than the pixel error threshold, the matching pair is considered abnormal and excluded. Fuse the remaining matching pairs with the candidate points in the corresponding three-dimensional point cloud data. This step involves adding new points in the three-dimensional space to the global point cloud map or updating the positions of existing feature points. By continuously repeating this process, a complete, accurate, and real-time global map can be gradually constructed.
[0050] It should be noted that in this application, the second preset number, the preset similarity threshold, and the pixel error threshold can all be flexibly set and adjusted according to the actual situation. For the setting of the second preset number, 50 can be selected as the number of candidate points. For the setting of the preset similarity threshold, 0.2 can be selected as the threshold. For the setting of the pixel error threshold, 2 pixels can be selected as the threshold.
[0051] S107: Determine the current position of the mobile device according to the global map and the current pose of the binocular camera, plan a path based on the target position and the current position, and perform positioning and navigation of the mobile device according to the planned path.
[0052] Determine the current position of the mobile device according to the global map and the current pose of the binocular camera, plan a path based on the target position and the current position, and perform positioning and navigation of the mobile device according to the planned path, including: Divide the global map into multiple grid cells, and each grid cell represents a spatial position that the mobile device can occupy. Mark the impassable grid cells according to the positions of the obstacles. The goal of path planning is to find a path that can bypass all impassable grid cells and minimize the sum of the Euclidean distances between the adjacent grid cells passed through.
[0053] Specifically, the size of the grid cells can be set according to the size of the mobile device. The division of the grid cells can transform complex map information into discrete data that is easy to process. According to the positions of the obstacles, the impassable grid cells are marked, and these grid cells will be regarded as obstacles during the path planning process and need to be bypassed. Each grid can be regarded as an area with a specific size and shape, and the midpoint coordinates of the grid are used to represent the position of the grid. During the path planning process, the Euclidean distance between the midpoints of adjacent grids is calculated, and a path is searched for to minimize the sum of these distances. Such a path will be able to effectively avoid obstacles while ensuring that the mobile device moves from the starting point to the ending point at the shortest distance.
[0054] An embodiment of this application also provides a visual fusion positioning and navigation device 200, as Figure 2 shown. The device includes: a preprocessing module 201, an extraction module 202, a matching module 203, a reading module 204, a prediction module 205, a construction module 206, and a positioning and navigation module 207.
[0055] The preprocessing module 201 is used to collect environmental images using the binocular camera mounted on the mobile device and preprocess the collected images.
[0056] The extraction module 202 is used to extract feature points and their descriptors from the preprocessed images.
[0057] The matching module 203 is used to perform feature point matching between the left and right images of the binocular camera using the descriptors and restore the three-dimensional point cloud data of the environment according to the matching results.
[0058] The reading module 204 is used to read in real time the IMU data obtained by the IMU sensor mounted on the mobile device, including acceleration and angular velocity information.
[0059] The prediction module 205 is used to predict the pose of the binocular camera using the IMU data and update the positions of the feature points in the current frame according to the prediction results.
[0060] The construction module 206 is used to fuse the updated positions of the feature points with the three-dimensional point cloud data to construct a global map.
[0061] The positioning and navigation module 207 is used to determine the current position of the mobile device according to the global map and the current pose of the binocular camera, and plan a path according to the target position and the current position, and perform positioning and navigation of the mobile device according to the planned path.
[0062] Some of the modules in the device described in this application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, classes, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0063] The devices or modules illustrated in the above application embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. For the convenience of description, when describing the above devices, they are divided into various modules according to functions and described separately. When implementing the embodiments of this application, the functions of each module can be implemented in the same or multiple software and / or hardware. Of course, the module that implements a certain function can also be implemented by combining multiple sub-modules or sub-units.
[0064] The methods, devices or modules described in this application can be implemented in the form of computer-readable program code. The controller can be implemented in any appropriate manner. For example, the controller can take the form of a microprocessor or a processor, and a computer-readable medium that stores computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, application specific integrated circuit (ASIC), programmable logic controller, and embedded microcontroller. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, the method steps can be logically programmed to enable the controller to be implemented in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, and embedded microcontrollers to achieve the same function. Therefore, such a controller can be regarded as a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as both software modules for implementing the method and the structures within the hardware component.
[0065] As Figure 3As shown in the figure, an embodiment of the present application further provides a visual fusion positioning and navigation server, including a memory 301 and a processor 302; the memory 301 is used to store computer-executable instructions; the processor 302 is used to execute the computer-executable instructions to implement a visual fusion positioning and navigation method described above in the embodiments of the present application.
[0066] An embodiment of the present application further provides a computer-readable storage medium, which stores executable instructions. When a computer executes the executable instructions, it can implement a visual fusion positioning and navigation method described above in the embodiments of the present application.
[0067] From the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary hardware. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product, or can also be reflected in the implementation process of data migration. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a mobile terminal, a server, or a network device, etc.) to execute the method described in the embodiments of the present application.
[0068] The various embodiments in this specification are described in a progressive manner. For the same or similar parts between the various embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. All or part of the present application can be used in many general or special computer system environments or configurations.
[0069] The above embodiments are only used to illustrate the technical solutions of the present application, rather than limiting the present application; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the present application.
Claims
1. A visual fusion positioning and navigation method, characterized in that, Including: Collecting environmental images using the binocular camera mounted on the mobile device and preprocessing the collected images; Extracting feature points and their descriptors from the preprocessed images; Using the descriptors to perform feature point matching between the left and right images of the binocular camera, and restoring the three-dimensional point cloud data of the environment according to the matching results; Realtime reading the IMU data obtained by the IMU sensor mounted on the mobile device, including acceleration and angular velocity information; Predicting the pose of the binocular camera using the IMU data, and updating the feature point positions of the current frame according to the prediction results; The predicting the pose of the binocular camera using the IMU data and updating the feature point positions of the current frame according to the prediction results includes: defining the pose of the binocular camera and constructing a dynamic model to describe the change of the pose of the binocular camera over time; discretizing the dynamic model using the fourth-order Runge-Kutta method to predict the pose of the binocular camera at the next moment and obtaining the prediction results; using three-dimensional point back-projection to transform the feature points of the current frame back into points in three-dimensional space; using the prediction results and coordinate system transformation to transform the points in three-dimensional space from the camera coordinate system at the current moment to the camera coordinate system at the next moment; using the reprojection equation to project the transformed points in three-dimensional space back onto the two-dimensional image plane to obtain the updated feature point coordinates, so as to update the feature point positions of the current frame; Define the pose of the binocular camera as: ; where is the pose of the binocular camera at the current moment, used to describe the rigid body transformation in three-dimensional space, is the set of all rigid body transformations, is 's three-dimensional rotation orthogonal matrix, is 's column vector, representing the position of the origin of the camera coordinate system in the world coordinate system, is 's row vector, is the homogeneous term of homogeneous coordinates; The dynamic model is: ; where is the time derivative of the pose of the binocular camera at the current moment on , representing the rate of change of the pose of the binocular camera with time, is the extended element, , is the skew-symmetric matrix of the angular velocity , , , is the three-dimensional acceleration vector, is the angular velocity on the axis component, is the angular velocity on the axis component; The expression for discretizing the dynamic model using the fourth-order Runge-Kutta method is: ; where is the predicted pose of the binocular camera at the next moment, which is used as the prediction result, is the time interval from the current moment to the next moment, function is the matrix exponential, , , and are the slope terms, , , , , is the extended element of the binocular camera at the current moment, is the extended element of the binocular camera at the intermediate moment, is the extended element of the binocular camera before the next moment; The expression for three-dimensional point back-projection is: ; where is the point in three-dimensional space, is the back-projection function, are the coordinates of the feature points in the current frame, is the depth value corresponding to the feature points, is the inverse matrix of the intrinsic parameters of the binocular camera, is to expand the two-dimensional coordinates into the homogeneous coordinate form; the expression of the coordinate system transformation is: ; where, is the point in the transformed three-dimensional space, is the pose of the binocular camera at the predicted next moment, is the inverse matrix of the pose of the binocular camera at the current moment; the reprojection equation is: ; where, are the updated coordinates of the feature points in the current frame, is the projection function, are the first three components of the point in the transformed three-dimensional space, is the homogeneous coordinate component of the point in the transformed three-dimensional space, is the intrinsic parameter matrix of the binocular camera; Fusing the updated feature point positions with the three-dimensional point cloud data to construct a global map; Determining the current position of the mobile device according to the global map and the current pose of the binocular camera, planning a path according to the target position and the current position, and performing positioning and navigation of the mobile device according to the planned path.
2. The visual fusion positioning and navigation method according to claim 1, wherein The preprocessing the collected images includes: Obtaining the grayscale histogram of the collected images; Calculating the mean value of the grayscale histogram, and if the mean value is less than the first preset threshold, determining that the image is too dark; Calculating the standard deviation of the grayscale histogram, and if the standard deviation is less than the second preset threshold, determining that the image has insufficient contrast; If it is determined that the image is too dark or has insufficient contrast, adjusting the brightness or contrast of the image.
3. The visual fusion positioning and navigation method according to claim 1, wherein The extracting feature points and their descriptors from the preprocessed images includes: Calculating the gradient magnitude and direction of the preprocessed images to obtain gradient information; Calculating the corner response value of each pixel point of the image according to the gradient information; Regarding the pixel points with corner response values greater than the third preset threshold as potential feature points; Only retaining the feature points with the largest corner response value in the local area as the target feature points; For each target feature point, selecting an image block with a fixed size around it as the neighborhood; Randomly selecting the first preset number of pixel points in the neighborhood; Sampling for each pixel point to obtain sampling points; Recording the index of the direction interval to which the gradient direction of each sampling point belongs and the relative magnitude of the gradient magnitude; Constructing a descriptor vector with a fixed length to extract descriptors; wherein, each element of the descriptor vector corresponds to the accumulation of the gradient magnitudes of the sampling points in a direction interval.
4. The visual fusion positioning and navigation method according to claim 1, wherein The fusing the updated feature point positions with the three-dimensional point cloud data to construct a global map includes: Extract descriptors for the feature points of each updated frame; Generate a global point cloud map using the 3D point cloud data; In the global point cloud map, for each extracted descriptor, search for the second preset number of candidate points closest to it; Calculate the similarity between each extracted descriptor and the descriptors of the candidate points, and retain the matching pairs with similarity less than the preset similarity threshold; The calculation formula for similarity is as follows: ; where is the similarity between the extracted descriptor and the descriptor of the candidate point, is the extracted descriptor, is the descriptor of the candidate point, is the Euclidean distance; The expression for projecting candidate points from three-dimensional space onto the two-dimensional image plane of the current frame using the predicted pose of the binocular camera at the next moment is: ; where is the candidate point, is the position after the candidate point is projected onto the current frame through the predicted pose of the binocular camera at the next moment; Calculate the reprojection error, and eliminate the abnormal matching pairs with reprojection error greater than the pixel error threshold; The calculation formula of the reprojection error is as follows: ; where is the reprojection error, is the coordinate of the feature point of the updated current frame; Fuse the retained matching pairs with the corresponding candidate points in the 3D point cloud data to construct a global map.
5. The visual fusion positioning and navigation method according to claim 4, wherein Determine the current position of the mobile device according to the global map and the current pose of the binocular camera, and plan a path based on the target position and the current position, and perform positioning and navigation of the mobile device according to the planned path, including: Divide the global map into multiple grid cells, and each grid cell represents a spatial position that the mobile device can occupy; Mark the impassable grid cells according to the positions of the obstacles; The goal of path planning is to find a path that can bypass all impassable grid cells and minimize the sum of the Euclidean distances between adjacent grid cells passed through.
6. A visual fusion positioning and navigation device, characterized in that, Including: A preprocessing module for collecting environmental images using the binocular camera mounted on the mobile device and preprocessing the collected images; An extraction module for extracting feature points and their descriptors from the preprocessed images; A matching module for performing feature point matching between the left and right images of the binocular camera using the descriptors and restoring the 3D point cloud data of the environment according to the matching results; A reading module for real-time reading of the IMU data obtained by the IMU sensor mounted on the mobile device, including acceleration and angular velocity information; A prediction module for predicting the pose of the binocular camera using the IMU data and updating the positions of the feature points of the current frame according to the prediction results; Predicting the pose of the binocular camera using IMU data and updating the position of the feature points in the current frame according to the prediction results, including: defining the pose of the binocular camera and constructing a dynamic model to describe the change of the pose of the binocular camera over time; discretizing the dynamic model using the fourth-order Runge-Kutta method to predict the pose of the binocular camera at the next moment and obtaining the prediction result; using three-dimensional point back-projection to transform the feature points in the current frame back into points in three-dimensional space; using the prediction result and coordinate system transformation to transform the points in three-dimensional space from the camera coordinate system at the current moment to the camera coordinate system at the next moment; using the reprojection equation to project the transformed points in three-dimensional space back onto the two-dimensional image plane to obtain the updated feature point coordinates to update the position of the feature points in the current frame; defining the pose of the binocular camera as: ; where is the pose of the binocular camera at the current moment, used to describe the rigid body transformation in three-dimensional space, is the set of all rigid body transformations, is the three-dimensional rotation orthogonal matrix of is the column vector of is the row vector of is the homogeneous term of the homogeneous coordinates; the dynamic model is: ; where is the time derivative of the pose of the binocular camera at the current moment on indicating the rate of change of the pose of the binocular camera over time, is the extended element, , is the angular velocity the skew-symmetric matrix of , , is the three-dimensional acceleration vector, is the angular velocity on the axis component, is the angular velocity on the axis component; the expression for discretizing the dynamic model using the fourth-order Runge-Kutta method is: ; where is the predicted pose of the binocular camera at the next moment, which is used as the prediction result, is the time interval from the current moment to the next moment, function is the matrix exponential, , , and is the slope term, , , , , is the extended element of the binocular camera at the current moment element, is the extended element of the binocular camera at the intermediate moment element, is the extended element of the binocular camera before the next moment; the expression of the 3D point back-projection is: ; where, ; is the point in 3D space, is the back-projection function, is the coordinate of the feature point in the current frame, is the depth value corresponding to the feature point, is the inverse matrix of the internal parameters of the binocular camera, is to expand the 2D coordinate into the homogeneous coordinate form; the expression of the coordinate transformation is: ; where, is the point in the transformed 3D space, is the pose of the binocular camera at the predicted next moment, is the inverse matrix of the pose of the binocular camera at the current moment; the reprojection equation is: ; where, is the updated coordinate of the feature point in the current frame, is the projection function, are the first three components of the point in the transformed 3D space, is the homogeneous coordinate component of the point in the transformed 3D space, is the internal parameter matrix of the binocular camera; A construction module for fusing the updated feature point positions with the 3D point cloud data to construct a global map; A positioning and navigation module for determining the current position of the mobile device according to the global map and the current pose of the binocular camera, and planning a path based on the target position and the current position, and performing positioning and navigation of the mobile device according to the planned path.
7. A visual fusion positioning and navigation server, characterized in that, Including a memory and a processor; The memory is used to store computer-executable instructions; The processor is used to execute the computer-executable instructions to implement the method according to any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores executable instructions, and when the computer executes the executable instructions, it can implement the method according to any one of claims 1-5.
Citation Information
Patent Citations
Tightly coupled binocular vision-inertial SLAM method using combined point-line features
CN109579840A
Indoor unmanned aerial vehicle positioning navigation path planning method and system
CN119290000A