Methods, apparatus, storage media, and electronic devices for determining positioning modes
By performing image quality detection and adaptive positioning mode switching in front of the visual inertial odometry, the problem of inaccurate positioning of binocular visual inertial odometry when the image is abnormal is solved, ensuring the safety and accurate positioning of unmanned equipment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING SANKUAI ONLINE TECH CO LTD
- Filing Date
- 2022-07-22
- Publication Date
- 2026-07-17
AI Technical Summary
When binocular visual inertial odometry is abnormal in either monocular or binocular images, the positioning effect is poor or even impossible, leading to safety risks for unmanned equipment.
Before visual inertial odometry pose positioning calculation, the image quality is checked, and the target positioning mode is determined based on the quality check results and type of monocular or binocular images. The positioning mode is adaptively switched to ensure accuracy.
It effectively avoids positioning failures caused by anomalies in monocular images and improves positioning accuracy when binocular images are abnormal, thus ensuring the safety and positioning accuracy of unmanned equipment.
Smart Images

Figure CN117474987B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of visual inertial positioning technology, and more specifically, to a method, apparatus, storage medium, and electronic device for determining a positioning mode. Background Technology
[0002] Visual-Inertial Odometry (VIO) algorithms, based on camera and Inertial Measurement Unit (IMU) data, utilize small, lightweight cameras to provide rich environmental information to assist high-frequency IMUs in pose measurement, effectively improving the accuracy of pose estimation affected by noise and bias. Specifically, binocular visual-inertial odometry (Binocular VIO) uses a pair of simultaneously exposed camera images and IMU data as input, addressing the limitation of monocular VIO in feature point triangulation when stationary.
[0003] In related technologies, binocular VIO requires a certain shared field of view between the two cameras to perform binocular matching and feature point triangulation. Current binocular VIO solutions can be broadly categorized into optimization-based and filtering-based approaches. Optimization-based solutions include VINS-fusion, while filtering-based solutions include S-MSCKF and OpenVINS. Each solution can be divided into front-end and back-end processing parts. The front-end performs feature point extraction, feature point tracking, and outlier filtering; the back-end uses camera observations of feature points for triangulation and utilizes reprojection errors for state updates. However, VINS-fusion and S-MSCKF extract feature points from the left-eye image before performing binocular tracking. When an anomaly occurs in one eye, such as the right eye image, the binocular tracking process will fail because the feature points in the left eye image cannot find corresponding feature points in the right eye image. Consequently, the VIO algorithm cannot complete its localization function due to the inability to complete front-end feature point tracking. The OpenVINS solution differs from the two solutions mentioned above in its front-end feature point tracking method. OpenVINS extracts feature points from the left-eye image. If insufficient feature points are tracked in the right-eye image after binocular tracking, it further extracts feature points from the right-eye image. Therefore, OpenVINS can handle situations where anomalies occur in one-eye images. However, when anomalies occur in both binocular images, the accuracy of feature point tracking will significantly decrease. This leads to poor localization results or even completely incorrect pose estimations. Using incorrect pose estimations as localization results can cause risks to the unmanned equipment due to positioning errors. Summary of the Invention
[0004] To address the problems existing in related technologies, this disclosure provides a method, apparatus, storage medium, and electronic device for determining positioning patterns.
[0005] To achieve the above objectives, a first aspect of this disclosure provides a method for determining a positioning mode, the method being applied to a device for pose positioning based on visual inertial odometry, the method comprising:
[0006] Acquire images captured by the camera;
[0007] Determine the image type of the image;
[0008] The image is subjected to monocular image quality detection to obtain monocular image quality inspection results. The monocular image quality inspection results are used to characterize whether monocular visual inertial odometry pose calculation can be completed based on the image.
[0009] The target localization mode is determined based on the monocular image quality inspection results and the image type.
[0010] Optionally, performing monocular image quality detection on the image to obtain a monocular image quality detection result includes:
[0011] If the image is not an abnormal image and the image meets the monocular feature point extraction conditions, the image is determined to pass the monocular image quality detection, wherein the abnormal image includes at least one of overexposed image, underexposed image, and solid color image;
[0012] If the image is not the abnormal image and the image does not meet the monocular feature point extraction conditions, or if the image is the abnormal image, it is determined that the image has failed the monocular image quality detection.
[0013] Wherein, if the image passes the monocular image quality detection, it indicates that at least monocular visual inertial odometry pose calculation can be completed based on the image; if the image fails the monocular image quality detection, it indicates that monocular visual inertial odometry pose calculation cannot be completed based on the image.
[0014] Optionally, determining the image type of the image includes:
[0015] When the number of images is 1, the image type is determined to be a monocular image type;
[0016] The step of determining the target localization mode based on the monocular image quality inspection result and the image type includes:
[0017] If the image type is the monocular image type, and the monocular image quality inspection result is that the image passes the monocular image quality inspection, then the monocular positioning mode is determined as the target positioning mode.
[0018] If the monocular image quality inspection result is that the image fails the monocular image quality inspection, then the non-localization mode is determined as the target localization mode.
[0019] Optionally, determining the image type of the image includes:
[0020] When the number of images is 2, the image type of the first image and the second image in the images is determined to be a binocular image type;
[0021] The step of determining the target localization mode based on the monocular image quality inspection result and the image type includes:
[0022] When the image type is the binocular image type, the target positioning mode is determined based on the first monocular image quality inspection result of the first image and the second monocular image quality inspection result of the second image.
[0023] Optionally, determining the target localization mode based on the first monocular image quality inspection result of the first image and the second monocular image quality inspection result of the second image includes:
[0024] If the first monocular image quality inspection result indicates that the first image passes the monocular image quality inspection, and the second monocular image quality inspection result indicates that the second image passes the monocular image quality inspection, then binocular image quality inspection is performed on the first image and the second image to obtain binocular image quality inspection results. The binocular image quality inspection results are used to indicate whether binocular visual inertial odometry pose calculation can be completed based on the first image and the second image.
[0025] If the first image and the second image pass the binocular image quality detection, the binocular positioning mode is determined as the target positioning mode;
[0026] If the first image and the second image fail the binocular image quality detection, the monocular positioning mode is determined as the target positioning mode.
[0027] Optionally, determining the target localization mode based on the first monocular image quality inspection result of the first image and the second monocular image quality inspection result of the second image includes:
[0028] If the first monocular image quality inspection result indicates that the first image has passed the monocular image quality inspection, and the second monocular image quality inspection result indicates that the second image has not passed the monocular image quality inspection, then the monocular positioning mode is determined as the target positioning mode.
[0029] If the first monocular image quality inspection result indicates that the first image has failed the monocular image quality inspection, and the second monocular image quality inspection result indicates that the second image has failed the monocular image quality inspection, then the non-localization mode is determined as the target localization mode.
[0030] Optionally, the step of performing binocular image quality detection on the first image and the second image to obtain binocular image quality detection results includes:
[0031] The first image and the second image are preprocessed to obtain a first preprocessed image, a second preprocessed image, and optical flow tracing information between the first preprocessed image and the second preprocessed image.
[0032] Extract multiple first target feature points from the first preprocessed image;
[0033] For each first target feature point, based on the optical flow tracing information, a second target feature point corresponding to the first target feature point is determined from the second preprocessed image, and based on the optical flow tracing information, a third target feature point corresponding to the second target feature point is determined from the first preprocessed image.
[0034] If the distance between the first target feature point and the corresponding third target feature point is less than the first threshold, and the first target feature point and the corresponding second target feature point satisfy the epipolar constraint screening condition, then the first target feature point and the corresponding second target feature point are determined as valid pairing points;
[0035] If the number of valid pairing points is greater than a second threshold, the first image and the second image are determined to have passed the binocular image quality detection.
[0036] Optionally, the camera includes a first camera and a second camera, and the method for determining whether the first target feature point and the second target feature point satisfy the epipolar constraint screening condition includes:
[0037] Based on the intrinsic parameters of the first camera, determine the first coordinates of the first target feature point on the normalized plane corresponding to the first camera;
[0038] Based on the intrinsic parameters of the second camera, determine the second coordinates of the second target feature point on the normalized plane corresponding to the second camera;
[0039] Based on the extrinsic parameters between the first camera and the second camera, the epipolar line corresponding to the second coordinate is determined according to the epipolar constraint.
[0040] When the distance from the first coordinate to the epipolar line is less than the third threshold, it is determined that the first target feature point and the second target feature point satisfy the epipolar constraint screening condition.
[0041] Optionally, the method for determining whether the image meets the monocular feature point extraction conditions includes:
[0042] The image is subjected to histogram equalization to obtain a preprocessed image;
[0043] The preprocessed image is divided into grids to obtain multiple grid sub-images;
[0044] From each of the grid sub-images, a set of candidate feature points with response values greater than a preset response threshold is selected, and a set of target feature points is selected from the set of candidate feature points, wherein the distance between each target feature point in the set of target feature points is greater than a fourth threshold.
[0045] Calculate the mean response value for each set of target feature points, and calculate the standard deviation based on the mean response values of all the target feature points.
[0046] If the standard deviation is less than the fifth threshold and the total number of target feature points in all target feature point sets is greater than the sixth threshold, then the image is determined to meet the monocular feature point extraction conditions.
[0047] Optionally, the method further includes:
[0048] When the device is in the monocular positioning mode, at preset time intervals, for the most recently acquired image, the step of determining the image type of the image is triggered.
[0049] When the device is in the non-positioning mode, for each acquired image, the step of determining the image type of the image is triggered.
[0050] Optionally, the method further includes:
[0051] When the device is in the binocular positioning mode, if feature point matching fails during the binocular visual inertial odometry pose calculation based on the first image and the second image, the non-positioning mode is determined as the target positioning mode.
[0052] A second aspect of this disclosure provides a positioning mode determination device, the device performing pose positioning based on visual inertial odometry, the device comprising:
[0053] The receiving module is configured to acquire images captured by the camera;
[0054] The determination module is configured to determine the image type of the image;
[0055] The monocular quality inspection module is configured to perform monocular image quality inspection on the image and obtain a monocular image quality inspection result. The monocular image quality inspection result is used to characterize whether monocular visual inertial odometry pose calculation can be completed based on the image.
[0056] The determination module is configured to determine the target positioning mode based on the monocular image quality inspection results and the image type.
[0057] Optionally, the monocular quality inspection module includes:
[0058] The first execution submodule is configured to determine that the image passes the monocular image quality detection when the image is not an abnormal image and the image meets the monocular feature point extraction conditions, wherein the abnormal image includes at least one of an overexposed image, an over-dark image, and a solid color image.
[0059] The second execution submodule is configured to determine that the image has failed the monocular image quality detection when the image is not the abnormal image and the image does not meet the monocular feature point extraction conditions, or when the image is the abnormal image.
[0060] Wherein, if the image passes the monocular image quality detection, it indicates that at least monocular visual inertial odometry pose calculation can be completed based on the image; if the image fails the monocular image quality detection, it indicates that monocular visual inertial odometry pose calculation cannot be completed based on the image.
[0061] Optionally, the determination module includes:
[0062] The first determining submodule is configured to determine that the image type is a monocular image type when the number of images is 1.
[0063] The determining module includes:
[0064] The second determining submodule is configured to determine the monocular positioning mode as the target positioning mode if the monocular image quality inspection result is that the image passes the monocular image quality inspection when the image type is the monocular image type.
[0065] The third determining submodule is configured to determine the non-localization mode as the target localization mode if the monocular image quality inspection result is that the image fails the monocular image quality inspection.
[0066] Optionally, the determination module includes:
[0067] The fourth determining submodule is configured to determine, when the number of images is 2, that the image type of the first image and the second image in the images is a binocular image type;
[0068] The determining module includes:
[0069] The fifth determining submodule is configured to determine the target positioning mode based on the first monocular image quality inspection result of the first image and the second monocular image quality inspection result of the second image when the image type is the binocular image type.
[0070] Optionally, the fifth determining submodule includes:
[0071] The binocular image quality inspection submodule is configured to perform binocular image quality inspection on the first image and the second image if the first monocular image quality inspection result indicates that the first image passes the monocular image quality inspection and the second monocular image quality inspection result indicates that the second image passes the monocular image quality inspection, thereby obtaining a binocular image quality inspection result. The binocular image quality inspection result is used to indicate whether binocular visual inertial odometry pose calculation can be completed based on the first image and the second image.
[0072] The sixth determining submodule is configured to determine the binocular positioning mode as the target positioning mode when the first image and the second image pass the binocular image quality detection.
[0073] The seventh determining submodule is configured to determine the monocular positioning mode as the target positioning mode if the first image and the second image fail the binocular image quality detection.
[0074] Optionally, the fifth determining submodule includes:
[0075] The eighth determining submodule is configured to determine the monocular positioning mode as the target positioning mode if the first monocular image quality inspection result indicates that the first image has passed the monocular image quality inspection and the second monocular image quality inspection result indicates that the second image has not passed the monocular image quality inspection.
[0076] The ninth determining submodule is configured to determine the non-localization mode as the target localization mode if the first monocular image quality inspection result indicates that the first image has failed the monocular image quality inspection, and the second monocular image quality inspection result indicates that the second image has failed the monocular image quality inspection.
[0077] Optionally, the binocular image quality inspection submodule is configured to:
[0078] The first preprocessing submodule is configured to preprocess the first image and the second image to obtain a first preprocessed image, a second preprocessed image, and optical flow tracing information between the first preprocessed image and the second preprocessed image.
[0079] An extraction submodule is configured to extract a plurality of first target feature points from the first preprocessed image;
[0080] The optical flow tracing submodule is configured to, for each first target feature point, determine a second target feature point corresponding to the first target feature point from the second preprocessed image based on the optical flow tracing information, and determine a third target feature point corresponding to the second target feature point from the first preprocessed image based on the optical flow tracing information.
[0081] The third execution submodule is configured to determine the first target feature point and the corresponding second target feature point as valid pairing points if the distance between the first target feature point and the corresponding third target feature point is less than a first threshold and the first target feature point and the corresponding second target feature point satisfy the epipolar constraint screening condition.
[0082] The fourth execution submodule is configured to determine, if the number of valid pairing points is greater than a second threshold, that the first image and the second image have passed the binocular image quality detection.
[0083] Optionally, the camera includes a first camera and a second camera, and a third execution submodule is further configured to:
[0084] Based on the intrinsic parameters of the first camera, determine the first coordinates of the first target feature point on the normalized plane corresponding to the first camera;
[0085] Based on the intrinsic parameters of the second camera, determine the second coordinates of the second target feature point on the normalized plane corresponding to the second camera;
[0086] Based on the extrinsic parameters between the first camera and the second camera, the epipolar line corresponding to the second coordinate is determined according to the epipolar constraint.
[0087] When the distance from the first coordinate to the epipolar line is less than the third threshold, it is determined that the first target feature point and the second target feature point satisfy the epipolar constraint screening condition.
[0088] Optionally, the second execution submodule is further configured to determine whether the image satisfies the monocular feature point extraction conditions by:
[0089] The image is subjected to histogram equalization to obtain a preprocessed image;
[0090] The preprocessed image is divided into grids to obtain multiple grid sub-images;
[0091] From each of the grid sub-images, a set of candidate feature points with response values greater than a preset response threshold is selected, and a set of target feature points is selected from the set of candidate feature points, wherein the distance between each target feature point in the set of target feature points is greater than a fourth threshold.
[0092] Calculate the mean response value for each set of target feature points, and calculate the standard deviation based on the mean response values of all the target feature points.
[0093] If the standard deviation is less than the fifth threshold and the total number of target feature points in all target feature point sets is greater than the sixth threshold, then the image is determined to meet the monocular feature point extraction conditions.
[0094] Optionally, the device further includes:
[0095] The first adaptive module is configured to, when the device is in the monocular positioning mode, trigger the step of determining the image type of the most recently acquired image at preset time intervals.
[0096] The second adaptive module is configured to, when the device is in the non-positioning mode, trigger the step of determining the image type of the image for each acquired image.
[0097] Optionally, the device further includes:
[0098] The third adaptive module is configured to, when the device is in the binocular positioning mode, if feature point matching fails during the binocular visual inertial odometry pose calculation based on the first image and the second image, determine the non-positioning mode as the target positioning mode.
[0099] A third aspect of this disclosure provides a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the steps of the method described in any of the first aspects.
[0100] A fourth aspect of this disclosure provides an electronic device, the electronic device comprising:
[0101] A memory on which computer programs are stored;
[0102] A processor for executing the computer program in the memory to implement the steps of the method of any one of the first aspects.
[0103] By adopting the above technical solution, at least the following beneficial technical effects can be achieved:
[0104] The system acquires images captured by a camera and determines the image type of the received images. Monocular image quality inspection is performed on the images to obtain the monocular image quality inspection result. Since the monocular image quality inspection result indicates whether at least monocular visual inertial odometry pose calculation can be completed based on the received images, a specific target positioning mode can be determined based on the monocular image quality inspection result and the image type. This allows the device to enter the target positioning mode and perform pose positioning calculations according to the target positioning mode. This method, by performing quality inspection on the received images before performing visual inertial odometry pose positioning calculations and determining the target positioning mode based on the monocular image quality inspection result and the image type, effectively avoids the problem of VINS-Fusion and S-MSCKF schemes failing to locate due to monocular image anomalies in related technologies. It also avoids the problem of low positioning accuracy caused by binocular image anomalies in the OpenVINS scheme. Therefore, this method can adaptively determine the target positioning mode based on image quality under various image anomaly conditions, adapting to the current image anomaly situation and obtaining highly accurate pose positioning results through the target positioning mode, thus ensuring the safety of the device.
[0105] Other features and advantages of this disclosure will be described in detail in the following detailed description section. Attached Figure Description
[0106] The accompanying drawings are provided to further illustrate the present disclosure and form part of the specification. They are used together with the following detailed description to explain the present disclosure, but do not constitute a limitation thereof. In the drawings:
[0107] Figure 1 This is a flowchart illustrating a method for determining a positioning pattern according to an exemplary embodiment of the present disclosure.
[0108] Figure 2 This is a flowchart illustrating another method for determining a positioning mode according to an exemplary embodiment of the present disclosure.
[0109] Figure 3 This is a block diagram illustrating a positioning mode determination device according to an exemplary embodiment of the present disclosure.
[0110] Figure 4 This is a block diagram illustrating an electronic device according to an exemplary embodiment of the present disclosure. Detailed Implementation
[0111] The specific embodiments of this disclosure will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustration and explanation only and are not intended to limit this disclosure.
[0112] It should be noted that all actions involving the acquisition of signals, information, or data in this disclosure are carried out in compliance with the relevant data protection laws and policies of the country where the location is situated, and with authorization from the owner of the relevant device.
[0113] This disclosure addresses the problem in related technologies where binocular visual inertial odometry (VIO) suffers from poor pose localization or even fails to locate when encountering anomalies in either monocular or binocular images. It proposes a method that performs image quality detection before calculating the VIO's pose localization, and adaptively determines / switches the VIO's localization mode based on the detected image quality. This effectively avoids the problem of being unable to locate when monocular images are abnormal, or the problem of poor localization performance when binocular images are abnormal.
[0114] The following provides a detailed description of the embodiments of the technical solution disclosed herein.
[0115] Figure 1 This is a flowchart illustrating a method for determining a positioning mode according to an exemplary embodiment of the present disclosure. This method for determining a positioning mode is applied to a device for pose positioning based on visual inertial odometry, such as an electronic device like a drone, unmanned vehicle, or robot. Figure 1 As shown, the method for determining this positioning mode may include the following steps:
[0116] S11. Acquire images captured by the camera.
[0117] The camera can be either a monocular camera or a binocular camera. In the case of a binocular camera, it includes a first camera for the left eye and a second camera for the right eye.
[0118] S12. Determine the image type of the image.
[0119] The image type may include at least one of monocular image type and binocular image type.
[0120] Since binocular visual inertial odometry uses a pair of simultaneously exposed camera images and IMU data as input data, one implementation for determining the image type of an image can be to determine the image type of the received image based on the number of images received in step S11.
[0121] Another method for determining the image type is to identify the image type based on an image type marker on the image. The image type marker can be provided by the camera system.
[0122] S13. Perform monocular image quality detection on the image to obtain monocular image quality inspection results. The monocular image quality inspection results are used to characterize whether monocular visual inertial odometry pose calculation can be completed based on the image.
[0123] For example, after determining the image type of the received images, monocular image quality inspection can be performed on each image to obtain a monocular image quality inspection result. The monocular image quality inspection result indicates whether at least monocular visual inertial odometry pose calculation can be completed based on the received images. For example, if one received image passes the monocular image quality inspection, then monocular visual inertial odometry pose calculation can be completed based on that image. As another example, if both received images pass the monocular image quality inspection, then in addition to being able to complete monocular visual inertial odometry pose calculation based on either of the two images, it is also possible to complete binocular visual inertial odometry pose calculation based on both images.
[0124] S14. Determine the target positioning mode based on the monocular image quality inspection results and the image type.
[0125] After determining the target positioning mode, the controllable device can enter or switch to the target positioning mode. When the device enters / switches to the target positioning mode, it performs pose positioning calculations based on the target positioning mode.
[0126] Using the above method, images captured by the camera are acquired, and the image type of the received images is determined. Monocular image quality inspection is performed on the images to obtain the monocular image quality inspection result. Since the monocular image quality inspection result is used to characterize whether at least monocular visual inertial odometry pose calculation can be completed based on the received images, a specific target positioning mode can be determined based on the monocular image quality inspection result and the image type. This allows the device to enter the target positioning mode and perform pose positioning calculation according to the target positioning mode. This method, because it performs quality inspection on the received images before performing visual inertial odometry pose positioning calculation and determines the target positioning mode based on the monocular image quality inspection result and the image type, effectively avoids the problem of VINS-Fusion and S-MSCKF schemes failing to locate due to monocular image anomalies in related technologies, and also avoids the problem of low positioning accuracy caused by binocular image anomalies in the OpenVINS scheme. Therefore, by adopting the method disclosed herein, the target positioning mode can be adaptively determined based on the image quality under various abnormal image conditions, thereby adapting to the current abnormal image conditions and obtaining highly accurate pose positioning results through the target positioning mode, thus ensuring the safety of the device.
[0127] Optionally, performing monocular image quality detection on the image to obtain a monocular image quality detection result includes:
[0128] If the image is not an anomalous image and meets the monocular feature point extraction conditions, the image is determined to pass the monocular image quality detection, wherein the anomalous image includes at least one of an overexposed image, an underexposed image, and a solid color image; if the image is not an anomalous image and does not meet the monocular feature point extraction conditions, or if the image is an anomalous image, the image is determined to fail the monocular image quality detection; wherein, if the image passes the monocular image quality detection, it indicates that at least monocular visual inertial odometry pose calculation can be completed based on the image; if the image fails the monocular image quality detection, it indicates that monocular visual inertial odometry pose calculation cannot be completed based on the image.
[0129] Since overexposed images, underexposed images, and solid color images can all reduce the accuracy of feature point tracking or cause tracking failure, resulting in poor pose localization or even completely wrong pose estimation results, overexposed images, underexposed images, and solid color images are identified as abnormal images in this embodiment of the disclosure.
[0130] One method for determining whether an image is abnormal is to calculate its grayscale mean and standard deviation. Since the grayscale mean represents the average brightness of the image, and the standard deviation represents the degree of brightness dispersion, overly dark images have smaller grayscale mean values, overly bright images have larger grayscale mean values, and solid color images (or images close to solid colors) have smaller grayscale standard deviations. Therefore, in this disclosure, if a threshold g_min < the image's grayscale mean < a threshold g_max, and the image's grayscale standard deviation > a threshold h_min, the image can be determined to be an abnormal image, not an overexposed, overly dark, or solid color image. Conversely, if an image is abnormal, it can be determined to be an abnormal image. It should be noted that an overexposed image is defined as an image with a grayscale mean greater than the threshold g_max. An overly dark image is defined as an image with a grayscale mean less than the threshold g_min. A solid color image is defined as an image with a grayscale standard deviation greater than the threshold h_min.
[0131] Another method for determining whether an image is abnormal is to determine the image's histogram and, based on the image characteristics of the histogram, determine whether the corresponding image is an abnormal image, such as an overexposed image, an over-dark image, or a solid color image.
[0132] If the image is determined not to be an abnormal image, it is necessary to further determine whether the image meets the monocular feature point extraction conditions. One implementation method for determining whether the image meets the monocular feature point extraction conditions includes:
[0133] The image is subjected to histogram equalization to obtain a preprocessed image; the preprocessed image is divided into grids to obtain multiple grid sub-images; a set of candidate feature points with response values greater than a preset response threshold is selected from each grid sub-image, and a set of target feature points is selected from the set of candidate feature points based on the distance between the points, wherein the distance between each target feature point in the set of target feature points is greater than a fourth threshold; the mean response value of each set of target feature points is calculated, and the standard deviation is calculated based on the mean response values of all the target feature points; if the standard deviation is less than a fifth threshold and the total number of target feature points in all the sets of target feature points is greater than a sixth threshold, then the image is determined to meet the monocular feature point extraction condition.
[0134] It should be explained that histogram equalization is a method in the field of image processing that uses the image histogram to adjust the contrast, which will not be described in detail in this disclosure.
[0135] In this disclosure, histogram equalization is performed on the image to increase its global contrast, resulting in a preprocessed image. The preprocessed image is then divided into multiple grid sub-images. For each grid sub-image, a set of candidate feature points with response values greater than a preset response threshold is selected based on the FAST algorithm. From this set of candidate feature points, a target feature point set is then selected, where the distance between all target feature points in the target feature point set is greater than a fourth threshold. This method not only ensures that the selected target feature points are evenly distributed across the image but also avoids the phenomenon of clustering of selected target feature points.
[0136] After obtaining the target feature point set corresponding to each grid sub-image, the mean response value of each target feature point set can be calculated based on the response value of each target feature point, and the standard deviation can be calculated based on the mean of all response values. If the standard deviation is less than the fifth threshold and the total number of target feature points in all target feature point sets is greater than the sixth threshold, it indicates that a sufficient number of uniformly distributed target feature points can be extracted from the image, and the average response values of the target feature points in each grid sub-image are relatively similar. This type of image is suitable for feature point extraction and tracking processing for monocular VIO, that is, it is suitable for completing monocular visual inertial odometry pose calculation. Therefore, it can be determined that this image meets the monocular feature point extraction conditions.
[0137] Conversely, if the standard deviation is greater than or equal to the fifth threshold, or the total number of target feature points in all target feature point sets is less than or equal to the sixth threshold, then the image can be determined not to meet the monocular feature point extraction conditions.
[0138] It should be noted that the FAST algorithm defines a pixel as a feature point if it is located in a different region from a sufficiently large number of its surrounding pixels. The response value is a parameter used to describe the discriminative power of a feature point. A larger response value indicates greater discriminative power for that feature point.
[0139] Any of the thresholds in this disclosure can be adaptively set based on needs, and this disclosure does not impose any specific limitations on them.
[0140] For example, if the image is not an abnormal image and the image meets the monocular feature point extraction conditions, it can be determined that the image passes the monocular image quality detection. If the image passes the monocular image quality detection, it indicates that at least monocular visual inertial odometry pose calculation can be completed based on the image.
[0141] For another example, if the image is not an anomalous image but does not meet the monocular feature point extraction conditions, or if the image is an anomalous image, it can be determined that the image has failed the monocular image quality detection. If the image fails the monocular image quality detection, it indicates that monocular visual inertial odometry pose calculation cannot be performed based on the image.
[0142] Optionally, determining the image type of the image includes:
[0143] When the number of images is 1, the image type is determined to be a monocular image type;
[0144] Accordingly, determining the target localization mode based on the monocular image quality inspection result and the image type includes:
[0145] When the image type is the monocular image type, if the monocular image quality inspection result is that the image passes the monocular image quality inspection, then the monocular positioning mode is determined as the target positioning mode; if the monocular image quality inspection result is that the image fails the monocular image quality inspection, then the non-positioning mode is determined as the target positioning mode.
[0146] For example, if the number of received images is 1, then the image type of that image can be determined to be a monocular image. Furthermore, if the monocular image quality inspection result for that image is that it passes the monocular image quality check, then the monocular positioning mode can be determined as the target positioning mode, and in the monocular positioning mode, the monocular visual inertial odometry pose calculation can be performed based on that image. Conversely, if the monocular image quality inspection result is that the image fails the monocular image quality check, then the non-positioning mode can be determined as the target positioning mode.
[0147] Monocular positioning mode refers to the mode in which the device calculates pose based on monocular visual inertial odometry. Non-positioning mode refers to the mode in which the device does not perform pose calculation.
[0148] Optionally, determining the image type of the image includes:
[0149] When the number of images is 2, the image type of the first image and the second image in the images is determined to be a binocular image type;
[0150] Accordingly, determining the target localization mode based on the monocular image quality inspection result and the image type includes:
[0151] When the image type is the binocular image type, the target positioning mode is determined based on the first monocular image quality inspection result of the first image and the second monocular image quality inspection result of the second image.
[0152] For example, when the number of received images is two, the image types of the first and second images can be determined to be binocular images. The target localization mode can be determined based on the quality inspection results of the first monocular image of the first image and the second monocular image of the second image. For instance, if the quality inspection result of the first monocular image of the first image indicates that the first image failed the monocular image quality inspection, and the quality inspection result of the second monocular image of the second image indicates that the second image failed the monocular image quality inspection, the non-localization mode can be determined as the target localization mode. As another example, if the quality inspection result of the first monocular image indicates that the first image passed the monocular image quality inspection, and the quality inspection result of the second monocular image indicates that the second image failed the monocular image quality inspection, the monocular localization mode can be determined as the target localization mode, and the monocular visual inertial odometry pose calculation can be performed based on the first image under the monocular localization mode.
[0153] For example, if the first monocular image quality inspection result indicates that the first image has failed the monocular image quality inspection, and the second monocular image quality inspection result indicates that the second image has passed the monocular image quality inspection, then the monocular positioning mode can be determined as the target positioning mode. In this monocular positioning mode, the monocular visual inertial odometry pose calculation can be performed based on the second image.
[0154] For example, if the first monocular image quality inspection result indicates that the first image has passed the monocular image quality inspection, and the second monocular image quality inspection result indicates that the second image has passed the monocular image quality inspection, it can be further determined whether the pose calculation of binocular visual inertial odometry can be completed based on the first image and the second image.
[0155] It should be noted that in the embodiments of this disclosure, the first image and the second image represent the left eye image and the right eye image, but this disclosure does not limit the first image / second image to be a left eye image or a right eye image.
[0156] Optionally, determining the target localization mode based on the first monocular image quality inspection result of the first image and the second monocular image quality inspection result of the second image includes:
[0157] If the first monocular image quality inspection result indicates that the first image passes the monocular image quality inspection, and the second monocular image quality inspection result indicates that the second image passes the monocular image quality inspection, then a binocular image quality inspection is performed on the first image and the second image to obtain a binocular image quality inspection result. The binocular image quality inspection result is used to indicate whether binocular visual inertial odometry pose calculation can be completed based on the first image and the second image. If the first image and the second image pass the binocular image quality inspection, the binocular positioning mode is determined as the target positioning mode. If the first image and the second image fail the binocular image quality inspection, the monocular positioning mode is determined as the target positioning mode.
[0158] For example, if the first monocular image quality inspection result indicates that the first image passes the monocular image quality inspection, and the second monocular image quality inspection result indicates that the second image passes the monocular image quality inspection, then a binocular image quality inspection is performed on the first and second images to obtain a binocular image quality inspection result that indicates whether binocular visual inertial odometry pose calculation can be completed based on the first and second images. If the first and second images pass the binocular image quality inspection, then the binocular positioning mode can be determined as the target positioning mode. In the binocular positioning mode, the binocular visual inertial odometry pose calculation is completed based on the first and second images. If the first and second images fail the binocular image quality inspection, then the monocular positioning mode can be determined as the target positioning mode. In the monocular positioning mode, the monocular visual inertial odometry pose calculation is completed based on either the first or second image.
[0159] Optionally, the step of performing binocular image quality detection on the first image and the second image to obtain binocular image quality detection results includes:
[0160] The first image and the second image are preprocessed to obtain a first preprocessed image, a second preprocessed image, and optical flow tracing information between the first preprocessed image and the second preprocessed image. Multiple first target feature points are extracted from the first preprocessed image. For each first target feature point, based on the optical flow tracing information, a second target feature point corresponding to the first target feature point is determined from the second preprocessed image, and a third target feature point corresponding to the second target feature point is determined from the first preprocessed image based on the optical flow tracing information. If the distance between the first target feature point and the corresponding third target feature point is less than a first threshold, and the first target feature point and the corresponding second target feature point satisfy the epipolar constraint screening condition, then the first target feature point and the corresponding second target feature point are determined as valid pairing points. If the number of valid pairing points is greater than a second threshold, the first image and the second image are determined to have passed the binocular image quality detection.
[0161] The preprocessing of the first and second images includes histogram equalization to make their average gray levels similar, thereby satisfying the gray-level invariance assumption required for optical flow tracking. Based on this assumption, an improved LK optical flow tracking algorithm (KLT optical flow algorithm) based on image pyramids is used to perform optical flow tracking on the first and second preprocessed images, obtaining optical flow tracking information between them. This information includes the optical flow and affine transformation matrix between the two images.
[0162] In some embodiments, the first preprocessed image and the second preprocessed image can be the histogram equalization results obtained during monocular image quality detection of the first image and the second image. In another embodiment, in order to make the average gray levels of the processed first preprocessed image and the second preprocessed image as similar as possible, a target parameter (such as using the same average gray level value) can be set to perform histogram equalization on the first image and the second image.
[0163] After preprocessing the first image and the second image to obtain the first preprocessed image, the second preprocessed image, and the optical flow tracking information between the first preprocessed image and the second preprocessed image, multiple first target feature points can be extracted from the first preprocessed image. The method of extracting the first target feature points is similar to the method of determining the set of all target feature points of the image mentioned above. Alternatively, the set of all target feature points of the first image obtained during the monocular image quality detection of the first image can be directly used as multiple first target feature points of the first preprocessed image.
[0164] Furthermore, for each first target feature point in the first preprocessed image, based on optical flow tracing information, a second target feature point corresponding to the first target feature point is tracked and determined from the second preprocessed image, and based on optical flow tracing information, a third target feature point corresponding to the second target feature point is tracked and determined from the first preprocessed image. If the distance between the first target feature point and the corresponding third target feature point is less than a first threshold, and the first target feature point and the corresponding second target feature point satisfy the epipolar constraint screening condition, then the first target feature point and the corresponding second target feature point are determined as valid pairing points.
[0165] If the number of valid paired points in the first and second preprocessed images is greater than a second threshold, then the first and second images can be determined to have passed the binocular image quality detection. Conversely, if the number of valid paired points in the first and second preprocessed images is less than or equal to the second threshold, then the first and second images can be determined to have failed the binocular image quality detection.
[0166] Optionally, the camera includes a first camera and a second camera, and the method for determining whether the first target feature point and the second target feature point satisfy the epipolar constraint screening condition includes:
[0167] Based on the intrinsic parameters of the first camera, determine the first coordinates of the first target feature point on the normalized plane corresponding to the first camera; based on the intrinsic parameters of the second camera, determine the second coordinates of the second target feature point on the normalized plane corresponding to the second camera; based on the extrinsic parameters between the first camera and the second camera, determine the epipolar line corresponding to the second coordinates according to the epipolar constraint; when the distance from the first coordinate to the epipolar line is less than a third threshold, determine that the first target feature point and the second target feature point satisfy the epipolar constraint screening condition.
[0168] In this system, the intrinsic parameters k1 of the first camera, k2 of the second camera, and the extrinsic parameters R and t between the first and second cameras are all known parameters. For specific parameter definitions, please refer to the definitions in related technical documents.
[0169] Epipolar constraint is a point-to-line constraint that provides important constraints on corresponding points, compressing the process of finding corresponding points from the entire image to finding corresponding points on a straight line.
[0170] Based on the intrinsic parameter k1 of the first camera, the first coordinates of the first target feature point p1 on the normalized plane corresponding to the first camera can be determined. Based on the intrinsic parameter k2 of the second camera, the second coordinates of the second target feature point p2 on the normalized plane corresponding to the second camera can be determined. Based on the extrinsic parameters R and T between the first and second cameras, and according to the epipolar constraint, the epipolar line l(a, b, c) = t∧×R×x2 corresponding to the second coordinate x2 is determined, where t^ represents the antisymmetric matrix of t; the distance from the first coordinate x1 to the epipolar line l(a, b, c) = t^×R×x2 is also considered. When the value is less than the third threshold, the first target feature point p1 and the second target feature point p2 are determined to satisfy the epipolar constraint screening condition. Here, a, b, and c are the 3D data of the vector in the formula l(a, b, c) = t^×R×x2, and T is the transpose sign. Conversely, the distance from the first coordinate x1 to the epipolar line l(a, b, c) = t^×R×x2... When the value is greater than or equal to the third threshold, it is determined that the first target feature point p1 and the second target feature point p2 do not meet the epipolar constraint screening conditions.
[0171] Optionally, the method for determining the positioning mode further includes:
[0172] When the device is in the monocular positioning mode, at preset time intervals, for the most recently acquired image, the step of determining the image type of the image is triggered; when the device is in the non-positioning mode, for each acquired image, the step of determining the image type of the image is triggered.
[0173] For example, when the device is in monocular positioning mode, at preset intervals such as 30 seconds, step S12 can be triggered to determine the image type of the most recently acquired image. This involves performing monocular image quality detection on the image, obtaining the monocular image quality inspection result, and determining the target positioning mode based on the result and image type. The device is then controlled to enter or switch to the target positioning mode. For instance, after step S11, it is determined whether the timer has reached 30 seconds. If so, step S12 is executed. If not, pose calculation continues according to the monocular positioning mode.
[0174] Similarly, when the device is in non-positioning mode, for each image acquired in step S11, step S12 is triggered to determine the image type, and then monocular image quality detection is performed on the image to obtain the monocular image quality inspection result. Based on the monocular image quality inspection result and the image type, the target positioning mode is determined, and the device is controlled to enter or switch to the target positioning mode. For example, S12 is executed after each step S11.
[0175] This method, which triggers the S12 step of determining the image type only when an image is acquired in non-positioning mode and every 30 seconds in monocular positioning mode, and then performs monocular image quality inspection on the image to obtain the monocular image quality inspection result, and determines the target positioning mode based on the monocular image quality inspection result and the image type, and controls the device to enter or switch to the target positioning mode, can reduce the additional computational consumption caused by image quality inspection.
[0176] Optionally, the method for determining the positioning mode further includes:
[0177] When the device is in the binocular positioning mode, if feature point matching fails during the binocular visual inertial odometry pose calculation based on the first image and the second image, the non-positioning mode is determined as the target positioning mode.
[0178] When the device is in binocular positioning mode, if feature point matching fails during the binocular visual inertial odometry pose calculation based on the first and second images, the control device exits the binocular positioning mode and enters the non-positioning mode. This allows the device to adaptively switch between non-positioning mode, monocular positioning mode, and binocular positioning mode.
[0179] Figure 2 This is a flowchart illustrating another method for determining a positioning mode according to an exemplary embodiment of this disclosure. Figure 2 As shown, the method includes the following steps:
[0180] S21. Acquire images captured by the camera;
[0181] S22. Determine the image type of the image;
[0182] If it is a monocular image type, then execute S23. If it is a binocular image type, then the characterization image includes a first image and a second image, and execute S24.
[0183] S23. Perform monocular image quality inspection on the image to obtain the monocular image quality inspection result;
[0184] S24. Perform monocular image quality detection on the first image and the second image in the image respectively to obtain the quality inspection result of the first monocular image and the quality inspection result of the second monocular image.
[0185] S25. Determine whether the monocular image quality inspection result indicates that the image has passed the monocular image quality inspection.
[0186] If the monocular image quality inspection result indicates that the image passes the monocular image quality inspection, then execute S26; if the monocular image quality inspection result indicates that the image fails the monocular image quality inspection, then execute S27.
[0187] S26. Set the monocular positioning mode as the target positioning mode;
[0188] S27. Set the non-positioning mode as the target positioning mode;
[0189] S28. Determine the quality inspection result of the first monocular image and the quality inspection result of the second monocular image to indicate that one of the first image and the second image has passed the monocular image quality inspection.
[0190] In this case, S26 is executed.
[0191] S29. Determine the quality inspection result of the first monocular image and the quality inspection result of the second monocular image to indicate that neither the first image nor the second image passed the monocular image quality inspection.
[0192] In this case, S27 is executed.
[0193] S30, the first monocular image quality inspection result, and the second monocular image quality inspection result indicate that both the first image and the second image have passed the monocular image quality inspection;
[0194] In this case, S31 is executed.
[0195] S31. Determine whether the first image and the second image pass the binocular image quality detection;
[0196] If the first image and the second image pass the binocular image quality detection, then execute S32; if the first image and the second image fail the binocular image quality detection, then execute S26.
[0197] S32. Set the binocular positioning mode as the target positioning mode.
[0198] The specific implementation methods of each step in the above embodiments have been described in detail in the embodiments of the aforementioned related methods, and will not be repeated here.
[0199] Figure 3 This is a block diagram illustrating a positioning mode determination device according to an exemplary embodiment of the present disclosure. Figure 3 As shown, the positioning mode determining device 300 includes:
[0200] The receiving module 310 is configured to acquire images captured by the camera;
[0201] The determination module 320 is configured to determine the image type of the image;
[0202] The monocular quality inspection module 330 is configured to perform monocular image quality inspection on the image and obtain a monocular image quality inspection result. The monocular image quality inspection result is used to characterize whether monocular visual inertial odometry pose calculation can be completed based on the image.
[0203] The determination module 340 is configured to determine the target positioning mode based on the monocular image quality inspection results and the image type.
[0204] By employing the aforementioned device, the received image undergoes quality inspection before visual inertial odometry pose localization calculation. Based on the monocular image quality inspection results and the image type, the target localization mode is determined. Therefore, this method effectively avoids the problem of VINS-Fusion and S-MSCKF schemes failing to locate due to monocular image anomalies, and also avoids the low accuracy of localization results caused by binocular image anomalies in the OpenVINS scheme. Thus, this method can adaptively determine the target localization mode based on image quality under various image anomaly conditions, adapting to the current image anomaly situation and obtaining highly accurate pose localization results through the target localization mode, thereby ensuring the safety of the device.
[0205] Optionally, the monocular quality inspection module 330 includes:
[0206] The first execution submodule is configured to determine that the image passes the monocular image quality detection when the image is not an abnormal image and the image meets the monocular feature point extraction conditions, wherein the abnormal image includes at least one of an overexposed image, an over-dark image, and a solid color image.
[0207] The second execution submodule is configured to determine that the image has failed the monocular image quality detection when the image is not the abnormal image and the image does not meet the monocular feature point extraction conditions, or when the image is the abnormal image.
[0208] Wherein, if the image passes the monocular image quality detection, it indicates that at least monocular visual inertial odometry pose calculation can be completed based on the image; if the image fails the monocular image quality detection, it indicates that monocular visual inertial odometry pose calculation cannot be completed based on the image.
[0209] Optionally, the determination module 320 includes:
[0210] The first determining submodule is configured to determine that the image type is a monocular image type when the number of images is 1.
[0211] The determining module 340 includes:
[0212] The second determining submodule is configured to determine the monocular positioning mode as the target positioning mode if the monocular image quality inspection result is that the image passes the monocular image quality inspection when the image type is the monocular image type.
[0213] The third determining submodule is configured to determine the non-localization mode as the target localization mode if the monocular image quality inspection result is that the image fails the monocular image quality inspection.
[0214] Optionally, the determination module 320 includes:
[0215] The fourth determining submodule is configured to determine, when the number of images is 2, that the image type of the first image and the second image in the images is a binocular image type;
[0216] The determining module includes:
[0217] The fifth determining submodule is configured to determine the target positioning mode based on the first monocular image quality inspection result of the first image and the second monocular image quality inspection result of the second image when the image type is the binocular image type.
[0218] Optionally, the fifth determining submodule includes:
[0219] The binocular image quality inspection submodule is configured to perform binocular image quality inspection on the first image and the second image if the first monocular image quality inspection result indicates that the first image passes the monocular image quality inspection and the second monocular image quality inspection result indicates that the second image passes the monocular image quality inspection, thereby obtaining a binocular image quality inspection result. The binocular image quality inspection result is used to indicate whether binocular visual inertial odometry pose calculation can be completed based on the first image and the second image.
[0220] The sixth determining submodule is configured to determine the binocular positioning mode as the target positioning mode when the first image and the second image pass the binocular image quality detection.
[0221] The seventh determining submodule is configured to determine the monocular positioning mode as the target positioning mode if the first image and the second image fail the binocular image quality detection.
[0222] Optionally, the fifth determining submodule includes:
[0223] The eighth determining submodule is configured to determine the monocular positioning mode as the target positioning mode if the first monocular image quality inspection result indicates that the first image has passed the monocular image quality inspection and the second monocular image quality inspection result indicates that the second image has not passed the monocular image quality inspection.
[0224] The ninth determining submodule is configured to determine the non-localization mode as the target localization mode if the first monocular image quality inspection result indicates that the first image has failed the monocular image quality inspection, and the second monocular image quality inspection result indicates that the second image has failed the monocular image quality inspection.
[0225] Optionally, the binocular image quality inspection submodule is configured to:
[0226] The first preprocessing submodule is configured to preprocess the first image and the second image to obtain a first preprocessed image, a second preprocessed image, and optical flow tracing information between the first preprocessed image and the second preprocessed image.
[0227] An extraction submodule is configured to extract a plurality of first target feature points from the first preprocessed image;
[0228] The optical flow tracing submodule is configured to, for each first target feature point, determine a second target feature point corresponding to the first target feature point from the second preprocessed image based on the optical flow tracing information, and determine a third target feature point corresponding to the second target feature point from the first preprocessed image based on the optical flow tracing information.
[0229] The third execution submodule is configured to determine the first target feature point and the corresponding second target feature point as valid pairing points if the distance between the first target feature point and the corresponding third target feature point is less than a first threshold and the first target feature point and the corresponding second target feature point satisfy the epipolar constraint screening condition.
[0230] The fourth execution submodule is configured to determine, if the number of valid pairing points is greater than a second threshold, that the first image and the second image have passed the binocular image quality detection.
[0231] Optionally, the camera includes a first camera and a second camera, and a third execution submodule is further configured to:
[0232] Based on the intrinsic parameters of the first camera, determine the first coordinates of the first target feature point on the normalized plane corresponding to the first camera;
[0233] Based on the intrinsic parameters of the second camera, determine the second coordinates of the second target feature point on the normalized plane corresponding to the second camera;
[0234] Based on the extrinsic parameters between the first camera and the second camera, the epipolar line corresponding to the second coordinate is determined according to the epipolar constraint.
[0235] When the distance from the first coordinate to the epipolar line is less than the third threshold, it is determined that the first target feature point and the second target feature point satisfy the epipolar constraint screening condition.
[0236] Optionally, the second execution submodule is further configured to determine whether the image satisfies the monocular feature point extraction conditions by:
[0237] The image is subjected to histogram equalization to obtain a preprocessed image;
[0238] The preprocessed image is divided into grids to obtain multiple grid sub-images;
[0239] From each of the grid sub-images, a set of candidate feature points with response values greater than a preset response threshold is selected, and a set of target feature points is selected from the set of candidate feature points, wherein the distance between each target feature point in the set of target feature points is greater than a fourth threshold.
[0240] Calculate the mean response value for each set of target feature points, and calculate the standard deviation based on the mean response values of all the target feature points.
[0241] If the standard deviation is less than the fifth threshold and the total number of target feature points in all target feature point sets is greater than the sixth threshold, then the image is determined to meet the monocular feature point extraction conditions.
[0242] Optionally, the device 300 further includes:
[0243] The first adaptive module is configured to, when the device is in the monocular positioning mode, trigger the step of determining the image type of the most recently acquired image at preset time intervals.
[0244] The second adaptive module is configured to, when the device is in the non-positioning mode, trigger the step of determining the image type of the image for each acquired image.
[0245] Optionally, the device 300 further includes:
[0246] The third adaptive module is configured to, when the device is in the binocular positioning mode, if feature point matching fails during the binocular visual inertial odometry pose calculation based on the first image and the second image, determine the non-positioning mode as the target positioning mode.
[0247] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0248] This disclosure also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the positioning mode determination method in the above embodiments.
[0249] Figure 4 This is a block diagram illustrating an electronic device 700 according to an exemplary embodiment of the present disclosure. Figure 4 As shown, the electronic device 700 may include a processor 701 and a memory 702. The electronic device 700 may also include one or more of a multimedia component 703, an input / output (I / O) interface 704, and a communication component 705. Furthermore, the electronic device 700 includes a camera and an IMU unit.
[0250] The processor 701 controls the overall operation of the electronic device 700 to complete all or part of the steps in the method for determining the positioning mode described above. The memory 702 stores various types of data to support the operation of the electronic device 700. This data may include, for example, instructions for any application or method operating on the electronic device 700, and application-related data such as contact data, sent and received messages, pictures, audio, video, etc. The memory 702 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. Multimedia component 703 may include a screen and an audio component. The screen may be, for example, a touchscreen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signals may be further stored in memory 702 or transmitted via communication component 705. The audio component also includes at least one speaker for outputting audio signals. I / O interface 704 provides an interface between processor 701 and other interface modules, such as a keyboard, mouse, buttons, etc. These buttons may be virtual or physical buttons. Communication component 705 is used for wired or wireless communication between the electronic device 700 and other devices. Wireless communication, such as Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, 4G, NB-IoT, eMTC, or other 5G technologies, or combinations thereof, is not limited here. Therefore, the corresponding communication component 705 may include: a Wi-Fi module, a Bluetooth module, an NFC module, etc.
[0251] In an exemplary embodiment, the electronic device 700 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the positioning mode determination method described above.
[0252] In another exemplary embodiment, a computer-readable storage medium including program instructions is also provided, which, when executed by a processor, implement the steps of the positioning mode determination method described above. For example, the computer-readable storage medium may be the memory 702 including program instructions described above, which may be executed by the processor 701 of the electronic device 700 to complete the positioning mode determination method described above.
[0253] In another exemplary embodiment, a computer program product is also provided, the computer program product comprising a computer program executable by a programmable device, the computer program having a code portion for performing the above-described method for determining the positioning mode when executed by the programmable device.
[0254] The preferred embodiments of this disclosure have been described in detail above with reference to the accompanying drawings. However, this disclosure is not limited to the specific details of the above embodiments. Within the scope of the technical concept of this disclosure, various simple modifications can be made to the technical solutions of this disclosure, and these simple modifications all fall within the protection scope of this disclosure.
[0255] It should also be noted that the various specific technical features described in the above embodiments can be combined in any suitable manner without contradiction. To avoid unnecessary repetition, this disclosure will not describe the various possible combinations separately.
[0256] Furthermore, various different embodiments of this disclosure can be combined in any way, as long as they do not violate the spirit of this disclosure, they should also be regarded as the content disclosed in this disclosure.
Claims
1. A method for determining a positioning pattern, characterized in that, The method is applied to a device for pose localization based on visual inertial odometry, and the method includes: Acquire images captured by the camera; Determine the image type of the image; The image is subjected to monocular image quality detection to obtain monocular image quality inspection results. The monocular image quality inspection results are used to characterize whether monocular visual inertial odometry pose calculation can be completed based on the image. The target localization mode is determined based on the monocular image quality inspection results and the image type. The step of performing monocular image quality detection on the image to obtain the monocular image quality detection result includes: If the image is not an abnormal image and the image meets the monocular feature point extraction conditions, the image is determined to pass the monocular image quality detection, wherein the abnormal image includes at least one of overexposed image, underexposed image, and solid color image; If the image is not the abnormal image and the image does not meet the monocular feature point extraction conditions, or if the image is the abnormal image, it is determined that the image has failed the monocular image quality detection. Wherein, if the image passes the monocular image quality detection, it indicates that at least monocular visual inertial odometry pose calculation can be completed based on the image; if the image fails the monocular image quality detection, it indicates that monocular visual inertial odometry pose calculation cannot be completed based on the image. The step of determining the image type includes: When the number of images is 1, the image type is determined to be a monocular image type; The step of determining the target localization mode based on the monocular image quality inspection result and the image type includes: If the image type is the monocular image type, and the monocular image quality inspection result is that the image passes the monocular image quality inspection, then the monocular positioning mode is determined as the target positioning mode. If the monocular image quality inspection result indicates that the image fails the monocular image quality inspection, then the non-localization mode is determined as the target localization mode; or, When the number of images is 2, the image type of the first image and the second image in the images is determined to be a binocular image type; The step of determining the target localization mode based on the monocular image quality inspection result and the image type includes: When the image type is the binocular image type, the target positioning mode is determined based on the first monocular image quality inspection result of the first image and the second monocular image quality inspection result of the second image.
2. The method according to claim 1, characterized in that, Determining the target localization mode based on the first monocular image quality inspection result of the first image and the second monocular image quality inspection result of the second image includes: If the first monocular image quality inspection result indicates that the first image passes the monocular image quality inspection, and the second monocular image quality inspection result indicates that the second image passes the monocular image quality inspection, then binocular image quality inspection is performed on the first image and the second image to obtain binocular image quality inspection results. The binocular image quality inspection results are used to indicate whether binocular visual inertial odometry pose calculation can be completed based on the first image and the second image. If the first image and the second image pass the binocular image quality detection, the binocular positioning mode is determined as the target positioning mode; If the first image and the second image fail the binocular image quality detection, the monocular positioning mode is determined as the target positioning mode.
3. The method according to claim 1 or 2, characterized in that, Determining the target localization mode based on the first monocular image quality inspection result of the first image and the second monocular image quality inspection result of the second image includes: If the first monocular image quality inspection result indicates that the first image has passed the monocular image quality inspection, and the second monocular image quality inspection result indicates that the second image has not passed the monocular image quality inspection, then the monocular positioning mode is determined as the target positioning mode. If the first monocular image quality inspection result indicates that the first image has failed the monocular image quality inspection, and the second monocular image quality inspection result indicates that the second image has failed the monocular image quality inspection, then the non-localization mode is determined as the target localization mode.
4. The method according to claim 2, characterized in that, The step of performing binocular image quality detection on the first image and the second image to obtain binocular image quality detection results includes: The first image and the second image are preprocessed to obtain a first preprocessed image, a second preprocessed image, and optical flow tracing information between the first preprocessed image and the second preprocessed image. Extract multiple first target feature points from the first preprocessed image; For each first target feature point, based on the optical flow tracing information, a second target feature point corresponding to the first target feature point is determined from the second preprocessed image, and based on the optical flow tracing information, a third target feature point corresponding to the second target feature point is determined from the first preprocessed image. If the distance between the first target feature point and the corresponding third target feature point is less than the first threshold, and the first target feature point and the corresponding second target feature point satisfy the epipolar constraint screening condition, then the first target feature point and the corresponding second target feature point are determined as valid pairing points; If the number of valid pairing points is greater than a second threshold, the first image and the second image are determined to have passed the binocular image quality detection.
5. The method according to claim 4, characterized in that, The camera includes a first camera and a second camera. The method for determining whether the first target feature point and the second target feature point satisfy the epipolar constraint screening condition includes: Based on the intrinsic parameters of the first camera, determine the first coordinates of the first target feature point on the normalized plane corresponding to the first camera; Based on the intrinsic parameters of the second camera, determine the second coordinates of the second target feature point on the normalized plane corresponding to the second camera; Based on the extrinsic parameters between the first camera and the second camera, the epipolar line corresponding to the second coordinate is determined according to the epipolar constraint. When the distance from the first coordinate to the epipolar line is less than the third threshold, it is determined that the first target feature point and the second target feature point satisfy the epipolar constraint screening condition.
6. The method according to claim 1, characterized in that, The methods for determining whether the image meets the monocular feature point extraction conditions include: The image is subjected to histogram equalization to obtain a preprocessed image; The preprocessed image is divided into grids to obtain multiple grid sub-images; From each of the grid sub-images, a set of candidate feature points with response values greater than a preset response threshold is selected, and a set of target feature points is selected from the set of candidate feature points, wherein the distance between each target feature point in the set of target feature points is greater than a fourth threshold. Calculate the mean response value for each set of target feature points, and calculate the standard deviation based on the mean response values of all the target feature points. If the standard deviation is less than the fifth threshold and the total number of target feature points in all target feature point sets is greater than the sixth threshold, then the image is determined to meet the monocular feature point extraction conditions.
7. The method according to claim 1, characterized in that, The method further includes: When the device is in the monocular positioning mode, at preset time intervals, for the most recently acquired image, the step of determining the image type of the image is triggered. When the device is in the non-positioning mode, for each acquired image, the step of determining the image type of the image is triggered.
8. The method according to claim 2, characterized in that, The method further includes: When the device is in the binocular positioning mode, if feature point matching fails during the binocular visual inertial odometry pose calculation based on the first image and the second image, the non-positioning mode is determined as the target positioning mode.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the steps of the method described in any one of claims 1-8.
10. An electronic device, characterized in that, The electronic device includes: A memory on which computer programs are stored; A processor for executing the computer program in the memory to implement the steps of the method according to any one of claims 1-8.