Relocalization methods, apparatus and equipment for cross-modal feature fusion

By employing a cross-modal feature fusion method, geometric-visual joint features are constructed using point cloud data and two-dimensional visual feature descriptors. This solves the problem of pose estimation error in single-sensor systems, achieving efficient and accurate localization and pose estimation. It also supports loop closure detection and relocalization of 3D point cloud maps in visual SLAM systems.

CN120451271BActive Publication Date: 2026-01-06ORIENTAL JUZHI (BEIJING) TECHNOLOGY INNOVATION CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510597667.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2026-01-06
Estimated Expiration
2045-05-09

AI Technical Summary

Technical Problem

In existing technologies, pose estimation by a single sensor is prone to recursive cumulative errors, visual perception is prone to drift errors under 2D-2D geometric constraints, and lidar perception is not robust enough in repetitive structural scenes, has a high mismatch rate, and changes in sensor reliability can lead to matching failures.

Method used

By employing a cross-modal feature fusion method, point cloud data with fused depth information is combined with two-dimensional visual feature descriptors to construct geometric-visual joint features. These features are then matched using a keyframe database or a global point cloud map, and pose calculation is performed using the PNP or RANSAC algorithm.

Benefits of technology

It improves the processing efficiency and accuracy of real-time localization and pose estimation, enables loop closure detection in visual SLAM systems, and performs high-precision localization based on 3D point cloud maps, with strong illumination robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120451271B_ABST
    Figure CN120451271B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a cross-modal feature fusion relocation method, device and equipment, comprising: preprocessing the perception data of the data sensor of the mobile device to obtain point cloud data with fused depth information; performing feature fusion according to the geometric relationship between the point cloud data and the two-dimensional visual feature descriptor of each point cloud to obtain geometric-visual joint features corresponding to geometric patterns; the geometric patterns are composed of multiple point clouds; performing comprehensive matching of the geometric-visual joint features in the real-time constructed key frame database or the pre-constructed global point cloud map database; and performing pose solving on the target pattern pair composed of the target geometric pattern and the target matching pattern according to the successful comprehensive matching to obtain the calibration pose information of the mobile device. The method can be applied to loop detection in a visual SLAM system, and can also be used for map reuse and relocation based on the generated three-dimensional point cloud map, realizing high-precision positioning and having strong light robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of localization and mapping technology, and in particular to a relocalization method, apparatus and device for cross-modal feature fusion. Background Technology

[0002] Simultaneous Localization and Mapping (SLAM) is a key technology in fields such as autonomous driving, intelligent robots, drones, and virtual reality. It relies on sensors to perceive the surrounding environment and estimate pose. However, because visual sensors, laser sensors, and other external relative sensors, pose estimation suffers from recursive cumulative errors. Over long periods of operation, the estimated state deviates from the true value. Therefore, pose estimation correction is necessary.

[0003] In realizing the concept disclosed herein, the inventors discovered at least the following technical problems in the related technologies: In the related technologies, the system often corrects deviations from the pose by using loop closure detection of a single sensor or map reuse relocalization; the visual perception-based approach requires combining historical 3D landmarks to calculate the pose, while the scale uncertainty of pure 2D-2D geometric constraints easily leads to drift errors; the LiDAR perception approach is not robust enough in repetitive structure (e.g., walls, corridors) scenarios, resulting in a high mismatch rate; in some cases, ignoring the dynamic changes in sensor reliability (e.g., sparse LiDAR point clouds or camera image occlusion) can easily lead to matching failures. Summary of the Invention

[0004] To address or at least partially address the aforementioned technical problems, embodiments of this disclosure provide a relocation method, apparatus, and device for cross-modal feature fusion.

[0005] In a first aspect, embodiments of this disclosure provide a relocalization method based on cross-modal feature fusion. The method includes: preprocessing the perceived data from the data sensors of a mobile device to obtain point cloud data with fused depth information; fusing features based on the geometric relationships between the point cloud data and the two-dimensional visual feature descriptors of each point cloud to obtain geometric-visual joint features corresponding to a geometric shape; the geometric shape being composed of multiple point clouds; performing comprehensive matching of the geometric-visual joint features in a real-time constructed keyframe database or a pre-constructed global point cloud map database; and calculating the pose of a target shape pair composed of the successfully matched target geometric shape and the matched target shape to obtain the calibration pose information of the mobile device.

[0006] In some embodiments, the point cloud data is feature point cloud data with fused depth information from the camera imaging plane viewpoint of the mobile device; the feature point cloud data is generated from at least one of visual perception data and radar perception data. Specifically, feature fusion is performed based on the geometric relationships between the point cloud data and the two-dimensional visual feature descriptors of each point cloud to obtain the geometric-visual joint features corresponding to the geometric shape. This includes: performing position matching and association between the feature point cloud data with fused depth information and the two-dimensional visual feature descriptors extracted from image pixels from the camera imaging plane viewpoint; selecting point clouds based on the position matching and association results to construct geometric shapes and perform feature fusion to obtain the geometric-visual joint features corresponding to the geometric shapes.

[0007] In some embodiments, point clouds are selected based on the results of location matching and association to construct geometric shapes for feature fusion, resulting in geometric-visual joint features corresponding to the geometric shapes, including:

[0008] For the first point cloud object whose position matches the feature point cloud data with fused depth information and the corresponding two-dimensional visual feature descriptor, select the first point cloud object to construct a geometric figure;

[0009] For a second point cloud object whose feature point cloud data with fused depth information does not match the position corresponding to the two-dimensional visual feature descriptor, a position nearest neighbor association is performed based on the first position distribution of the feature point cloud data and the second position distribution of the point cloud corresponding to the two-dimensional visual feature descriptor, and a geometric figure is constructed based on the result of the position nearest neighbor association.

[0010] The constructed geometric figure is spliced ​​according to at least the following dimensions: point cloud position, edge length, and the two-dimensional visual feature descriptor corresponding to the point cloud contained therein, to obtain the geometric-visual joint feature corresponding to the geometric figure.

[0011] In some embodiments, positional nearest neighbor association is performed based on the first positional distribution of the feature point cloud data and the second positional distribution of the corresponding point cloud of the two-dimensional visual feature descriptor, and a geometric figure is constructed based on the result of the positional nearest neighbor association, including:

[0012] When the first position distribution of the first target point cloud and the second position distribution of the second target point cloud are within a preset nearest neighbor range in the second point cloud object, the point clouds within the preset nearest neighbor range in the first and second target point clouds are marked as first associated point clouds. The multidimensional features of the first associated point clouds are shared within the preset nearest neighbor range; geometric figures are constructed by selecting the first associated point clouds within the preset nearest neighbor range; or...

[0013] When there is a positional overlap between the first positional distribution of the first target point cloud and the second positional distribution of the second target point cloud in the second point cloud object, an expansion within a preset neighborhood is performed based on the positional overlap in the first target point cloud and the second target point cloud, and the point cloud in the preset neighborhood is marked as the second associated point cloud. The multidimensional features of the second associated point cloud are shared within the preset neighborhood; the second associated point cloud is selected in the preset neighborhood to construct a geometric figure.

[0014] In some embodiments, the above-mentioned selection of point clouds to construct geometric figures includes: arbitrarily selecting a preset number of point cloud data to form a set of candidate geometric figures; sorting the edges between point clouds in each candidate geometric figure according to the edge size in the set; determining in sequence whether any edge in each candidate geometric figure exceeds a set threshold according to the edge sorting; removing candidate geometric figures with edges exceeding the set threshold, and constructing the remaining candidate geometric figures in the set as the geometric figures to be matched.

[0015] In some embodiments, the comprehensive matching includes: a first matching for the geometric relationship and a second matching for the two-dimensional visual feature descriptor, wherein the first matching and the second matching are performed in a preset order.

[0016] In some embodiments, the geometric figure is a triangle formed by arbitrarily selecting three point clouds; the geometric-visual joint features corresponding to the geometric figure include: the positions of the three point cloud vertices in the triangle, the lengths of the three connecting sides, and the two-dimensional visual feature descriptors corresponding to the three point cloud vertices; in the keyframe database or the global point cloud map database, the geometric-visual joint features of each reference triangle are stored based on the storage hash index obtained by processing the diagonal value of the longest side. Specifically, in a real-time constructed keyframe database or a pre-constructed global point cloud map database, the above-mentioned geometric-visual joint feature comprehensive matching is performed, including: for each triangle to be matched, the following comprehensive matching processing operations are performed: the diagonal value of the longest side of the current triangle to be matched is processed to obtain the hash index to be matched; the hash index to be matched is matched with the storage hash index to locate the target storage area; within the target storage area, angle matching is performed based on the diagonal value of the longest side of the current triangle to be matched; for the first set of matching reference triangles obtained by successful angle matching, side length matching is performed sequentially based on the longest and second longest sides of the current triangle to be matched; for the second set of matching reference triangles obtained by successful side length matching, feature matching is performed based on the two-dimensional visual feature descriptors of each point cloud vertex of the triangle to be matched; and a target graphic pair is generated based on the target matching reference triangle obtained by successful feature matching and the current triangle.

[0017] In some embodiments, the calibration pose information of the mobile device is obtained by performing pose calculation on the target graphic pair consisting of the successfully matched target geometry and the target matching graphic. This includes: calling the PNP algorithm to perform pose calculation based on the correspondence between the target geometry and the target matching graphic in the target graphic pair, thereby obtaining the calibration pose information of the mobile device at the target time; or, in the case of multiple target graphic pairs, calling the RANSAC algorithm to remove mismatched pairs, and performing pose optimization on the filtered matching pairs based on a nonlinear optimization algorithm, thereby obtaining the calibration pose information of the mobile device at the target time.

[0018] In some embodiments, the aforementioned perception data includes at least one of the following: visual perception data and radar perception data; wherein the aforementioned visual perception data includes at least one of the following: image data captured by a depth camera and image data captured by a multi-view camera. The aforementioned preprocessing of the perception data from the data sensors to obtain point cloud data with fused depth information includes at least one of the following: extracting and describing pixel features from the image data captured by the depth camera, and projecting the pixels onto a three-dimensional space according to their corresponding depth values ​​to obtain feature point cloud data with fused depth information; or, extracting and describing pixel features from the image data captured by the multi-view camera, and generating depth based on the stereo vision corresponding to the multi-view cameras and projecting it onto a three-dimensional space to obtain feature point cloud data with fused depth information; or, using the point cloud data from the radar perception data as a three-dimensional feature point cloud, projecting the three-dimensional feature point cloud back onto the imaging plane corresponding to the camera and performing feature description to obtain feature point cloud data with fused depth information.

[0019] Secondly, embodiments of this disclosure provide a cross-modal feature fusion relocalization device. This relocalization device is integrated into a mobile device or is an independent entity capable of communication with the mobile device. The relocalization device includes: a point cloud data generation module, a feature fusion module, a comprehensive matching module, and a pose calculation module. The point cloud data generation module preprocesses the sensor data of the mobile device to obtain point cloud data with fused depth information. The feature fusion module fuses features based on the geometric relationships between the point cloud data and the two-dimensional visual feature descriptors of each point cloud to obtain geometric-visual joint features corresponding to a geometric shape; the geometric shape is composed of multiple point clouds. The comprehensive matching module performs comprehensive matching of the geometric-visual joint features in a real-time constructed keyframe database or a pre-constructed global point cloud map database. The pose calculation module calculates the pose of a target shape formed by the successfully matched target geometric shape and the matched target shape to obtain the calibration pose information of the mobile device.

[0020] Thirdly, embodiments of this disclosure provide a mobile device. The mobile device integrates the repositioning device provided in the second aspect of the embodiments; or, the mobile device includes: a first point cloud data generation module, a first feature fusion module, a first comprehensive matching module, and a first pose calculation module. The first point cloud data generation module is used to preprocess the perceived data from the mobile device's data sensors to obtain point cloud data with fused depth information. The first feature fusion module is used to perform feature fusion based on the geometric relationships between the point cloud data and the two-dimensional visual feature descriptors of each point cloud to obtain geometric-visual joint features corresponding to the geometric figure; the geometric figure is composed of multiple point clouds. The first comprehensive matching module is used to perform comprehensive matching of the geometric-visual joint features in a real-time constructed keyframe database or a pre-constructed global point cloud map database. The first pose calculation module is used to perform pose calculation based on the target geometric figure successfully matched and the target matching figure forming a target figure pair to obtain the calibration pose information of the mobile device.

[0021] In some embodiments, the mobile device is at least one of the following: a smart robot, a vehicle, a flight device, or a wearable device.

[0022] Fourthly, embodiments of this disclosure provide a service device. This service device provides service support for the positioning of a mobile device, and integrates the repositioning device provided in the second aspect of the embodiments; or, the service device includes: a second point cloud data generation module, a second feature fusion module, a second comprehensive matching module, and a second pose calculation module. The second point cloud data generation module preprocesses the perceived data from the mobile device's data sensors to obtain point cloud data with fused depth information. The second feature fusion module performs feature fusion based on the geometric relationships between the point cloud data and the two-dimensional visual feature descriptors of each point cloud to obtain geometric-visual joint features corresponding to a geometric shape; the geometric shape is composed of multiple point clouds. The second comprehensive matching module performs comprehensive matching of the geometric-visual joint features in a real-time constructed keyframe database or a pre-constructed global point cloud map database. The second pose calculation module performs pose calculation based on the target geometric shape and the target matched shape forming a target shape pair to obtain the calibration pose information of the mobile device.

[0023] Fifthly, embodiments of this disclosure provide an electronic device. The electronic device includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, communication interface, and memory communicate with each other via the communication bus; the memory stores computer programs; and the processor, when executing the program stored in the memory, implements the cross-modal feature fusion relocation method as described above.

[0024] Sixthly, embodiments of this disclosure provide a computer-readable storage medium. The computer-readable storage medium stores a computer program that, when executed by a processor, implements the cross-modal feature fusion relocation method as described above.

[0025] The technical solutions provided in the embodiments of this disclosure have at least some or all of the following advantages:

[0026] By leveraging the geometric relationships between point cloud data incorporating depth information and fusing features with the 2D visual feature descriptors of each point cloud, multi-dimensional fusion of 3D geometric features and 2D visual feature descriptors is achieved. Based on the simultaneous multi-dimensional feature description of point clouds using 3D geometric features incorporating depth information and 2D visual feature descriptors, this method enables location identification during subsequent matching with a keyframe database or global point cloud map database. Matching based on the geometric-visual joint features corresponding to the geometric shapes overcomes the inability to distinguish similar scenes such as white walls and corridors when using simple geometric feature matching. Furthermore, since the matching of 2D visual feature descriptors also incorporates the depth information corresponding to 3D geometric features, redundancy is reduced during matching of numerous point cloud locations, improving matching efficiency and accuracy, and avoiding drift errors caused by the scale uncertainty of pure 2D-2D geometric constraints. Overall, this method improves the processing efficiency and accuracy of real-time localization and pose estimation. It can not only be applied to loop closure detection in visual SLAM systems but also enables map reuse and relocalization based on pre-generated 3D point cloud maps (supporting color point cloud maps), achieving high-precision localization with strong illumination robustness. Attached Figure Description

[0027] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0028] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, those skilled in the art can obtain other drawings based on these drawings without creative effort.

[0029] Figure 1 A flowchart of a relocation method for cross-modal feature fusion according to an embodiment of the present disclosure is illustrated schematically.

[0030] Figure 2 A detailed implementation flowchart of step S120 according to an embodiment of the present disclosure is shown schematically.

[0031] Figure 3 A detailed implementation flowchart of step S220 according to an embodiment of the present disclosure is shown schematically.

[0032] Figure 4 The diagram illustrates a comparison between the position recognition and loop closure detection results obtained by processing a real-world radar-visual dataset using (a) the relocation method provided in this embodiment and (b) the position recognition and loop closure detection results obtained using only the STD triangle descriptor.

[0033] Figure 5 A structural block diagram of a relocation device for cross-modal feature fusion according to an embodiment of the present disclosure is shown schematically.

[0034] Figure 6 A schematic block diagram of an electronic device provided in an embodiment of the present disclosure is shown. Detailed Implementation

[0035] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this disclosure. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0036] During the research and development process, it was found that in related technologies, positioning systems often correct deviations in pose by using loop closure detection from a single sensor or map reuse relocalization; visual perception-based methods require combining historical 3D landmarks to calculate pose, but the scale uncertainty of pure 2D-to-2D geometric constraints can easily cause drift errors; LiDAR perception methods are not robust enough in repetitive structural scenarios (such as walls and corridors), resulting in a high mismatch rate; in some cases, ignoring dynamic changes in sensor reliability (such as sparse LiDAR point clouds or camera image occlusion) can also easily lead to matching failures.

[0037] For example, some LiDAR sensing methods rely on the geometric features of LiDAR point cloud data without considering visual information. In some scenarios, relying solely on geometric features may not be effective in distinguishing locations with similar geometric structures, such as corridors and walls, leading to a high false-match rate. Furthermore, when applied to vision systems, this method lacks the ability to directly locate within a 3D point cloud map based on visual observation. The vision system cannot directly correlate a simple 3D point cloud with the current image, thus hindering pose optimization (or relocalization).

[0038] For example, some purely visual perception methods encode based on extracted ORB features and a bag-of-words model. This approach relies heavily on the training quality of the bag-of-words model: building such a model requires a large amount of training data, and the training quality directly affects the accuracy of loop closure detection. If the training data differs significantly from the actual application scenario, it may lead to a decrease in detection performance.

[0039] In purely visual perception methods, location identification and matching based on ORB features typically involves extracting higher-dimensional features, such as using bag-of-words models or semantic recognition, due to the massive amount of point cloud data. This avoids using binary descriptors corresponding to location points directly, as this would result in extremely high computational costs and very low retrieval efficiency, and is generally not supported by existing vision systems. In other words, in purely visual perception methods, due to the enormous amount of point cloud data, high-dimensional feature extraction is always necessary before utilization; directly using the point cloud data would cause the entire system to crash and fail to perform location matching calculations. Furthermore, purely visual perception solutions cannot be directly reused for 3D point cloud maps; for example, maps generated from laser point clouds cannot be reused, limiting their applicability.

[0040] Furthermore, since visual perception methods mostly rely on two-dimensional image features for location recognition, the lack of depth information makes it difficult to be directly compatible with three-dimensional point cloud maps constructed by LiDAR. It is necessary to calculate the pose through relevant algorithms and in combination with historical three-dimensional landmarks, and the scale uncertainty of pure 2D-2D geometric constraints can easily cause drift errors.

[0041] In view of this, embodiments of this disclosure provide a cross-modal feature fusion relocalization method, apparatus, and device, which fuses three-dimensional geometric features with two-dimensional visual feature descriptors in a multi-dimensional manner. This not only applies to loop closure detection in visual SLAM systems but also enables map reuse and relocalization based on pre-generated three-dimensional point cloud maps (supporting color point cloud maps), achieving high-precision localization. The method provided in this disclosure integrates the advantages of both three-dimensional geometric features and two-dimensional visual features through hybrid hash matching, improving the processing efficiency and accuracy of real-time localization and pose estimation. It also exhibits strong illumination robustness and supports the reuse of pure geometric point cloud maps (such as maps generated by LiDAR) and direct matching of color point cloud maps (such as color point cloud maps generated by depth cameras (RGB-D), LV-SLAM (LiDAR-Vision-Inertial Odometry) systems, etc.).

[0042] By simultaneously describing the multi-dimensional features of point clouds using 3D geometric features incorporating depth information and 2D visual feature descriptors, this method can overcome the inability to distinguish similar scenes such as white walls and corridors when matching with keyframe databases or global point cloud map databases for location recognition. This is achieved by matching based on the geometric-visual joint features corresponding to the geometric shapes. Furthermore, since the matching of 2D visual feature descriptors also incorporates the depth information corresponding to 3D geometric features, redundancy is reduced during the matching of numerous point cloud locations, improving matching efficiency and accuracy. This avoids the drift error easily caused by the scale uncertainty of pure 2D-2D geometric constraints. Overall, this method improves the processing efficiency and accuracy of real-time localization and pose estimation. It can be applied not only to loop closure detection in visual SLAM systems but also to map reuse and relocalization based on pre-generated 3D point cloud maps (supporting color point cloud maps), achieving high-precision localization with strong illumination robustness.

[0043] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.

[0044] The first exemplary embodiment of this disclosure provides a relocalization method based on cross-modal feature fusion. The method provided in this embodiment can be applied to an electronic device with computing capabilities. This electronic device can be integrated into a mobile device, or it can be a device independent of and capable of communicating with the mobile device, used to provide computing services for the positioning of the mobile device.

[0045] Figure 1 A flowchart of a relocation method for cross-modal feature fusion according to an embodiment of the present disclosure is illustrated schematically.

[0046] Reference Figure 1 As shown, the cross-modal feature fusion relocation method provided in this embodiment includes the following steps: S110, S120, S130 and S140.

[0047] In step S110, the perceived data from the mobile device's data sensors is preprocessed to obtain point cloud data with fused depth information.

[0048] In some embodiments, the data sensors may encompass one or more types of Time-of-Flight (TOF) sensors (which measure distance by measuring the time required for light, sound waves, or other types of signals to travel from transmission to reception), such as optical, acoustic, and radar sensors; visual sensors; and radar. The perception data corresponding to the mobile device may be, but is not limited to, one or more of the following: image data captured by a depth camera (RGB-D camera) (an example of visual perception data); image data captured by a multi-view camera (another example of visual perception data); and radar perception data. In some embodiments, the radar perception data may be sparse point clouds output by LV-SLAM (LiDAR-Vision-Inertial Odometry), or point cloud data output by one or more radars (e.g., LiDAR, millimeter-wave radar, etc.).

[0049] In some embodiments, the above preprocessing includes, but is not limited to, point cloud generation, plane segmentation, feature extraction, and feature description.

[0050] In step S110 above, the image data captured by the depth camera is preprocessed with the sensor data to obtain point cloud data fused with depth information, including:

[0051] Pixel feature points are extracted and feature descriptions are performed on image data captured by a depth camera, and the pixels are projected into three-dimensional space according to their corresponding depth values ​​to obtain feature point cloud data with fused depth information.

[0052] During the capture of a depth camera, each pixel will have a corresponding depth value. Based on the image data captured by the depth camera, pixel feature points are extracted and feature descriptions are performed. Combined with the depth values, the data is projected into a three-dimensional space, which can obtain feature point cloud data with fused depth information from the camera imaging plane view of the mobile device.

[0053] In step S110 above, the image data captured by the multi-view camera is preprocessed with the sensor data to obtain point cloud data with fused depth information, including:

[0054] Pixel feature points are extracted and feature descriptions are performed on image data captured by multi-view cameras. Based on the stereo vision corresponding to the multi-view cameras, depth is generated and projected into three-dimensional space to obtain feature point cloud data with fused depth information.

[0055] Multi-view stereo vision composed of multi-view cameras can generate pixel-corresponding depth. Based on the image data captured by the multi-view cameras, pixel feature points are extracted and feature descriptions are performed, and combined with depth values ​​and projected into three-dimensional space. Feature point cloud data with fused depth information can be obtained from the camera imaging plane view of the mobile device.

[0056] In step S110 above, the radar sensing data is preprocessed to obtain point cloud data with fused depth information, including:

[0057] Point cloud data from radar sensing data is used as three-dimensional feature point cloud. The three-dimensional feature point cloud is then projected back onto the imaging plane of the camera and its features are described to obtain feature point cloud data with fused depth information.

[0058] Point cloud data in radar perception data is a three-dimensional feature point cloud carrying depth information. It is the perception result in the world coordinate system. By projecting the three-dimensional feature point cloud back to the imaging plane of the camera, the corresponding position under the camera imaging plane view is obtained and the features are described. Feature point cloud data with fused depth information can be obtained under the camera imaging plane view of the mobile device.

[0059] In step S120, feature fusion is performed based on the geometric relationship between the point cloud data and the two-dimensional visual feature descriptor of each point cloud to obtain the geometric-visual joint features corresponding to the geometric figure; the geometric figure is composed of multiple point clouds.

[0060] In some embodiments, the point cloud data is feature point cloud data that incorporates depth information from the camera imaging plane view of the mobile device; the feature point cloud data is generated from at least one of visual perception data and radar perception data.

[0061] The information corresponding to the above two-dimensional visual feature descriptors can be various types of information describing point cloud features in a two-dimensional perspective. These can be CNN (convolutional neural network) features of 64, 128, or 256 dimensions (which can represent differences in point clouds with color and grayscale), or various binary descriptors, or other improved two-dimensional feature description information.

[0062] Since the two-dimensional visual feature descriptors extracted from image pixels under the camera imaging plane viewpoint are used as another dimension of point cloud features, geometric figures are constructed and features are fused after position matching and association with the previously obtained feature point cloud data, thus realizing the multi-dimensional fusion of three-dimensional geometric features and two-dimensional visual feature descriptors and subsequent comprehensive matching.

[0063] By selecting several points from the feature point cloud data to form a geometric shape, different types of geometric shapes can be obtained by varying the number of points selected. These geometric shapes can include, but are not limited to, triangles, quadrilaterals, pentagons, hexagons, and other polygons, or even irregular shapes. To reduce the amount of matching computation, regular shapes are preferred. Keeping the number of points constant, changing the combination of the selected point positions can yield multiple geometric shapes of the same type.

[0064] Three-dimensional geometric features are obtained by describing the geometric relationships of various geometric figures. These geometric relationships include, but are not limited to, at least one of the following: vertex position, side length, angle distribution, and the correspondence between sides and angles. Considering that matching based solely on three-dimensional geometric features in a database can be affected by factors such as lighting and occlusion, this embodiment of the present disclosure fuses features based on the geometric relationships between point cloud data and the two-dimensional visual feature descriptors of each point cloud to obtain the geometric-visual joint features corresponding to the geometric figures. This achieves multi-dimensional fusion of three-dimensional geometric features and two-dimensional visual feature descriptors. Based on three-dimensional geometric features incorporating depth information, multi-dimensional feature description of the point cloud is performed simultaneously with two-dimensional vision. Therefore, when subsequently matching with a keyframe database or a global point cloud map database for location recognition, comprehensive matching of multi-dimensional features is required.

[0065] Taking a triangle as an example of a geometric figure, the geometric-visual joint feature of a triangle is represented as [p, l, I], where p represents the position p(p1, p2, p3) of the three point cloud vertices that make up the triangle; l represents the side length l(l1, l2, l3) of the three point cloud vertices that make up the triangle; and I represents the two-dimensional visual feature descriptor I(I1, I2, I3) corresponding to the three point cloud vertices that make up the triangle.

[0066] Figure 2 A detailed implementation flowchart of step S120 according to an embodiment of the present disclosure is shown schematically.

[0067] In some embodiments, refer to Figure 2 As shown, in step S120 above, feature fusion is performed based on the geometric relationship between point cloud data and the two-dimensional visual feature descriptor of each point cloud to obtain the geometric-visual joint features corresponding to the geometric figure, including the following steps: S210 and S220.

[0068] In step S210, position matching and association are performed between the feature point cloud data with fused depth information and the two-dimensional visual feature descriptors extracted from image pixels from the camera imaging plane viewpoint.

[0069] In some embodiments, the following situation may exist between the position (also described as the point cloud position) of the two-dimensional visual feature descriptor extracted from the image pixel under the camera imaging plane view and the point cloud position corresponding to the feature point cloud data fused with depth information: some point cloud positions match and other point cloud positions do not match.

[0070] The position matching process involves matching the point cloud positions corresponding to two modalities. One modality is the position corresponding to the feature point cloud data obtained by fusing depth information, and the other modality is the two-dimensional visual feature descriptor extracted from the image pixels (which can be image pixels in an image obtained based on visual perception data) from the camera imaging plane viewpoint. The corresponding point cloud positions are matched based on the two modalities. For the matched parts, the subsequent step S311 is executed to construct the geometric figure. For the mismatched parts, position association processing is performed, and the subsequent step S312 is executed to construct the geometric figure as well. Then, step S320 is executed.

[0071] As described in step S110 above, the point cloud data is feature point cloud data with fused depth information from the camera imaging plane view of the mobile device. Therefore, the geometric relationships are also the edge, corner, and other relationships between feature point cloud data with fused depth information from the camera imaging plane view. By incorporating a two-dimensional visual feature descriptor from another modality into the edge and corner relationships of the feature point cloud data with fused depth information and performing position matching and association to construct a geometric figure for feature fusion, the point cloud features in the imaging plane can be comprehensively and complementaryly represented from multiple dimensions. On the one hand, since the geometric-visual joint feature fusion integrates three-dimensional geometric features with depth information and features that can characterize color and light, it can achieve a more comprehensive and complementary representation of point cloud features in the imaging plane. The system utilizes two-dimensional feature descriptions, such as line differences, to support the reuse of pure geometric point cloud maps (e.g., LiDAR-generated maps) and direct matching of color point cloud maps (e.g., color point cloud maps generated by depth cameras (RGB-D), LV-SLAM (LiDAR-Vision-Inertial Odometry) systems, etc.). When encountering sparse LiDAR point clouds, it performs comprehensive matching of geometric and visual joint features, enabling accurate location recognition and matching based on the modalities corresponding to the two-dimensional visual feature descriptors. When encountering camera image occlusion, feature fusion allows for the mutual complementarity of two-dimensional vision and geometry, resulting in more comprehensive features and reducing the probability of mismatches during comprehensive matching.

[0072] In step S220, point clouds are selected to construct geometric figures based on the results of position matching and association, and feature fusion is performed to obtain the geometric-visual joint features corresponding to the geometric figures.

[0073] Figure 3 A detailed implementation flowchart of step S220 according to an embodiment of the present disclosure is shown schematically.

[0074] In some embodiments, refer to Figure 3As shown, in step S220 above, point cloud is selected to construct geometric figures based on the results of position matching and association, and feature fusion is performed to obtain the geometric-visual joint features corresponding to the geometric figures. This includes the following steps: S311, S312, and S320. Steps S311 and S312 are parallel steps, and one can be selected for execution. In some scenarios, both S311 and S312 may be included.

[0075] In step S311, for the first point cloud object whose position matches the feature point cloud data with fused depth information and the two-dimensional visual feature descriptor, the first point cloud object is selected to construct a geometric figure.

[0076] For example, the feature point cloud data with fused depth information is described as a set M, and the point cloud positions corresponding to the two-dimensional visual feature descriptors extracted from the image pixels under the camera imaging plane view are described as N. There are some point cloud position sets Mc with fused depth information that completely overlap with the point cloud position sets Nc corresponding to the two-dimensional visual feature descriptors. These are considered to be position matches. Then Mc and Nc can be regarded as the first point cloud objects, or each point cloud in Mc and Nc can be regarded as the first point cloud object, and point clouds are selected from the set DY1 composed of the first point cloud objects to construct geometric figures.

[0077] In step S312, for the second point cloud object whose position does not match the feature point cloud data with fused depth information and the corresponding position of the two-dimensional visual feature descriptor, position nearest neighbor association is performed based on the first position distribution of the feature point cloud data and the second position distribution of the point cloud corresponding to the two-dimensional visual feature descriptor, and a geometric figure is constructed based on the result of position nearest neighbor association.

[0078] Continuing with the example above, there are also instances where the positions in the point cloud location set Md, which incorporates some depth information, do not completely overlap with the positions in the point cloud location set Nd corresponding to the two-dimensional visual feature descriptor. These are considered to be position mismatches, and Md and Nd can be regarded as second point cloud objects, or each point cloud in Md and Nd can be regarded as a second point cloud object.

[0079] In some embodiments, step S312 above, which involves performing positional nearest neighbor association based on the first positional distribution of the feature point cloud data and the second positional distribution of the corresponding point cloud of the two-dimensional visual feature descriptor, and constructing a geometric figure based on the result of the positional nearest neighbor association, includes:

[0080] When the first position distribution of the first target point cloud and the second position distribution of the second target point cloud are within a preset nearest neighbor range in the second point cloud object, the point clouds within the preset nearest neighbor range in the first and second target point clouds are marked as first associated point clouds. The multidimensional features of the first associated point clouds are shared within the preset nearest neighbor range; geometric figures are constructed by selecting the first associated point clouds within the preset nearest neighbor range; or...

[0081] When there is a positional overlap between the first positional distribution of the first target point cloud and the second positional distribution of the second target point cloud in the second point cloud object, an expansion within a preset neighborhood is performed based on the positional overlap in the first target point cloud and the second target point cloud, and the point cloud in the preset neighborhood is marked as the second associated point cloud. The multidimensional features of the second associated point cloud are shared within the preset neighborhood; the second associated point cloud is selected in the preset neighborhood to construct a geometric figure.

[0082] In the above embodiments, some feature point cloud data in the second point cloud object are in a close neighbor relationship with the second position distribution of the point cloud corresponding to the two-dimensional visual feature descriptor, or have a close neighbor relationship with partially overlapping positions. By marking point clouds within a preset close neighbor range as associated point clouds (corresponding to the first associated point cloud or the second associated point cloud), and sharing multiple features of associated point clouds within the close neighbor range, it is possible to effectively avoid the situation where single-modal feature data points are missing in some scenarios. Based on the sharing and comprehensive matching of multimodal feature close neighbors, the problem of mismatch can be reduced.

[0083] In some embodiments, step S220 above, selecting point clouds to construct geometric figures, includes:

[0084] A set of candidate geometric figures is formed by arbitrarily selecting a preset number of point cloud data points; the preset number can be different, such as 3, 4, 5, or 6, to form candidate geometric figures such as triangles, quadrilaterals, pentagons, and hexagons. A set containing multiple candidate geometric figures is obtained by combining point clouds from different locations. In this set, some candidate geometric figures may share edges or points. In fact, in the randomly generated method, there are a large number of candidate geometric figures sharing edges and points.

[0085] In the above set, the edges connecting the point clouds in each candidate geometry are sorted according to the size of the edges;

[0086] Based on the order of the connected edges, determine in sequence whether any connected edge in each candidate geometry exceeds the set threshold;

[0087] Candidate geometric figures with more than a set threshold of connected edges are removed, and the remaining candidate geometric figures in the above set are used to construct the geometric figures to be matched.

[0088] In this embodiment, based on a set threshold, the candidate geometric figures in the above set are screened based on the edge connection size, and the candidate geometric figures with any edge connection exceeding the set threshold are excluded. The remaining candidate geometric figures in the above set are determined as the geometric figures to be matched, which can reduce the number of geometric figures to be matched and improve the matching efficiency. At the same time, since the geometric figures corresponding to the overly long side lengths can actually be decomposed into geometric figures with smaller side lengths during matching, when there are a large number of shared edges and shared points among the candidate geometric figures in the above set, data redundancy can be reduced by excluding the geometric figures with overly long side lengths, and matching is performed only based on geometric figures with finer granularity, effectively improving the matching efficiency and reducing data redundancy.

[0089] In step S320, the constructed geometric figures are at least feature - stitched according to the following dimensions: point cloud position, edge connection side length, and the two - dimensional visual feature descriptor corresponding to the contained point cloud, to obtain the geometric - visual joint feature corresponding to the geometric figure.

[0090] Taking a triangle as an example of a geometric figure, p represents the positions p(p1, p2, p3) of the three point cloud vertices forming the triangle, and each position can be expressed using three - dimensional coordinates x, y, z; l represents the three side lengths l(l1, l2, l3) of the three point cloud vertices forming the triangle. In some embodiments, each triangle is arranged in ascending order of side length. For example, there is the following sorting relationship: l1 ≤ l2 < l3; I represents the two - dimensional visual feature descriptors I(I1, I2, I3) corresponding to the three point cloud vertices forming the triangle. The position of each point cloud vertex and the two - dimensional visual feature descriptor are corresponding one by one in the order of side length, where the point corresponds to the opposite side; then the geometric - visual joint feature of the triangle is represented as [p, l, I].

[0091] In step S130, in the key - frame database constructed in real - time or the pre - constructed global point cloud map database, the above - mentioned comprehensive matching of geometric - visual joint features is performed.

[0092] In some application scenarios of the present disclosure, for in - process analysis, when performing simultaneous localization and mapping processing on a mobile device during movement, steps S110 to S140 are executed. The perception data in step S110 is real - time perception data, that is, the time stamp corresponding to each perception data is the current moment; in step S130, it can be a key - frame database constructed in real - time or a pre - constructed global point cloud map database (for example, in an offline mode, the pre - constructed global point cloud map database is imported, and mapping is no longer performed during the real - time movement of the mobile device, but only the pose is determined). This method can correspond to the loop - closure detection mode: constructing a key - frame database in real - time, storing the geometric - visual joint features of geometric figures and the corresponding pose T.

[0093] In other application scenarios, post-event analysis is performed after the mobile device completes its task or performs mobile mapping. The analysis focuses on the mobile device's behavior during movement. The sensing data in step S110 is historical sensing data, and the timestamp corresponding to each sensing data point is a specific historical moment. Step S130 can utilize a pre-built global point cloud map database. This approach corresponds to a map reuse and relocalization mode: a global point cloud map database is pre-generated, storing the geometric-visual joint features of geometric figures and their corresponding poses T.

[0094] The above types of databases are described as: D{P i |P i =(l i ,I i ,p i ,T i ),i∈n},P i This represents the matching reference information stored in the database; i represents the frame number; n is the total number of frames; T i This represents the pose corresponding to the i-th frame. Based on the matching reference information in the database (described as matching pair information), it can be distinguished from the joint features of the point cloud to be matched by adding the subscript target.

[0095] In some embodiments, the comprehensive matching includes: a first matching for the geometric relationship and a second matching for the two-dimensional visual feature descriptor, wherein the first matching and the second matching are performed in a preset order.

[0096] In some embodiments, the geometric figure is a triangle formed by arbitrarily selecting three point clouds; the geometric-visual joint features corresponding to the geometric figure include: the positions of the three point cloud vertices in the triangle, the lengths of the three connecting edges, and the two-dimensional visual feature descriptors corresponding to the three point cloud vertices; in the keyframe database or the global point cloud map database, the geometric-visual joint features and corresponding poses (or matching pair information) of each reference triangle are stored based on the storage hash index obtained by processing the diagonal value of the longest side.

[0097] In step S130 above, the comprehensive matching of the aforementioned geometric-visual joint features is performed in the real-time constructed keyframe database or the pre-constructed global point cloud map database, including:

[0098] For each triangle to be matched, perform the following comprehensive matching operation:

[0099] The diagonal value of the longest side of the current triangle to be matched is processed to obtain the hash index of the triangle to be matched;

[0100] Match the hash index to be matched with the storage hash index to locate the target storage range;

[0101] Within the aforementioned target storage area, angle matching is performed based on the diagonal value of the longest side of the current triangle to be matched;

[0102] For the first set of matching reference triangles obtained from successful angle matching, side length matching is performed sequentially based on the longest and second longest sides of the current triangle to be matched;

[0103] For the second set of matching reference triangles obtained by successful side length matching, feature matching is performed based on the two-dimensional visual feature descriptors of each point cloud vertex of the triangle to be matched;

[0104] Based on the target matching reference triangle obtained from successful feature matching and the current triangle, a target graphic pair is generated.

[0105] In this embodiment, to accelerate overall computation and query efficiency, the cosine value of the diagonal θ of the longest side is stored. This cosine value satisfies the following expression:

[0106]

[0107] In the example using a triangle as the geometric figure, the positions of the three point cloud vertices are p1, p2, and p3, respectively. The opposite side of point cloud vertex p1 is l1, the opposite side of point cloud vertex p2 is l2, and the opposite side of point cloud vertex p3 is l3. The diagonal value θ of the longest side falls in the angle interval (0, 180°), and cos(θ) is in the monotonically decreasing interval (-1, 1).

[0108] The aforementioned diagonal cosine values ​​are quantized to 0.01 (the specific value is for example and can vary) to form a storage hash index, which can generate 201 intervals (including boundaries). Each interval corresponds to a storage unit. Some preset angle ranges can be described as storage intervals, and matching pairs with the same hash index are stored in the corresponding storage intervals.

[0109] During the query, the hash value corresponding to the cosine value of the largest angle is first calculated and matched within the storage range. After locating the target storage range, a fine-grained angle query is performed: angle matching is performed based on the diagonal value of the longest side of the current triangle to be matched; if the current query is for the cosine angle cos(θ)... q (q represents the current triangle's index) and the target reference cosine angle cos(θ) in the database. target The difference between the two angles is less than the angle matching threshold thr, and the specific expression is as follows:

[0110] cos(θ q )-cos(θ target ) <thr, (2)

[0111] The first set of matching reference triangles is obtained, within which the longest side of the current triangle to be matched is selected. and the second longest side The side lengths are matched sequentially, as shown in the following expression:

[0112]

[0113] If the longest side Matching reference information with a certain longest edge The result of the matching is less than the longest edge matching threshold. Then continue with the second longest edge. The specific expression for matching the side length is as follows:

[0114]

[0115] If the second longest side Matching reference information with a certain second longest edge The result of the matching is less than the second longest side matching threshold. The second matching reference triangle set is obtained; the two-dimensional visual feature descriptor I is then applied to this set. q Feature matching, the specific expression is as follows:

[0116]

[0117] If the two-dimensional visual feature descriptor I corresponding to the i-th vertex of the current triangle q-i Two-dimensional visual feature descriptor I for corresponding vertices of a certain reference triangle. target-i The sum of the squared differences between the two values, calculated by adding the values ​​according to the number of vertices, is less than the feature matching threshold r. thr If the feature matching result is considered to be a successful match, then the target triangle and the target matching reference triangle can be obtained as a target graphic pair.

[0118] Although the above example uses a triangle as a geometric shape, the solution of this disclosure can be applied to other geometric shapes. For triangles, multiple storage intervals can be divided and indexed by performing hash processing based on angles, so as to improve the overall computing efficiency and query efficiency, which has a positive effect on the processing time of real-time graph building.

[0119] In step S140, the pose calculation is performed on the target graphic pair formed by the successfully matched target geometry and the target matching graphic to obtain the calibration pose information of the mobile device.

[0120] In some embodiments, in step S140 above, the pose calculation is performed on the target graphic pair formed by the successfully matched target geometry and the target matching graphic to obtain the calibration pose information of the mobile device, including:

[0121] The PNP algorithm is invoked to calculate the pose based on the correspondence between the target geometry and the target matching geometry in the target image alignment, thereby obtaining the calibration pose information of the aforementioned mobile device at the target time; or,

[0122] In the case of multiple target image pairs, the RANSAC algorithm is called to remove mismatched pairs, and the selected matching pairs are then optimized for pose using a nonlinear optimization algorithm (such as the Levenberg-Marquardt algorithm) to obtain the calibration pose information of the mobile device at the target time.

[0123] In embodiments including steps S110 to S140, feature fusion is achieved by utilizing the geometric relationships between point cloud data with fused depth information and the two-dimensional visual feature descriptors of each point cloud. This enables multi-dimensional fusion of three-dimensional geometric features and two-dimensional visual feature descriptors. Multi-dimensional feature description of the point cloud is performed simultaneously based on the three-dimensional geometric features incorporating depth information and the two-dimensional visual feature descriptors. This allows for location identification during subsequent matching with a keyframe database or a global point cloud map database. Matching is performed based on the geometric-visual joint features corresponding to the geometric shapes, which overcomes the limitations of matching solely based on geometric features in similar scenarios such as white walls and corridors. This method overcomes the indistinguishability of scenes; simultaneously, because the matching of 2D visual feature descriptors also incorporates depth information corresponding to 3D geometric features, redundancy is reduced during the matching of numerous point cloud locations, improving matching efficiency and accuracy, and avoiding drift errors easily caused by scale uncertainty in pure 2D-to-2D geometric constraints. Overall, it improves the processing efficiency and accuracy of real-time localization and pose estimation. This method can not only be applied to loop closure detection in visual SLAM systems, but also perform map reuse and relocalization based on pre-generated 3D point cloud maps (supporting color point cloud maps), achieving high-precision localization and exhibiting strong illumination robustness. It can still perform vision-LiDAR-based position recognition and loop closure detection in structurally similar scenes. Furthermore, maps built based on this joint feature are compatible with conventional STD (geometric approach), visual observation, and other position recognition schemes, achieving more efficient map utilization.

[0124] Figure 4 The diagram illustrates a comparison between the position recognition and loop closure detection results obtained by processing a real-world radar-visual dataset using (a) the relocation method provided in this embodiment and (b) the position recognition and loop closure detection results obtained using only the STD triangle descriptor.

[0125] contrast Figure 4 As shown in (a) and (b), relocalization processing based on the radar-visual dataset collected by the LIVO system running at the front end, using only the STD triangle descriptor (corresponding to the geometric method) for position recognition and pose correction results in a large number of position recognition attempts and numerous errors, leading to poor and divergent positioning results. In contrast, the method provided in this embodiment for position recognition and loop closure detection can improve the accuracy of position recognition, achieve high-precision positioning, and has strong illumination robustness.

[0126] A second exemplary embodiment of this disclosure provides a relocation apparatus for cross-modal feature fusion.

[0127] Figure 5 A structural block diagram of a relocation device for cross-modal feature fusion according to an embodiment of the present disclosure is shown schematically.

[0128] Reference Figure 5 As shown in the embodiments of this disclosure, the cross-modal feature fusion relocalization device 500 includes: a point cloud data generation module 510, a feature fusion module 520, a comprehensive matching module 530, and a pose calculation module 540.

[0129] The aforementioned relocation device is integrated into the mobile device or is an independent entity that can communicate with the mobile device.

[0130] The point cloud data generation module 510 described above is used to preprocess the perception data from the data sensor of the mobile device to obtain point cloud data with fused depth information.

[0131] The aforementioned feature fusion module 520 is used to perform feature fusion with the two-dimensional visual feature descriptor of each point cloud based on the geometric relationship between the point cloud data, to obtain the geometric-visual joint features corresponding to the geometric figure; the aforementioned geometric figure is composed of multiple point clouds.

[0132] The aforementioned integrated matching module 530 is used to perform integrated matching of the aforementioned geometric-visual joint features in a real-time constructed keyframe database or a pre-constructed global point cloud map database.

[0133] The aforementioned pose calculation module 540 is used to perform pose calculation based on the target graphic pair formed by the successfully matched target geometry and the target matching graphic to obtain the calibration pose information of the aforementioned mobile device.

[0134] For more details of this embodiment, please refer to the relevant description of the first embodiment, which will not be repeated here.

[0135] In the aforementioned device, multi-dimensional feature description of point clouds is performed simultaneously using 3D geometric features incorporating depth information and 2D visual feature descriptors. This allows for location identification during subsequent matching with a keyframe database or global point cloud map database. Matching based on the geometric-visual joint features corresponding to the geometric shapes overcomes the inability to distinguish similar scenes such as white walls and corridors when matching solely based on geometric features. Furthermore, since the matching of 2D visual feature descriptors also incorporates depth information corresponding to 3D geometric features, redundancy is reduced during matching of numerous point cloud locations, improving matching efficiency and accuracy. This avoids drift errors caused by the scale uncertainty of pure 2D-2D geometric constraints. Overall, this method improves the processing efficiency and accuracy of real-time localization and pose estimation. It can be applied not only to loop closure detection in visual SLAM systems but also to map reuse and relocalization based on pre-generated 3D point cloud maps (supporting color point cloud maps), achieving high-precision localization with strong illumination robustness.

[0136] A third exemplary embodiment of this disclosure provides a mobile device. The mobile device integrates the repositioning device provided in the second embodiment; or, the mobile device includes: a first point cloud data generation module, a first feature fusion module, a first comprehensive matching module, and a first pose calculation module.

[0137] The aforementioned first point cloud data generation module is used to preprocess the sensor data of the mobile device to obtain point cloud data with fused depth information.

[0138] The first feature fusion module is used to perform feature fusion with the two-dimensional visual feature descriptor of each point cloud based on the geometric relationship between the point cloud data to obtain the geometric-visual joint features corresponding to the geometric figure; the geometric figure is composed of multiple point clouds.

[0139] The aforementioned first comprehensive matching module is used to perform comprehensive matching of the aforementioned geometric-visual joint features in a real-time constructed keyframe database or a pre-constructed global point cloud map database.

[0140] The first pose calculation module is used to calculate the pose of the target graphic pair formed by the successfully matched target geometry and the target matching graphic to obtain the calibration pose information of the mobile device.

[0141] In some embodiments, the mobile device is at least one of the following: intelligent robot (e.g., search and rescue robot, transport robot, sweeping robot, etc.), vehicle (e.g., vehicle supporting autonomous driving or assisted driving), flight equipment (e.g., drone, manned aircraft, etc.), wearable device.

[0142] For more details of this embodiment, please refer to the relevant description of the first embodiment, which will not be repeated here.

[0143] A fourth exemplary embodiment of this disclosure provides a service device.

[0144] The aforementioned service device is used to provide service support for the positioning of mobile devices. The aforementioned service device integrates the repositioning device provided in the second embodiment above; or, the aforementioned service device includes: a second point cloud data generation module, a second feature fusion module, a second comprehensive matching module, and a second pose calculation module.

[0145] The second point cloud data generation module is used to preprocess the sensor data of the mobile device to obtain point cloud data with fused depth information.

[0146] The second feature fusion module is used to perform feature fusion with the two-dimensional visual feature descriptor of each point cloud based on the geometric relationship between the point cloud data to obtain the geometric-visual joint features corresponding to the geometric figure; the geometric figure is composed of multiple point clouds.

[0147] The aforementioned second comprehensive matching module is used to perform comprehensive matching of the above-mentioned geometric-visual joint features in a real-time constructed keyframe database or a pre-constructed global point cloud map database.

[0148] The second pose calculation module is used to calculate the pose of the target graphic pair formed by the successfully matched target geometry and the target matching graphic to obtain the calibration pose information of the mobile device.

[0149] For more details of this embodiment, please refer to the relevant description of the first embodiment, which will not be repeated here.

[0150] Any number of the functional modules included in the aforementioned relocation device, mobile device, and service device can be combined into one module, or any one of these modules can be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules can be combined with at least part of the functionality of other modules and implemented in one module. At least one of the functional modules included in the aforementioned relocation device, mobile device, and service device can be at least partially implemented as hardware circuitry, such as Field Programmable Gate Array (FPGA), Programmable Logic Array (PLA), System-on-Chip, System-on-Substrate, System-on-Package, Application-Specific Integrated Circuit (ASIC), or implemented in hardware or firmware by any other reasonable means of integrating or packaging circuitry, or implemented in software, hardware, and firmware, or in any appropriate combination of any of these three implementation methods. Alternatively, at least one of the functional modules included in the aforementioned relocation device, mobile device, and service device can be at least partially implemented as a computer program module, which can perform corresponding functions when the computer program module is run.

[0151] The fifth exemplary embodiment of this disclosure provides an electronic device.

[0152] Figure 6 A schematic block diagram of an electronic device provided in an embodiment of the present disclosure is shown.

[0153] Reference Figure 6 As shown, the electronic device 600 provided in this embodiment includes a processor 601, a communication interface 602, a memory 603, and a communication bus 604. The processor 601, the communication interface 602, and the memory 603 communicate with each other through the communication bus 604. The memory 603 is used to store computer programs. When the processor 601 executes the program stored in the memory, it implements the cross-modal feature fusion relocation method as described above.

[0154] A sixth exemplary embodiment of this disclosure also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program that, when executed by a processor, implements the cross-modal feature fusion relocation method as described above.

[0155] The computer-readable storage medium may be included in the device or apparatus described in the above embodiments; or it may exist independently and not assembled into the device or apparatus. The computer-readable storage medium carries one or more programs that, when executed, implement the method according to the embodiments of this disclosure.

[0156] According to embodiments of this disclosure, the computer-readable storage medium can be a non-volatile computer-readable storage medium, such as including, but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0157] It should be noted that the collection, gathering, updating, analysis, processing, use, transmission, and storage of user personal information involved in the technical solutions provided in this disclosure comply with the provisions of relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. Necessary measures are taken to prevent unauthorized access to user personal information data and to safeguard user personal information security, network security, and national security.

[0158] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0159] The above description is merely a specific embodiment of this disclosure, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A cross-modal feature fusion-based relocalization method, characterized in that, The method comprises: preprocessing perception data of a data sensor of a mobile device to obtain point cloud data fused with depth information; the point cloud data is feature point cloud data fused with depth information under a camera imaging plane perspective of the mobile device; performing feature fusion according to geometric relationships between point cloud data and two-dimensional visual feature descriptors of each point cloud to obtain geometric-visual joint features corresponding to geometric patterns; the geometric patterns are composed of multiple point clouds; performing comprehensive matching of the geometric-visual joint features in a key frame database constructed in real time or a global point cloud map database constructed in advance; performing pose solving according to a target pattern pair composed of a target geometric pattern and a target matching pattern obtained through the comprehensive matching to obtain calibration pose information of the mobile device; wherein the feature fusion according to geometric relationships between point cloud data and two-dimensional visual feature descriptors of each point cloud to obtain geometric-visual joint features corresponding to geometric patterns comprises: performing position matching and association according to the feature point cloud data fused with depth information and two-dimensional visual feature descriptors extracted for image pixels under the camera imaging plane perspective; selecting point clouds to construct geometric patterns for feature fusion according to results of the position matching and association to obtain geometric-visual joint features corresponding to the geometric patterns, including: for first point cloud objects corresponding to position matching of the feature point cloud data fused with depth information and the two-dimensional visual feature descriptors, selecting the first point cloud objects to construct geometric patterns; for second point cloud objects corresponding to position mismatching of the feature point cloud data fused with depth information and the two-dimensional visual feature descriptors, performing position near-neighbor association according to first position distributions of the feature point cloud data and second position distributions of corresponding point clouds of the two-dimensional visual feature descriptors, and constructing geometric patterns based on results of the position near-neighbor association; performing feature splicing on the constructed geometric patterns in at least the following dimensions: point cloud position, edge length, and two-dimensional visual feature descriptors corresponding to the point clouds, to obtain geometric-visual joint features corresponding to the geometric patterns.

2. The relocation method of claim 1, wherein, The feature point cloud data is generated by at least one of visual perception data and radar perception data.

3. The relocation method of claim 1, wherein, The position near-neighbor association according to first position distributions of the feature point cloud data and second position distributions of corresponding point clouds of the two-dimensional visual feature descriptors, and the construction of geometric patterns based on results of the position near-neighbor association, comprise: when the first position distribution of a first target point cloud and the second position distribution of a second target point cloud in the second point cloud objects are within a preset near-neighbor range, marking point clouds within the preset near-neighbor range in the first target point cloud and the second target point cloud as first associated point clouds, and the multi-dimensional features of the first associated point clouds are shared within the preset near-neighbor range; selecting the first associated point clouds to construct geometric patterns in the preset near-neighbor range; or, When the first position distribution of the first target point cloud in the second point cloud object has a position coincidence part with the second position distribution of the second target point cloud, a preset neighborhood is expanded based on the position coincidence part in the first target point cloud and the second target point cloud, and a point cloud in the preset neighborhood is marked as a second associated point cloud, and the multi-dimensional features of the second associated point cloud are shared in the preset neighborhood; and a geometric figure is constructed by selecting the second associated point cloud in the preset neighborhood.

4. The relocation method according to any one of claims 1-3, characterized by, The method for constructing the geometric figure from the selected point cloud comprises: a set of candidate geometric figures is formed by randomly selecting a preset number of point cloud data; in the set, the edges between the point clouds in each candidate geometric figure are sorted according to the edge size; whether any edge in each candidate geometric figure exceeds a set threshold is determined in sequence according to the edge sorting; the candidate geometric figure with the edge exceeding the set threshold is removed, and the remaining candidate geometric figure in the set is constructed as a geometric figure to be matched.

5. The relocation method of claim 1, wherein, The comprehensive matching comprises: a first matching for the geometric relationship and a second matching for the two-dimensional visual feature descriptor, and the first matching and the second matching are performed in a preset order.

6. The relocation method according to any one of claims 1-3, wherein, The geometric figure is a triangle formed by randomly selecting three point clouds; the geometric-visual joint feature corresponding to the geometric figure comprises: the positions of the three point cloud vertices in the triangle, the lengths of the three edges, and the two-dimensional visual feature descriptors corresponding to the three point cloud vertices; in the key frame database or the global point cloud map database, the geometric-visual joint features of each reference triangle are stored based on the storage hash index obtained by processing the diagonal value of the longest side; In the real-time constructed key frame database or the pre-constructed global point cloud map database, the comprehensive matching of the geometric-visual joint feature comprises: For each triangle to be matched, the following comprehensive matching processing operations are performed: processing the diagonal value of the longest side of the current triangle to be matched to obtain a matching hash index; matching the matching hash index with the storage hash index to locate a target storage interval; in the target storage interval, angle matching is performed based on the diagonal value of the longest side of the current triangle to be matched; for the first matching reference triangle set obtained by successful angle matching, length matching is sequentially performed based on the longest side and the second longest side of the current triangle to be matched; for the second matching reference triangle set obtained by successful length matching, feature matching is performed based on the two-dimensional visual feature descriptors of each point cloud vertex of the triangle to be matched; a target figure pair is generated based on the target matching reference triangle obtained by successful feature matching and the current triangle.

7. The relocation method according to any one of claims 1-3, wherein, The calibration pose information of the mobile device is obtained by solving the pose based on the target figure pair formed by the target geometric figure and the target matching figure obtained by successful comprehensive matching, comprising: a PNP algorithm is called to solve the pose based on the correspondence between the target geometric figure and the target matching figure in the target figure pair to obtain the calibration pose information of the mobile device at the target time; or In the presence of multiple sets of target pattern pairs, a RANSAC algorithm is called to eliminate false matching pairs, and a pose optimization algorithm is used to solve the matching pairs after screening to obtain the calibration pose information of the mobile device at the target time.

8. The relocation method according to any one of claims 1-3, wherein, The perception data includes at least one of visual perception data and radar perception data, and the visual perception data includes at least one of image data captured by a depth camera and image data captured by a multi-view camera. The pre-processing of the perception data of the data sensor to obtain the point cloud data fused with depth information includes at least one of: pixel feature point extraction and feature description of the image data captured by the depth camera, and projection of the pixels to a three-dimensional space according to the corresponding depth values to obtain the feature point cloud data fused with depth information; or, pixel feature point extraction and feature description of the image data captured by the multi-view camera, and depth generation based on the multi-view stereo vision and projection to a three-dimensional space to obtain the feature point cloud data fused with depth information; or, projection of the point cloud data in the radar perception data to a three-dimensional feature point cloud, and feature description of the three-dimensional feature point cloud projected back to the imaging plane corresponding to the camera to obtain the feature point cloud data fused with depth information.

9. A repositioning device for cross-modal feature fusion, comprising: The repositioning device is integrated on the mobile device or is an independent body capable of communicating with the mobile device, and the repositioning device includes: a point cloud data generation module configured to pre-process the perception data of the data sensor of the mobile device to obtain point cloud data fused with depth information; the point cloud data is feature point cloud data fused with depth information in the perspective of the imaging plane of the camera of the mobile device; a feature fusion module configured to perform feature fusion according to the geometric relationship between the point cloud data and the two-dimensional visual feature descriptor of each point cloud to obtain geometric-visual joint features corresponding to geometric patterns; the geometric patterns are composed of multiple point clouds; a comprehensive matching module configured to perform comprehensive matching of the geometric-visual joint features in a key frame database constructed in real time or a global point cloud map database constructed in advance; a pose solving module configured to perform pose solving according to a target pattern pair composed of a target geometric pattern and a target matching pattern to obtain calibration pose information of the mobile device; wherein the feature fusion according to the geometric relationship between the point cloud data and the two-dimensional visual feature descriptor of each point cloud to obtain geometric-visual joint features corresponding to geometric patterns includes: position matching and association according to the feature point cloud data fused with depth information and the two-dimensional visual feature descriptor extracted for image pixels in the perspective of the imaging plane of the camera; selection of point clouds to construct geometric patterns for feature fusion according to the results of the position matching and association to obtain geometric-visual joint features corresponding to geometric patterns, including: selection of a first point cloud object corresponding to the position matching of the feature point cloud data fused with depth information and the two-dimensional visual feature descriptor to construct a geometric pattern. According to the first position distribution of the feature point cloud data and the second position distribution of the two-dimensional visual feature descriptor corresponding point cloud, position near neighbor association is performed, and a geometric figure is constructed based on the result after the position near neighbor association; The constructed geometric figure is at least spliced in the following dimensions: point cloud position, edge length, and two-dimensional visual feature descriptor corresponding to the contained point cloud, to obtain a geometric-visual joint feature corresponding to the geometric figure.

10. A mobile device, comprising: The mobile device is integrated with the repositioning device of claim 9; Or, The mobile device comprises a first point cloud data generation module, a first feature fusion module, a first comprehensive matching module, and a first pose solving module; The first point cloud data generation module is configured to preprocess the perception data of the data sensor of the mobile device to obtain point cloud data fused with depth information; the point cloud data is feature point cloud data fused with depth information under the camera imaging plane perspective of the mobile device; The first feature fusion module is configured to perform feature fusion on the basis of the geometric relationship between the point cloud data and the two-dimensional visual feature descriptor of each point cloud to obtain a geometric-visual joint feature corresponding to the geometric figure; the geometric figure is composed of multiple point clouds; The first comprehensive matching module is configured to perform comprehensive matching of the geometric-visual joint feature in a real-time constructed key frame database or a pre-constructed global point cloud map database; The first pose solving module is configured to perform pose solving on the basis of a target figure pair composed of a target geometric figure and a target matching figure to obtain calibration pose information of the mobile device; According to the first position distribution of the feature point cloud data and the second position distribution of the two-dimensional visual feature descriptor corresponding point cloud, position near neighbor association is performed, and a geometric figure is constructed based on the result after the position near neighbor association; The constructed geometric figure is at least spliced in the following dimensions: point cloud position, edge length, and two-dimensional visual feature descriptor corresponding to the contained point cloud, to obtain a geometric-visual joint feature corresponding to the geometric figure. The mobile device is at least one of the following: a smart robot, a vehicle, a flight device, and a wearable device. The mobile device is at least one of the following: a smart robot, a vehicle, a flight device, and a wearable device. ​ 11. The mobile device of claim 10, wherein, ​ 12. A service device, characterized by A service device integrated with the repositioning device of claim 9 is used to provide service support for positioning of a mobile device. Or, The service device comprises a second point cloud data generation module, a second feature fusion module, a second comprehensive matching module and a second pose solution module. The second point cloud data generation module is configured to preprocess sensing data of a data sensor of the mobile device to obtain point cloud data fused with depth information; the point cloud data is feature point cloud data fused with depth information under a camera imaging plane perspective of the mobile device. The second feature fusion module is configured to perform feature fusion on a geometric relationship between point cloud data and a two-dimensional visual feature descriptor of each point cloud to obtain geometric-visual joint features corresponding to geometric patterns; the geometric patterns are composed of multiple point clouds. The second comprehensive matching module is configured to perform comprehensive matching of the geometric-visual joint features in a key frame database constructed in real time or a global point cloud map database constructed in advance. The second pose solution module is configured to perform pose solution according to a target pattern pair composed of a target geometric pattern and a target matching pattern of which the comprehensive matching is successful, to obtain calibration pose information of the mobile device. The feature fusion on the geometric relationship between the point cloud data and the two-dimensional visual feature descriptor of each point cloud to obtain the geometric-visual joint features corresponding to the geometric patterns comprises: performing position matching and association on the feature point cloud data fused with depth information and the two-dimensional visual feature descriptor extracted for image pixels under the camera imaging plane perspective; selecting point clouds to construct geometric patterns for feature fusion according to a result of the position matching and association, to obtain the geometric-visual joint features corresponding to the geometric patterns, including: selecting a first point cloud object corresponding to position matching of the feature point cloud data fused with depth information and the two-dimensional visual feature descriptor to construct a geometric pattern; for a second point cloud object corresponding to position mismatching of the feature point cloud data fused with depth information and the two-dimensional visual feature descriptor, performing position neighbor association according to a first position distribution of the feature point cloud data and a second position distribution of a corresponding point cloud of the two-dimensional visual feature descriptor, and constructing a geometric pattern based on a result of the position neighbor association; performing feature splicing on the constructed geometric pattern in at least the following dimensions: point cloud position, edge length and the two-dimensional visual feature descriptor corresponding to the point cloud contained in the geometric pattern, to obtain the geometric-visual joint features corresponding to the geometric pattern.

13. An electronic device, comprising: The computer program is executed by the processor to implement the method of any one of claims 1-8. The computer program is executed by the processor to implement the method of any one of claims 1-8. ​ 14. A computer-readable storage medium having stored thereon a computer program, characterized in that, ​

Citation Information

Patent Citations

  • Semantic mapping and positioning method based on priori laser point cloud and depth map fusion

    CN112258618A

  • Semi-direct vision positioning method fusing point and line features

    CN115965686A