Multi-modal tight coupling SLAM method based on gradient descent optimization

By employing a gradient descent-optimized multimodal tightly coupled SLAM method, multimodal feature constraints are constructed using multi-sensor data and reflective markers. This addresses the challenge of adjusting driving strategies for intelligent vehicles in complex environments, enabling high-precision map modeling and stable vehicle pose estimation.

CN120877054AActive Publication Date: 2025-10-31NANJING YULING TECH CO LTD

Patent Information

Application Number
CN202511361174.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-23
Publication Date
2025-10-31
Estimated Expiration
2045-09-23

AI Technical Summary

Technical Problem

In complex driving scenarios, intelligent vehicles struggle to promptly identify obstacles such as walls and pillars in their surroundings, making it difficult to adjust driving strategies. In particular, SLAM methods are difficult to converge stably in narrow spaces and with metallic interference.

Method used

A multimodal tightly coupled SLAM method based on gradient descent optimization is adopted. Environmental feature data are collected through a camera, laser ranging unit and inertial measurement unit, reflective markers are detected to determine the position, a time-synchronized multimodal observation set is constructed, multimodal features are extracted and fused, an initial pose graph with multimodal feature constraints and marker anchor point constraints is constructed, and the pose graph is optimized through an error function and global consistency is verified.

Benefits of technology

It effectively avoids the lack of observation caused by narrow spaces and occlusion, enhances observability, reduces the risk of local optima, and ensures high-precision map modeling and vehicle pose estimation in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120877054A_ABST
    Figure CN120877054A_ABST
Patent Text Reader

Abstract

A multi-modal tight coupling SLAM method based on gradient descent optimization comprises the following steps: acquiring environment characteristic data in a movement process of intelligent mobile equipment in real time, and determining a spatial position of the intelligent mobile equipment by detecting a reflective mark so as to construct a time-synchronized multi-modal observation set. And identifying and extracting multi-modal features from the multi-modal observation set, so as to carry out feature fusion on the multi-modal features to obtain a united-coded joint feature set. And based on the joint feature set, constructing an initial pose map of the multi-modal feature constraint and the marked anchor point constraint. And calling an error function to model the multi-modal observation residual error and the marking residual error, and optimizing the initial pose image by iteratively adjusting the pose and the error of modeling estimation. And performing global consistency verification on the optimized initial pose map, and outputting an environment modeling result and the pose state of the intelligent mobile device, thereby solving the problem of difficult SLAM convergence caused by narrow space and metal interference in a complex scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence technology, and more specifically, relates to a multimodal tightly coupled SLAM method based on gradient descent optimization. Background Technology

[0002] With the development of artificial intelligence, autonomous driving technology for intelligent vehicles has received widespread attention. Intelligent driving decision-making and planning is one of the key technologies for realizing autonomous driving of intelligent vehicles. Intelligent driving decision-making and planning refers to the vehicle's planning of driving decisions based on perceived environmental information and pre-set goals, through decision-making algorithms and models. This includes path planning, behavioral decisions, and compliance with traffic rules.

[0003] However, in complex scenarios such as underground parking garages, if intelligent vehicles struggle to promptly identify environmental factors like walls and pillars that could obstruct or damage their movement, they may find it difficult to adjust their driving strategies in a timely manner. Therefore, in complex driving scenarios, environmental observation requires multimodal data support to help intelligent vehicles modify their driving strategies in real time through multimodal data analysis. This helps avoid observation gaps caused by narrow spaces or obstructions, thereby ensuring the safety of intelligent driving. Summary of the Invention

[0004] To address the shortcomings of existing technologies, the present invention aims to overcome the aforementioned deficiencies and propose a multimodal tightly coupled SLAM method based on gradient descent optimization.

[0005] The present invention adopts the following technical solution.

[0006] The first aspect of this invention discloses a multimodal tightly coupled SLAM method based on gradient descent optimization, the method comprising: Real-time acquisition of environmental feature data during the movement of intelligent mobile devices, and determination of the spatial position of the intelligent mobile devices by detecting reflective marks, in order to construct a time-synchronized multimodal observation set; Multimodal features are identified and extracted from the multimodal observation set, and feature fusion is performed on the multimodal features to obtain a unified encoded joint feature set; Based on the joint feature set, an initial pose graph with multimodal feature constraints and marker anchor point constraints is constructed; The error function is called to model the multimodal observation residuals and the label residuals, and the initial pose map is optimized by iteratively adjusting the pose and the modeling estimation error; Perform a global consistency check on the optimized initial pose graph, and output the environmental modeling results and the pose state of the smart mobile device after the check passes. The environmental feature data includes environmental image data and laser echoes. The multimodal features are corners and edges identified from the environmental image data, geometric features of environmental entities extracted from the laser echoes, and reflective marks identified from the high points of reflective marks.

[0007] Furthermore, the real-time acquisition of environmental feature data during the movement of the intelligent mobile device, and the determination of the spatial position of the intelligent mobile device by detecting reflective markers, to construct a time-synchronized multimodal observation set, includes: Control the camera to acquire environmental image data during the movement of the smart mobile device at a set frame rate; The laser ranging unit rotates and emits a laser at a fixed scanning frequency and receives the reflected signal to calculate the distance and angle between the smart mobile device and environmental entities during its movement, thereby generating a three-dimensional point set of the laser echo. The acceleration and angular velocity output by the intelligent mobile device body are obtained through the inertial measurement unit to calculate the inertial data of the intelligent mobile device; A time series is established according to a unified time step, and linear interpolation is performed on the environmental image data, 3D point set, and inertial data based on the time series. At the same time, the 3D point set is transformed into the coordinate system of the intelligent mobile device according to the calibrated extrinsic parameter matrix.

[0008] Furthermore, the real-time acquisition of environmental feature data during the movement of the intelligent mobile device, and the determination of the spatial position of the intelligent mobile device by detecting reflective markers to construct a time-synchronized multimodal observation set, also includes: Threshold segmentation is performed on the environmental image data to extract the set of pixels whose brightness exceeds a first threshold, and the weighted centroid of the region of the pixel set is calculated. The weighted centroid is converted into a ray direction vector, and the ray direction vector is normalized. At the same time, candidate reflection points with a reflection intensity not lower than the second threshold are selected from the three-dimensional point set. When the angle between the candidate reflection point and the ray direction vector is within a set threshold range, the candidate reflection point is marked for reflection, and the three-dimensional position coordinates of the reflection mark are obtained.

[0009] Furthermore, the step of identifying and extracting multimodal features from the multimodal observation set, and then fusing these multimodal features to obtain a unified encoded joint feature set, includes: Based on the reflection marker observation results in the multimodal observation set, the number of times the reflection marker appears in consecutive frames and the spatial consistency of the reflection marker between the three-dimensional point set and the environmental image data are verified in order to perform preliminary filtering of the reflection marker observation results. Edge points and corner points of environmental entities are extracted from the three-dimensional point set, and the coordinates of the reflection markers in the preliminary filtered reflection marker observation results are aligned with the observation timestamps of the edge points and corner points of the environmental entities to construct the unified coding joint feature set.

[0010] Furthermore, the step of constructing an initial pose graph with multimodal feature constraints and marker anchor point constraints based on the joint feature set includes: The motion state of the intelligent mobile device at the first moment is estimated based on the inertial data in the joint feature set, and the estimated position of the intelligent mobile device at the first moment is associated with the environmental image data and the three-dimensional point set at the first moment. The angular velocity of the smart mobile device at the first moment is converted into the device attitude, and the device attitude is bound with the associated environmental image data and 3D point set to determine the pose node of each frame of observation data, thus obtaining the pose node set.

[0011] Furthermore, the step of constructing an initial pose graph with multimodal feature constraints and marker anchor point constraints based on the joint feature set further includes: The matching pixels in adjacent environmental image frames are determined by feature point matching, and the corresponding points and pose nodes in the 3D point set are aligned and projected onto the environmental image coordinate system to calculate the visual residual between the 3D point set and pose nodes and the matching pixels, thereby obtaining the visual constraint edge. The geometric residuals of the wall corner points and column edge points in the joint feature set are defined to determine the geometric constraint edges. The inertial data and time step are used to determine the comparison residuals between the predicted pose and the pose nodes to obtain the inertial constraint edges.

[0012] Furthermore, the step of constructing an initial pose graph with multimodal feature constraints and marker anchor point constraints based on the joint feature set further includes: Each reflection marker point in the preliminary filtered reflection marker observation results is represented by three-dimensional coordinates, and the three-dimensional coordinate representation of the reflection marker point is transformed from the local coordinate system to the global coordinate system. The anchor point residual is calculated based on the corresponding reflection marker point and anchor point reference in the global coordinate system to generate reflection marker constraints. The visual constraint edges, geometric constraint edges, inertial constraint edges, and reflection mark constraints are merged and combined with the pose node set to generate the initial pose graph.

[0013] Furthermore, the step of calling the error function to model the multimodal observation residuals and the label residuals, and optimizing the initial pose map by iteratively adjusting the pose and modeling estimation errors, includes: Based on the visual residual, geometric residual, inertial contrast residual, and anchor point residual in the initial pose graph, an overall error function is constructed, and the error partial derivative of each node is calculated according to the overall error function to obtain the gradient vector of each node. Based on the gradient vector, the poses of each node in the initial pose graph are updated and normalized according to the calibrated learning rate, and the error change of the updated node poses is calculated. When the error change is lower than the convergence threshold, the optimized pose graph is output.

[0014] A second aspect of this invention discloses a multimodal tightly coupled SLAM device based on gradient descent optimization, used to implement the multimodal tightly coupled SLAM method based on gradient descent optimization as described in any of the first aspects, the device comprising: The data acquisition module is used to collect environmental feature data of the smart mobile device in real time during its movement, and to determine the spatial position of the smart mobile device by detecting reflective marks, so as to construct a time-synchronized multimodal observation set; The feature set construction module is used to identify and extract multimodal features from the multimodal observation set, and to perform feature fusion on the multimodal features to obtain a unified encoded joint feature set; The pose graph construction module is used to construct an initial pose graph with multimodal feature constraints and marker anchor point constraints based on the joint feature set. The pose graph optimization module is used to call an error function to model the multimodal observation residuals and label residuals, and to optimize the initial pose graph by iteratively adjusting the pose and modeling estimation errors. The pose graph verification module is used to perform global consistency verification on the optimized initial pose graph, and outputs the environmental modeling results and the pose state of the smart mobile device after the verification is passed. The environmental feature data includes environmental image data and laser echoes. The multimodal features are corners and edges identified from the environmental image data, geometric features of environmental entities extracted from the laser echoes, and reflective marks identified from the high points of reflective marks.

[0015] A third aspect of the present invention discloses a terminal, including a processor and a storage medium; The storage medium is used to store instructions; The processor is configured to operate according to the instructions to perform the steps of the method described in the first aspect.

[0016] A fourth aspect of the present invention discloses a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described in the first aspect.

[0017] The beneficial effects of the present invention are as follows: Compared with the prior art, the present invention has the following advantages: (1) This invention collects environmental feature data by using a camera, laser ranging unit and inertial measurement unit during the movement of the vehicle robot. At the same time, it observes the spatial position of the reflective mark by detecting the bright echo or bright spot. All the collected results are unified to the same time reference to ensure the consistency of the observation data. This provides complete original information for the subsequent establishment of unified constraints and effectively avoids the situation where the observation is missing and cannot be aligned due to narrow space and occlusion.

[0018] (2) This invention extracts and fuses features from a multimodal observation set to identify corners and edges in environmental images, extracts geometric features of walls and cylinders from laser echoes, and identifies reflection markers and their unique identifiers from synchronized high-brightness points. All features are then uniformly encoded to form a joint feature set. By increasing the recognition results of reflection markers, the observability in narrow and metallic environments is enhanced. Furthermore, based on the joint feature set, observation edges between the identified reflection markers are added to each frame pose, resulting in an initial pose map containing multimodal constraints and marker anchor point constraints. This solves the problem of uneven distribution of single environmental features and supplements the geometric convergence direction through anchor point observation.

[0019] (3) This invention constructs a unified error function to model the multimodal observation residuals and label residuals in a unified manner. During gradient descent, it utilizes the strong geometric convergence direction provided by the label constraint to avoid optimization getting trapped in local optima. It iteratively adjusts the pose and map estimation until the error converges, significantly reducing the risk of local optima in narrow environments and under metal reflection, and improving the stability of error convergence. Finally, a global consistency check is performed on the convergence optimization results to ensure that the map has no distortion in its overall structure and that the pose estimation is stable and reliable. If the check passes, the final high-precision parking lot environment map and vehicle pose results are output, solving the problem of SLAM convergence difficulties caused by narrow space and metal interference in complex scenarios such as underground parking lots. Attached Figure Description

[0020] Figure 1 This is a flowchart illustrating a multimodal tightly coupled SLAM method based on gradient descent optimization provided by the present invention.

[0021] Figure 2 This is a schematic diagram of a multimodal tightly coupled SLAM device based on gradient descent optimization provided by the present invention. Detailed Implementation

[0022] The present application will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention, and should not be construed as limiting the scope of protection of the present application.

[0023] like Figure 1 As shown, in one embodiment, a multimodal tightly coupled SLAM method based on gradient descent optimization includes the following steps: Step S110: Real-time acquisition of environmental feature data during the movement of the smart mobile device, and determination of the spatial location of the smart mobile device by detecting reflective markers, in order to construct a time-synchronized multimodal observation set.

[0024] The environmental feature data includes environmental image data and laser echoes.

[0025] In some embodiments, the multimodal tightly coupled SLAM method based on gradient descent optimization provided by the present invention includes the following steps in step S110: Step S111: Control the camera to acquire environmental image data during the movement of the smart mobile device at a set frame rate.

[0026] Step S112: The laser ranging unit rotates and emits laser light at a fixed scanning frequency and receives reflected signals to calculate the distance and angle between the smart mobile device and environmental entities during its movement, thereby generating a three-dimensional point set of laser echoes.

[0027] Step S113: Obtain the acceleration and angular velocity output by the smart mobile device body through the inertial measurement unit to calculate the inertial data of the smart mobile device.

[0028] Step S114: Establish a time series according to a unified time step, and perform linear interpolation on the environmental image data, 3D point set and inertial data according to the time series. At the same time, transform the 3D point set to the coordinate system of the smart mobile device according to the calibrated extrinsic parameter matrix.

[0029] In some embodiments, the multimodal tightly coupled SLAM method based on gradient descent optimization provided by the present invention further includes the following steps in step S110: Step S115: Perform threshold segmentation in the environmental image data to extract the set of pixels in the environmental image data whose brightness exceeds the first threshold, and calculate the weighted centroid of the pixel set region.

[0030] Step S116: Convert the weighted centroid into a ray direction vector and normalize the ray direction vector. At the same time, select candidate reflection points with a reflection intensity not lower than the second threshold from the three-dimensional point set.

[0031] Step S117: When the angle between the candidate reflection point and the light direction vector is within a set threshold range, the candidate reflection point is marked for reflection, and the three-dimensional position coordinates of the reflection mark are obtained.

[0032] In a specific embodiment, the present invention provides a multimodal tightly coupled SLAM method based on gradient descent optimization, comprising steps 1 to 5: Step 1: Multimodal sensor data and reflectance marker observations are acquired simultaneously.

[0033] During the movement of the vehicle-mounted robot, environmental feature data is collected using cameras, laser ranging units, and inertial measurement units. Simultaneously, its spatial position is observed by detecting bright echoes or bright spots of reflective markers. All collected results are standardized to the same time reference to ensure data consistency, providing complete raw information for establishing unified constraints and avoiding misalignment due to observational gaps caused by confined spaces or occlusions. This includes the following sub-steps: Sub-step 1.1: Camera image frame acquisition.

[0034] Specifically, the camera continuously acquires scene images at a set frame rate (20-60Hz), with each frame accompanied by a timestamp, and outputs an image frame matrix of size H×M. Each element in the matrix represents the brightness of the corresponding pixel, with a value range of 0-255, which is used for subsequent synchronization with point cloud and IMU (Inertial Measurement Unit) data.

[0035] Sub-step 1.2: Laser ranging point cloud acquisition.

[0036] Specifically, the laser unit rotates and emits a beam at a fixed scanning frequency (5-20Hz), receives reflected signals, and calculates distance and angle to form a three-dimensional point set. It outputs a point cloud matrix of size N×4, where each row contains the reflection intensity and the corresponding three-dimensional coordinates of the point. Each frame of the point cloud is accompanied by a timestamp for subsequent time synchronization with the image and IMU.

[0037] Sub-step 1.3: Inertial measurement data acquisition.

[0038] Specifically, the IMU continuously outputs acceleration and angular velocity on the vehicle-mounted robot body, and records a precise timestamp for each data point, outputting an inertial data sequence containing three-axis acceleration and three-axis angular velocity, with a sampling frequency of 100-400Hz.

[0039] Sub-step 1.4: Time alignment and coordinate unification.

[0040] Specifically, based on the aforementioned output image frame matrix, point cloud matrix, and inertial data sequence, a time series is established with a uniform time step (0.005-0.02s). Linear interpolation is then performed on the images, point clouds, and inertial data according to the timestamps. Next, using the calibrated extrinsic parameter matrix, the point cloud is transformed from the LiDAR coordinate system to the vehicle coordinate system, i.e., the coordinate system of the onboard robot, to ensure consistency with the camera and IMU. Finally, the aligned multimodal observation set is output.

[0041] Sub-step 1.5: Reflectance marker observation and extraction.

[0042] Specifically, threshold segmentation is performed on the time-aligned environmental image to extract the set of pixels with brightness greater than a set threshold (150-240), and the weighted centroid of the region corresponding to the extracted pixel set is calculated. , The expression is: ; In the formula, For image frame matrix, These are pixel coordinates.

[0043] Then, the weighted centroids ( , The reflection intensity is converted into a ray direction vector and normalized. Candidate points with a reflection intensity not lower than a set threshold (0.6-0.9) are selected from the point cloud matrix. If the angle between a candidate point and the ray direction is within the set range, the candidate point is determined to be the corresponding reflection marker, and its three-dimensional coordinates are recorded. Finally, the reflection marker observation set is output.

[0044] It should be noted that the weighted centroid ( , The weighted centroid () represents the pixel coordinates on the image plane, which is two-dimensional information. It needs to be converted into a three-dimensional ray direction vector. This conversion requires utilizing the camera's intrinsic parameters (intersection and principal point coordinates). First, the weighted centroid () , Subtracting the principal point coordinates from the pixel coordinates and then dividing by the focal length yields the corresponding normalized coordinates. Since the normalized vector length is 1, the weighted centroid can be calculated. , The conversion to the direction vector of light rays.

[0045] In this embodiment, the three-dimensional coordinates of the candidate point can form a direction vector in the camera coordinate system. The angle between the candidate point and the light direction vector is calculated using the cosine formula, and then the angle setting range is determined based on this angle.

[0046] Step S120: Identify and extract multimodal features from the multimodal observation set, and perform feature fusion on the multimodal features to obtain a unified encoded joint feature set.

[0047] Among them, the multimodal features are corners and edges identified from environmental image data, geometric features of environmental entities extracted from laser echoes, and reflective marks identified from the high points of reflective marks.

[0048] In some embodiments, the multimodal tightly coupled SLAM method based on gradient descent optimization provided by the present invention includes the following steps in step S120: Step S121: Based on the reflection marker observation results in the multimodal observation set, verify the number of times the reflection marker appears in consecutive frames and the spatial consistency between the reflection marker and the three-dimensional point set and the environmental image data, so as to perform preliminary filtering of the reflection marker observation results.

[0049] Step S122: Extract edge points and corner points of environmental entities from the 3D point set, and align the coordinates of the reflection markers in the preliminary filtered reflection marker observation results with the observation timestamps of the edge points and corner points of the environmental entities to construct a unified coded joint feature set.

[0050] In a specific embodiment, this invention provides a multimodal tightly coupled SLAM method based on gradient descent optimization. Step 2 involves multimodal feature fusion and reflection marker identification. The multimodal observation set output from Step 1 is subjected to feature extraction and fusion. Corner points and edges are identified from the image, geometric features of walls and cylinders are extracted from the laser echo, and reflection markers and their unique identifiers are identified from synchronized high-brightness points. All features are ultimately uniformly encoded to form a joint feature set. By increasing the recognition results of reflection markers, observability in narrow and metallic interference environments is enhanced. This includes the following sub-steps: Sub-step 2.1, candidate label validity test.

[0051] Specifically, based on the aforementioned set of reflection marker observations, the number of times each reflection marker appears in consecutive frames is checked; if it appears only once, it is removed. Next, the spatial consistency between the reflection markers in the point cloud and the environmental image is checked; if the deviation is greater than a set threshold (0.1-0.3m), it is removed, resulting in a preliminarily filtered set of reflection markers.

[0052] Sub-step 2.2: Joint modeling of labels and environmental features.

[0053] Specifically, cylindrical edge points and wall corner points are extracted from the point cloud matrix, and combined with the three-dimensional coordinates in the reflection marker observation set, the cylindrical edge points, wall corner points and the three-dimensional coordinates of the reflection markers are aligned according to the timestamp to obtain the aligned joint feature set.

[0054] Sub-step 2.3 introduces motion constraints and generates standardized feature vectors.

[0055] Specifically, the coordinates of feature points in the joint feature set at the same time point are bound to the inertial data sequence, so that each feature has a spatial location and corresponding motion state, in order to output a multimodal standardized feature vector set.

[0056] Step S130: Based on the joint feature set, construct an initial pose graph of multimodal feature constraints and marker anchor point constraints.

[0057] In some embodiments, the multimodal tightly coupled SLAM method based on gradient descent optimization provided by the present invention includes the following steps in step S130: Step S131: Estimate the motion state of the smart mobile device at the first moment based on the inertial data in the joint feature set, and associate the estimated position of the smart mobile device at the first moment with the environmental image data and the three-dimensional point set at the first moment.

[0058] Step S132: The angular velocity of the smart mobile device at the first moment is converted into the device attitude, and the device attitude is bound with the associated environmental image data and 3D point set to determine the pose node of each frame of observation data, thus obtaining the pose node set.

[0059] In some embodiments, the multimodal tightly coupled SLAM method based on gradient descent optimization provided by the present invention further includes the following steps in step S130: Step S133: Matching pixels in adjacent environmental image frames are determined by feature point matching, and the corresponding points and pose nodes in the 3D point set are aligned and projected onto the environmental image coordinate system to calculate the visual residual between the 3D point set, pose nodes and matching pixels, thereby obtaining the visual constraint edge.

[0060] Step S134: Define the geometric residuals for the wall corner points and column edge points in the joint feature set to determine the geometric constraint edges, and use inertial data and time step to determine the comparison residuals between the predicted pose and the pose nodes to obtain the inertial constraint edges.

[0061] In some embodiments, the multimodal tightly coupled SLAM method based on gradient descent optimization provided by the present invention further includes the following steps in step S130: Step S135: Perform three-dimensional coordinate representation on each reflection marker point in the preliminary filtered reflection marker observation results, and transform the three-dimensional coordinate representation of the reflection marker point from the local coordinate system to the global coordinate system. Calculate the anchor point residual based on the corresponding reflection marker point and anchor point reference in the global coordinate system to generate reflection marker constraints.

[0062] Step S136: Merge the visual constraint edges, geometric constraint edges, inertial constraint edges and reflection mark constraints, and combine them with the pose node set to generate an initial pose graph.

[0063] In a specific embodiment, this invention provides a multimodal tightly coupled SLAM method based on gradient descent optimization. Step 3 involves the initial construction of the pose graph and the introduction of marker observation constraints. Based on the joint feature set output in step 2, a pose graph containing odometry, inertial measurement, vision, and laser geometric edges is constructed. Simultaneously, observation edges between the pose graph and the identified reflection markers are added to each frame of the pose graph. This process yields an initial pose graph containing multimodal constraints and marker anchor point constraints, solving the problem of uneven distribution of single environmental features and supplementing the geometric convergence direction through anchor point observations. The process includes the following sub-steps: Sub-step 3.1: Pose node initialization.

[0064] Specifically, based on the aforementioned multimodal joint feature set, using the inertial data as the primary reference, the vehicle's motion state at the corresponding time is estimated, and the obtained vehicle position estimate is mapped one-to-one with the environmental image and point cloud set at that time. Simultaneously, the angular velocity information is converted into vehicle attitude, and the vehicle attitude is bound to other observation data at the corresponding time. In this way, each frame of observation data has a corresponding pose node, ensuring that the multimodal features are subsequently associated with the map structure.

[0065] Sub-step 3.2: Establishment of multimodal observation constraint edges.

[0066] Specifically, firstly, matching pixels are found from adjacent image frames through feature point matching. Then, the corresponding points and node poses in the point cloud are aligned and projected into the image coordinate system. The visual residual between the projected coordinates and the matching pixels is calculated. Finally, visual constraint edges are generated based on this visual residual. The calculation expression is: ; In the formula, For adjacent image frames at time t and time t+1 and Matching pixels; for The corresponding 3D point in the point cloud set; Let be the quaternion pose at time t; express The resulting rotation matrix; This is the projection function.

[0067] Subsequently, for wall corner points in the environmental feature set, the signed distance from the point to the plane is defined as the wall geometric residual. Then, based on the wall geometric residuals, the wall geometric constraint edges are generated. The expression for calculating the wall geometric residuals is: ; In the formula, These are the 3D points in the point cloud set corresponding to the corner points of the wall. This indicates the position of the point in the global coordinate system; This is the wall normal vector; Used as a reference point on the wall.

[0068] Similarly, for cylindrical edge points in the environmental feature set, the difference between the radial distance from the point to the cylindrical axis and the cylindrical radius is defined as the cylindrical geometric residual, and then the cylindrical geometric constraint edge is generated from the cylindrical geometric residual.

[0069] In this embodiment, the predicted pose of the vehicle is determined using IMU data and time steps, and then compared with the node pose to obtain the inertial residual. Then, inertial constraint edges are generated based on the inertial data residuals. The expression for calculating the inertial residuals is as follows: ; In the formula, ( , () represents the predicted pose obtained through IMU pre-integration prediction; , () represents the node pose obtained in sub-step 3.1; This is quaternion multiplication; It is a quaternion logarithmic mapping used to transform the rotation difference into the minimum rotation vector.

[0070] Sub-step 3.3, Introduction of reflection marker anchor point constraints.

[0071] Specifically, each marker point in the aforementioned set of pre-filtered reflection markers is represented using three-dimensional coordinates, and its transformation from the local coordinate system to the global coordinate system is expressed as follows: ; In the formula, These are the coordinates of the reflection marker points detected in the local coordinate system; For node pose quaternions The resulting rotation matrix; Let be the three-dimensional position of the pose node at time t; This represents the position of the reflection marker in the global coordinate system.

[0072] It should be noted that the attitude quaternions are known. The expression is: ; In the formula, The real part represents the cosine component of the rotation; The imaginary part represents the component of rotation multiplied by the sine term; all units are dimensionless.

[0073] Then, the rotation matrix obtained from the attitude quaternion The expression is: ;

[0074] Next, the reflection marker points in the global coordinate system are compared with their corresponding anchor reference positions (the marker installation positions known beforehand during the parking lot setup or the baseline established during system initialization). The difference between the two is calculated to obtain the anchor residual. Based on the unique identifier of each reflection marker, if the identifier is repeatedly observed at different time points, anchor constraint edges are established between these nodes and the reference anchor position according to the anchor residual.

[0075] Sub-step 3.4: Initial pose graph generation.

[0076] Specifically, the generated multimodal constraint set (including visual constraint edges, wall geometry constraint edges, cylindrical geometry constraint edges, and inertial constraint edges) is joined with the reflection marker constraint set (anchor point constraint edges). The pose node set is then combined with this union to generate the initial pose graph. Finally, a consistency check is performed on the generated initial pose graph: ensuring that each node is connected by at least one constraint edge; that reflection marker points under the same identifier are associated with only one reference anchor point; and that isolated nodes or observation data that do not meet the set geometric thresholds are deleted.

[0077] Step S140: The error function is called to model the multimodal observation residuals and the label residuals, and the initial pose map is optimized by iteratively adjusting the pose and the modeling estimation error.

[0078] In some embodiments, the multimodal tightly coupled SLAM method based on gradient descent optimization provided by the present invention includes the following steps in step S140: Step S141: Based on the visual residual, geometric residual, inertial contrast residual and anchor point residual in the initial pose graph, construct the overall error function, and calculate the error partial derivative of each node according to the overall error function to obtain the gradient vector of each node.

[0079] Step S142: Based on the gradient vector, update and normalize the poses of each node in the initial pose graph according to the calibrated learning rate, and calculate the error change of the updated node poses. When the error change is lower than the convergence threshold, output the optimized pose graph.

[0080] In a specific embodiment, this invention provides a multimodal tightly coupled SLAM method based on gradient descent optimization. Step 4 involves tightly coupled iterative optimization based on gradient descent. The initial pose map output in step 3 is used as the optimization object, and a unified error function is constructed to model the multimodal observation residuals and label residuals in a unified manner. During gradient descent, the strong geometric convergence direction provided by the label constraint is utilized to avoid optimization getting trapped in local optima, and the pose and map estimation are iteratively adjusted until the error converges. By adding label observation constraints, the risk of local optima in narrow environments and under metallic reflection is significantly reduced, improving convergence stability. This includes the following sub-steps: Sub-step 4.1: Unified modeling of the error function.

[0081] Specifically, the visual residual, wall geometric residual, cylindrical geometric residual, inertial residual, and reflection mark anchor point residual obtained above are combined into a unified overall error to obtain the overall error function.

[0082] Sub-step 4.2: Gradient calculation and direction determination.

[0083] Specifically, based on the overall error function, the partial derivative of the error function for each node is calculated to obtain a gradient vector containing the partial derivatives of the position and attitude of each node. The expression is: ; In the formula, This indicates the overall error's sensitivity to the unknown. This indicates the sensitivity of the overall error to attitude. In gradient calculation, the reflection-labeled residual is used to provide a stable convergence direction to suppress local optima.

[0084] Sub-step 4.3: Iterative update of pose variables.

[0085] Specifically, according to the learning rate calibrated in the experiment. (Constant, range 10) -4 ~10 -2 The pose is updated using the following expression: ; In the formula, ( , The updated node pose is then normalized to ensure it is a unit quaternion.

[0086] Sub-step 4.4: Convergence determination and final result generation.

[0087] Specifically, based on the updated node pose obtained in sub-step 4.3, the error change of the updated node pose is calculated. That is, the absolute value of the difference between the overall error of the updated node pose and the overall error before the update is calculated by calling the overall error function. If the error change is less than a set threshold (constant, 10), the error is considered to be less than the absolute value of the error before the update. -6 ~10 -4 If convergence is not achieved, the iteration continues.

[0088] Step S150: Perform global consistency verification on the optimized initial pose graph, and output the environment modeling results and pose state of the smart mobile device after the verification is passed.

[0089] In a specific embodiment, this invention provides a multimodal tightly coupled SLAM method based on gradient descent optimization. Step 5 involves global consistency verification and final result output. A global consistency check is performed on the optimization results output in step 4 to ensure that the map has no distortion in its overall structure and that pose estimation is stable and reliable. If the verification passes, the final high-precision parking lot environment map and vehicle pose results are output, solving the problem of SLAM convergence difficulties caused by the narrow space and metal interference in underground parking lots. This includes the following sub-steps: Sub-step 5.1, global consistency check.

[0090] Specifically, check whether the residuals are within a reasonable range on the overall map structure, and confirm that the map boundaries are continuous, the column and wall geometry conforms to the structural characteristics, and the pose trajectory is smooth and without abrupt changes. If all the check conditions are met, it is marked as "passed"; if there is local distortion or trajectory instability, it is marked as "failed".

[0091] Sub-step 5.2, final result output.

[0092] Specifically, if the verification result obtained in sub-step 5.1 is "passed", then the globally consistent high-precision map and the corresponding pose estimate are directly exported; if the verification result is "failed", then the iteration optimization of the convergence residual continues.

[0093] The following describes a multimodal tightly coupled SLAM device based on gradient descent optimization provided by the present invention. The multimodal tightly coupled SLAM device based on gradient descent optimization described below and the multimodal tightly coupled SLAM method based on gradient descent optimization described above can be referred to in correspondence.

[0094] like Figure 2 As shown, in one embodiment, a multimodal tightly coupled SLAM device based on gradient descent optimization includes a data acquisition module, a feature set construction module, a pose graph construction module, a pose graph optimization module, and a pose graph verification module.

[0095] The data acquisition module is used to collect environmental feature data of the smart mobile device in real time during its movement, and to determine the spatial position of the smart mobile device by detecting reflective markers, so as to construct a time-synchronized multimodal observation set.

[0096] The feature set construction module is used to identify and extract multimodal features from the multimodal observation set, and to perform feature fusion on the multimodal features to obtain a unified encoded joint feature set.

[0097] The pose graph construction module is used to construct an initial pose graph based on a joint feature set, which includes multimodal feature constraints and marker anchor constraints.

[0098] The pose graph optimization module is used to call the error function to model the multimodal observation residuals and label residuals, and to optimize the initial pose graph by iteratively adjusting the pose and modeling estimation errors.

[0099] The pose graph verification module is used to perform global consistency verification on the optimized initial pose graph, and outputs the environmental modeling results and the pose state of the smart mobile device after the verification is passed.

[0100] The environmental feature data includes environmental image data and laser echoes. The multimodal features are corners and edges identified from the environmental image data, geometric features of environmental entities extracted from the laser echoes, and reflective marks identified from the high points of reflective marks.

[0101] The applicant of this invention has provided a detailed description of the embodiments of the invention in conjunction with the accompanying drawings. However, those skilled in the art should understand that the above embodiments are merely preferred embodiments of the invention. The detailed description is only intended to help readers better understand the spirit of the invention and is not intended to limit the scope of protection of the invention. On the contrary, any improvements or modifications made based on the inventive spirit of the invention should fall within the scope of protection of the invention.

Claims

1. A multimodal tightly coupled SLAM method based on gradient descent optimization, characterized in that, The method includes: Real-time acquisition of environmental feature data during the movement of intelligent mobile devices, and determination of the spatial position of the intelligent mobile devices by detecting reflective marks, in order to construct a time-synchronized multimodal observation set; Multimodal features are identified and extracted from the multimodal observation set, and feature fusion is performed on the multimodal features to obtain a unified encoded joint feature set; Based on the joint feature set, an initial pose graph with multimodal feature constraints and marker anchor point constraints is constructed; The error function is called to model the multimodal observation residuals and the label residuals, and the initial pose map is optimized by iteratively adjusting the pose and the modeling estimation error; Perform a global consistency check on the optimized initial pose graph, and output the environmental modeling results and the pose state of the smart mobile device after the check passes. The environmental feature data includes environmental image data and laser echoes. The multimodal features are corners and edges identified from the environmental image data, geometric features of environmental entities extracted from the laser echoes, and reflective marks identified from the high points of reflective marks.

2. The multimodal tightly coupled SLAM method based on gradient descent optimization according to claim 1, characterized in that, The real-time acquisition of environmental feature data during the movement of the intelligent mobile device, and the determination of the spatial position of the intelligent mobile device by detecting reflective markers, to construct a time-synchronized multimodal observation set, includes: Control the camera to acquire environmental image data during the movement of the smart mobile device at a set frame rate; The laser ranging unit rotates and emits a laser at a fixed scanning frequency and receives the reflected signal to calculate the distance and angle between the smart mobile device and environmental entities during its movement, thereby generating a three-dimensional point set of the laser echo. The acceleration and angular velocity output by the intelligent mobile device body are obtained through the inertial measurement unit to calculate the inertial data of the intelligent mobile device; A time series is established according to a unified time step, and linear interpolation is performed on the environmental image data, 3D point set, and inertial data based on the time series. At the same time, the 3D point set is transformed into the coordinate system of the intelligent mobile device according to the calibrated extrinsic parameter matrix.

3. The multimodal tightly coupled SLAM method based on gradient descent optimization according to claim 2, characterized in that, The method of real-time acquisition of environmental feature data during the movement of the intelligent mobile device, and determination of the spatial position of the intelligent mobile device by detecting reflective markers to construct a time-synchronized multimodal observation set, also includes: Threshold segmentation is performed on the environmental image data to extract the set of pixels whose brightness exceeds a first threshold, and the weighted centroid of the region of the pixel set is calculated. The weighted centroid is converted into a ray direction vector, and the ray direction vector is normalized. At the same time, candidate reflection points with a reflection intensity not lower than the second threshold are selected from the three-dimensional point set. When the angle between the candidate reflection point and the ray direction vector is within a set threshold range, the candidate reflection point is marked for reflection, and the three-dimensional position coordinates of the reflection mark are obtained.

4. The multimodal tightly coupled SLAM method based on gradient descent optimization according to claim 1, characterized in that, The step of identifying and extracting multimodal features from the multimodal observation set, and then fusing the multimodal features to obtain a unified encoded joint feature set, includes: Based on the reflection marker observation results in the multimodal observation set, the number of times the reflection marker appears in consecutive frames and the spatial consistency of the reflection marker between the three-dimensional point set and the environmental image data are verified in order to perform preliminary filtering of the reflection marker observation results. Edge points and corner points of environmental entities are extracted from the three-dimensional point set, and the coordinates of the reflection markers in the preliminary filtered reflection marker observation results are aligned with the observation timestamps of the edge points and corner points of the environmental entities to construct the unified coding joint feature set.

5. The multimodal tightly coupled SLAM method based on gradient descent optimization according to claim 4, characterized in that, The construction of the initial pose graph based on the joint feature set, including multimodal feature constraints and labeled anchor point constraints, includes: The motion state of the intelligent mobile device at the first moment is estimated based on the inertial data in the joint feature set, and the estimated position of the intelligent mobile device at the first moment is associated with the environmental image data and the three-dimensional point set at the first moment. The angular velocity of the smart mobile device at the first moment is converted into the device attitude, and the device attitude is bound with the associated environmental image data and 3D point set to determine the pose node of each frame of observation data, thus obtaining the pose node set.

6. The multimodal tightly coupled SLAM method based on gradient descent optimization according to claim 5, characterized in that, The step of constructing an initial pose graph with multimodal feature constraints and labeled anchor point constraints based on the joint feature set further includes: The matching pixels in adjacent environmental image frames are determined by feature point matching, and the corresponding points and pose nodes in the 3D point set are aligned and projected onto the environmental image coordinate system to calculate the visual residual between the 3D point set and pose nodes and the matching pixels, thereby obtaining the visual constraint edge. The geometric residuals of the wall corner points and column edge points in the joint feature set are defined to determine the geometric constraint edges. The inertial data and time step are used to determine the comparison residuals between the predicted pose and the pose nodes to obtain the inertial constraint edges.

7. The multimodal tightly coupled SLAM method based on gradient descent optimization according to claim 6, characterized in that, The step of constructing an initial pose graph with multimodal feature constraints and labeled anchor point constraints based on the joint feature set further includes: Each reflection marker point in the preliminary filtered reflection marker observation results is represented by three-dimensional coordinates, and the three-dimensional coordinate representation of the reflection marker point is transformed from the local coordinate system to the global coordinate system. The anchor point residual is calculated based on the corresponding reflection marker point and anchor point reference in the global coordinate system to generate reflection marker constraints. The visual constraint edges, geometric constraint edges, inertial constraint edges, and reflection mark constraints are merged and combined with the pose node set to generate the initial pose graph.

8. The multimodal tightly coupled SLAM method based on gradient descent optimization according to claim 1, characterized in that, The method of calling the error function to model the multimodal observation residuals and the labeled residuals, and optimizing the initial pose map by iteratively adjusting the pose and modeling estimation errors, includes: Based on the visual residual, geometric residual, inertial contrast residual, and anchor point residual in the initial pose graph, an overall error function is constructed, and the error partial derivative of each node is calculated according to the overall error function to obtain the gradient vector of each node. Based on the gradient vector, the poses of each node in the initial pose graph are updated and normalized according to the calibrated learning rate, and the error change of the updated node poses is calculated. When the error change is lower than the convergence threshold, the optimized pose graph is output.

9. A terminal, comprising a processor and a storage medium; characterized in that: The storage medium is used to store instructions; The processor is configured to operate according to the instructions to execute the steps of the gradient descent-optimized multimodal tightly coupled SLAM method according to any one of claims 1-8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the steps of the gradient descent-optimized multimodal tightly coupled SLAM method as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Dynamic vision SLAM (Simultaneous Localization and Mapping) method

    CN118447090A

  • Photovoltaic cleaning robot based on automatic path planning and operation method thereof

    CN119472645A

  • Interaction method and system based on deep learning

    CN119919749A

  • Multi-data source slam method, device, and computer readable storage medium

    WO2022036792A1

  • Three-dimensional map construction method and apparatus, storage medium and chip system

    WO2024222149A1

Cited By

  • Spacecraft magnetic moment attitude control method based on non-inverse iterative algorithm

    CN122009527A