High-precision AI positioning method and system based on spatial multi-modal data fusion
By collecting multimodal data to analyze environmental attributes in real time, dynamically adjusting the fusion weight, and optimizing positioning with recursive neural network, the problem of positioning stability and environmental adaptability in satellite signal denial scenarios is solved, achieving high-precision and robust positioning effects.
Patent Information
- Application Number
- CN202510998413.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-07-21
AI Technical Summary
The existing fusion schemes are prone to fluctuation in the transition phase of complex environments due to excessive dependence on inertial navigation in satellite signal denial scenarios, and the static configuration of fusion weights cannot respond to fluctuations in real time.
By collecting satellite positioning signals, inertial measurement unit data, visual image sequences and lidar point cloud data, the environment physical attribute parameters are analyzed in real time, the fusion weight of each modal feature is dynamically adjusted, and positioning optimization is used using recursive neural networks, and error compensation and cross-modal feature alignment are combined with high-precision maps.
It realizes maintaining high-precision positioning in satellite signal denial scenarios, reducing the cumulative impact of inertial navigation errors, responding to environmental parameter fluctuations in real time, and improving positioning stability and accuracy in the transition stage of complex environments.
Smart Images

Figure CN120491127A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of fusion positioning technology, and specifically to a high-precision AI positioning method and system based on spatial multimodal data fusion. Background Art
[0002] As a core capability of intelligent mobile carriers, spatial positioning technology plays a fundamental role in autonomous driving, drone navigation, augmented reality, and other fields. As application scenarios expand into urban canyons, underground spaces, and harsh weather environments, traditional single-sensor positioning solutions face challenges in both accuracy and reliability. Multimodal data fusion technology, integrating heterogeneous data sources such as satellite signals, inertial measurement, visual perception, and lidar, has become a mainstream development direction for improving positioning performance in complex environments.
[0003] Among existing technologies, sensor fusion solutions based on the Kalman filter framework have become a relatively mature technical path. This approach uses absolute position information provided by the Global Navigation Satellite System (GNSS) as a benchmark, combined with high-frequency pose estimation from an inertial measurement unit (IMU) to achieve data complementarity. Some improved solutions further incorporate visual odometry or lidar point cloud matching results, integrating multi-source observations into a unified solution model by expanding the filter state. By establishing a statistical model of sensor errors, this technology can effectively improve positioning consistency in open environments.
[0004] Although existing fusion solutions have made certain progress, their technical implementation still has inherent constraints: in terms of environmental adaptability, the system over-relies on the cumulative error compensation mechanism of inertial navigation in satellite signal denial scenarios, such as tunnels and indoor environments. The long-term positioning stability is constrained by the accuracy of the motion model, making it difficult to maintain continuous and reliable positioning output; in terms of dynamic adjustment capabilities, the fusion weight parameters are usually statically configured based on preset environmental classifications, and cannot respond in real time to continuous fluctuations in physical environmental parameters such as sudden changes in illumination and attenuation of atmospheric transmittance, resulting in fluctuations in the system's positioning accuracy during the transition phase of complex environments. These factors mean that the positioning accuracy and robustness of existing solutions in complex scenarios still have room for improvement. Summary of the Invention
[0005] The purpose of this invention is to provide a high-precision AI positioning method and system based on spatial multimodal data fusion to solve the following technical problems: The existing fusion schemes are overly dependent on inertial navigation in satellite signal denial scenarios, resulting in limited long-term positioning stability. In addition, the static configuration of fusion weights cannot respond to environmental parameter fluctuations in real time, making the accuracy prone to fluctuations in complex environment transition phases.
[0006] The purpose of the present invention can be achieved through the following technical solutions: The high-precision AI positioning method based on spatial multimodal data fusion includes the following steps: S1. Collect spatial multimodal data, including satellite positioning signals, inertial measurement unit data, visual image sequences, lidar point clouds, and high-precision map data; S2, generates a set of physical property parameters of the current environment by real-time analysis of the illumination distribution in the visual image, the atmospheric particle density distribution in the lidar point cloud, and the multipath reflection path of the satellite signal; S3: Input the original observation value of the satellite positioning signal, the motion trajectory of the inertial measurement unit, the key feature points of the visual image, and the geometric structure of the lidar point cloud into the correction function bound to the physical property parameters of the environment, and output the standardized feature vector after environmental error compensation; S4. Input the normalized feature vector into a recursive neural network with spatiotemporal memory capability. The recursive neural network outputs the six-degree-of-freedom pose of the carrier by dynamically adjusting the fusion weight coefficients of each modal feature.
[0007] As a further solution of the present invention: in S2, the process of generating the physical property parameter set is: Extract scene illumination intensity gradient maps frame by frame from visual image sequences, and generate illumination stability quantitative indicators by calculating the rate of change of illumination intensity gradients between adjacent frames. Identify the three-dimensional spatial density distribution of atmospheric suspended particles from the lidar point cloud and construct an atmospheric transmittance attenuation function based on the inverse relationship between the point cloud reflection intensity and the transmission distance; Separate the propagation delay difference between the direct path and the multipath reflection path from the satellite positioning signal, and combine it with the building facade geometry data in the high-precision map to generate a multipath interference probability distribution map; The illumination stability quantitative index, the atmospheric transmittance attenuation function and the multipath interference probability distribution map are combined into an environmental physical property parameter set.
[0008] As a further solution of the present invention: in S3, the process of outputting the normalized feature vector after environmental error compensation is: For the original observation value of the satellite positioning signal, call the multipath interference probability distribution map, superimpose the building reflection weight on the signal propagation path, and reconstruct the solution parameters of the pseudo-range observation equation; For the motion trajectory of the inertial measurement unit, the air resistance compensation is performed on the motion acceleration according to the atmospheric transmittance attenuation function, and the kinematic constraints in the carrier dynamics equation are reconstructed; For key feature points of visual images, the feature point extraction threshold is dynamically adjusted based on the quantitative index of illumination stability. When the illumination gradient change rate exceeds the set threshold, the static object outline in the high-precision map is used as an auxiliary matching benchmark. For the geometric structure of the lidar point cloud, the refractive index of the point cloud coordinates is corrected by combining the atmospheric transmittance attenuation function, the three-dimensional spatial topological relationship is reconstructed, and the spatial consistency alignment of multimodal features is achieved.
[0009] As a further solution of the present invention: in S4, the process of the recursive neural network dynamically adjusting the fusion weight coefficients of each modal feature is: A gating mechanism based on a set of environmental physical property parameters is established. When the output value of the atmospheric transmittance attenuation function is lower than a critical threshold, the weight coefficient of the lidar point cloud feature is increased. When the proportion of strong interference areas in the multipath interference probability distribution map exceeds the set ratio, the weight coefficient of the satellite positioning signal feature is reduced; When the illumination stability quantitative index fluctuates continuously for more than a set number of times, the matching confidence between the visual image features and the static features on the HD map is used as the basis for weight allocation; The fusion weight coefficients of each modal feature are updated in real time through the normalized scalar value output by the gating mechanism and fed back to the computing unit of the recurrent neural network.
[0010] As a further solution of the present invention: in S4, the spatiotemporal memory capability is specifically: The recursive neural network maintains a memory matrix of the carrier's motion state, which stores the six-degree-of-freedom pose at each historical moment and the corresponding set of environmental physical property parameters. When it is detected that the similarity between the current environmental physical attribute parameter set and a record in the historical memory matrix exceeds the set threshold, the fusion weight coefficient of the historical moment is called as the initialization benchmark; When the carrier enters an area without satellite signals, such as a tunnel or indoors, a motion trajectory prediction model is constructed based on the characteristics of the most recent valid satellite positioning signal in the memory matrix to maintain positioning continuity.
[0011] As a further solution of the present invention: in S4, the six-degree-of-freedom pose of the output carrier includes: When continuous strong light or thick fog is detected, the coupling association between visual image features and lidar point cloud features is locked by disabling cross-modal attention calculation; When the satellite signal interruption duration exceeds the maximum reliable operating time limit of the inertial measurement unit, the topological road network in the high-precision map is used as the motion constraint boundary to limit the carrier's position estimation range; After each pose update, the current environment physical property parameter set is bound to the six-degree-of-freedom pose data and stored in an anti-tampering cache to form a traceable positioning evidence chain.
[0012] As a further solution of the present invention: the data structure of the anti-tampering cache area is: A blockchain-based storage structure indexed by timestamps, where each node contains a snapshot of the environment's physical property parameter set, a fusion weight coefficient set, and a six-degree-of-freedom pose; The nodes generate verification hash values through the carrier motion continuity equation. Any tampering of node data will cause the subsequent node hash chain to break. When the positioning system is restarted or the positioning anomaly is recovered, the recursive neural network is reinitialized from the nearest valid hash node, the corresponding environmental physical property parameter set is loaded, and the historical fusion weight coefficient configuration is inherited.
[0013] As a further solution of the present invention: in S1, the application of the high-precision map data in the spatial multimodal data includes: In S2, the electromagnetic wave reflection coefficient of the building material in the high-precision map is extracted to verify the rationality of the satellite signal multipath reflection path; In S3, the lane line curvature in the high-precision map is coupled with the trajectory curvature calculated by the inertial measurement unit to generate a trajectory smoothness constraint function; In S4, the output six-degree-of-freedom pose is spatially consistent with the topological relationship of the static objects in the high-precision map. If the verification deviation exceeds the set tolerance, the recursive neural network weight reset is triggered.
[0014] As a further solution of the present invention: the process of the spatial consistency check is: Extract the 3D coordinates of all static objects within a set radius centered on the output 6DOF pose from the HD map; Project the LiDAR point cloud into the three-dimensional coordinate space and calculate the shortest distance set between the point cloud and the surface of each feature model; When the standard deviation of the distance set exceeds the set multiple of the historical mean, it is judged as a posture abnormality. When the posture abnormality is triggered, the fusion weight coefficient of the recurrent neural network at the current moment is reset to the preset safety value, and the six-degree-of-freedom posture is recalculated.
[0015] The present invention also includes a high-precision AI positioning system based on spatial multimodal data fusion, which is used to implement the above-mentioned high-precision AI positioning method based on spatial multimodal data fusion, including: Spatial data acquisition module, used to collect spatial multimodal data, including satellite positioning signals, inertial measurement unit data, visual image sequences, lidar point clouds and high-precision map data; The environmental attribute analysis module is used to generate a set of physical attribute parameters of the current environment by real-time analysis of the illumination distribution in the visual image, the atmospheric particle density distribution in the lidar point cloud, and the multipath reflection path of the satellite signal; The cross-modal alignment module is used to input the raw observation values of the satellite positioning signal, the motion trajectory of the inertial measurement unit, the key feature points of the visual image, and the geometric structure of the lidar point cloud into the correction function bound to the physical property parameters of the environment, and output a standardized feature vector after environmental error compensation; The fusion positioning output module is used to input the standardized feature vector into a recursive neural network with spatiotemporal memory capability. The recursive neural network outputs the six-degree-of-freedom pose of the carrier by dynamically adjusting the fusion weight coefficients of each modal feature.
[0016] Beneficial effects of the present invention: The present invention solves the problems of limited environmental adaptability and insufficient dynamic adjustment capabilities in the existing technology by constructing a dynamic environmental physical property model, performing cross-modal feature alignment under physical constraints, and generating environmentally adaptive fusion positioning results. By real-time analysis of multimodal data such as visual images, lidar point clouds, and satellite signals, a set of environmental physical property parameters is generated to achieve active perception and modeling of the environment; through a correction function bound to the environmental physical property parameters, environmental error compensation is performed on each modal feature to achieve deep alignment of cross-modal features; through the gating mechanism of the recursive neural network, the fusion weight coefficient is dynamically adjusted based on the environmental physical parameters to achieve adaptive optimization of the positioning process. The present invention enables the positioning system to maintain high-precision positioning in satellite signal denial scenarios, effectively reducing the impact of accumulated inertial navigation errors, while responding to fluctuations in environmental parameters such as illumination and atmospheric transmittance in real time, improving positioning stability in the transition phase of complex environments, and significantly enhancing the positioning accuracy and robustness of the system in extremely complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The present invention will be further described below with reference to the accompanying drawings.
[0018] Figure 1 Schematic diagram of the process of the high-precision AI positioning method based on spatial multimodal data fusion of the present invention; Figure 2 It is a module schematic diagram of the high-precision AI positioning system based on spatial multimodal data fusion of the present invention. DETAILED DESCRIPTION
[0019] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.
[0020] See also Figure 1As shown, the present invention is a high-precision AI positioning method based on spatial multimodal data fusion, comprising the following steps: S1. Collect multimodal spatial data, including satellite positioning signals, motion data output by inertial measurement units, visual image sequences, point cloud data from lidar scans, and environmental prior information from high-precision maps. These data depict the spatial state of the vehicle from different dimensions. Satellite positioning signals provide a basic position reference, inertial data reflects real-time motion acceleration and angular velocity, visual images capture scene texture and lighting characteristics, laser point clouds represent the three-dimensional spatial geometry, and high-precision maps contain precise information about static environments such as roads and buildings.
[0021] S2. Generate scene lighting stability indicators by real-time analysis of the light intensity distribution of pixels in visual images; identify the spatial density distribution of atmospheric suspended particles from the lidar point cloud and construct an atmospheric transmittance attenuation function; separate the propagation delay difference between the direct path and the multipath reflection path from the satellite positioning signal, and combine it with the geometric data of the building facades in the high-precision map to generate a multipath interference probability distribution map. These indicators, functions and distribution maps are combined to form a set of physical property parameters of the current environment.
[0022] S3. For the original observation values of the satellite positioning signal, call the multipath interference probability distribution map, superimpose the building reflection weight on the signal propagation path, and reconstruct the solution parameters of the pseudo-range observation equation; for the motion trajectory of the inertial measurement unit, compensate the motion acceleration for air resistance according to the atmospheric transmittance attenuation function, and reconstruct the kinematic constraints in the carrier dynamics equation; for the key feature points of the visual image, dynamically adjust the extraction threshold according to the illumination stability index, and enable the static terrain contour in the high-precision map as an auxiliary matching benchmark when the illumination gradient change rate is large; for the geometric structure of the lidar point cloud, combine the atmospheric transmittance attenuation function to correct the refractive index of the point cloud coordinates, and reconstruct the three-dimensional spatial topological relationship. After these processes, the standardized feature vector after environmental error compensation is output.
[0023] S4. The standardized feature vector is input into a recurrent neural network with spatiotemporal memory. This network dynamically adjusts the fusion weight coefficients of each modal feature through a gating mechanism conditioned by a set of environmental physical property parameters: The weight of the lidar point cloud feature is increased when the output value of the atmospheric transmittance attenuation function is low; the weight of the satellite positioning signal feature is reduced when the proportion of strong interference areas in the multipath interference probability distribution map is high; and the confidence level of the match between the visual image features and the static features on the HD map is used as the basis for weight allocation when the illumination stability index fluctuates continuously. Simultaneously, the network maintains a memory matrix of the carrier's motion state, storing the six-degree-of-freedom poses and corresponding environmental parameters at historical moments, and ultimately outputs the carrier's six-degree-of-freedom pose.
[0024] In S2, the process of generating the physical property parameter set is as follows: A scene illumination gradient map is extracted frame by frame from a visual image sequence. For example, a vehicle may pass through a tree-lined avenue in the early morning, an open road at noon, and a backlit section in the evening. On the tree-lined avenue, sunlight filtering through leaves creates dappled light and shadow, resulting in a distinct illumination gradient at the boundary between light and dark areas in the image. Upon entering the open road, the overall illumination becomes uniform, and the gradient decreases. In the evening, backlit, the scenery ahead of the vehicle creates high-contrast outlines due to strong illumination, causing the gradient to change significantly again. By calculating the rate of change of the illumination gradient for the same area in adjacent frames under these scenarios, a quantitative indicator of illumination stability can be generated, which accurately reflects the dynamic changes in illumination conditions. For example, if the rate of change of the gradient between adjacent frames increases sharply at the moment of backlighting, the indicator value will increase accordingly, intuitively reflecting illumination instability.
[0025] Identify the three-dimensional spatial density distribution of atmospheric suspended particles from LiDAR point clouds. In light rain, the laser beam emitted by the LiDAR scatters off atmospheric particles such as raindrops, resulting in abnormal reflection intensities in some point clouds. By analyzing the distribution characteristics of reflection intensity in the point cloud data, differences in the density of atmospheric suspended particles in different regions can be identified. Roadside vegetation may contain more water vapor particles due to transpiration, resulting in more pronounced attenuation of point cloud reflection intensity in these areas. Based on the inverse relationship between point cloud reflection intensity and transmission distance—that is, the farther away from the LiDAR, the greater the attenuation of reflection intensity due to particle scattering—an atmospheric transmittance attenuation function is constructed. This function quantifies the transmission loss of the laser signal at different distances, providing a basis for subsequent point cloud correction.
[0026] The propagation delay differences between the direct path and multipath reflection paths are separated from the satellite positioning signal. When a vehicle travels through a mixed neighborhood with both low-rise and high-rise buildings, the satellite signal propagation paths are complex and diverse. Some signals reach the receiver directly, forming direct paths, while others are reflected off the slanted surfaces of low-rise buildings and the vertical walls of high-rise buildings, forming multiple multipath paths. Using signal processing techniques to separate the propagation time differences between these paths, combined with the building facade geometry data in the area from high-precision maps (e.g., a high-rise building with a height of 80 meters and a wall angle of 90 degrees, and a low-rise building with a height of 10 meters and a wall angle of 60 degrees), the source and strength of each reflection path can be accurately determined, generating a multipath interference probability distribution map. Around high-rise buildings, due to the numerous reflection paths and strong signals, the interference probability in the corresponding areas of the distribution map is significantly higher than in other areas.
[0027] The above-generated quantitative indicators of illumination stability, atmospheric transmittance attenuation function and multipath interference probability distribution map are integrated to form a parameter set that can reflect the current physical state of the environment in real time, providing an accurate basis for error compensation of each modal data.
[0028] In S3, the process of outputting the normalized feature vector after environmental error compensation is as follows: The original observations of satellite positioning signals are corrected using a multipath interference probability distribution map. Around super-high-rise buildings with high multipath interference probability, building reflection weights are superimposed on the signal propagation path based on the interference probability values in that area of the distribution map. Areas with greater reflected signal contribution receive more significant weight adjustments to the corresponding pseudorange observations. This reconstructs the solution parameters of the pseudorange observation equation and reduces positioning errors caused by multipath effects.
[0029] For the IMU's trajectory, air drag compensation is applied to the acceleration using an atmospheric transmittance attenuation function. During light rain, when atmospheric transmittance is low, the increased air density increases the vehicle's air drag. The acceleration data recorded by the IMU includes the effects of this additional drag. The drag coefficient at different speeds is calculated using the attenuation function, compensating the acceleration data and reconstructing the kinematic constraints within the vehicle's dynamics equations, ensuring that the trajectory is more accurately aligned with the vehicle's actual motion.
[0030] For key feature points in visual images, the extraction threshold is dynamically adjusted based on a quantitative indicator of lighting stability. When a vehicle rapidly exits a tree-lined avenue into midday sunlight, the sudden increase in light intensity causes the rate of change in the illumination gradient between frames to exceed the set threshold. The threshold previously applied to feature point extraction for the tree-lined avenue is no longer applicable. The threshold is automatically lowered to accommodate the bright light environment. At the same time, the static feature outlines of that section of road in the HD map, such as the precise positions of guardrails and light poles, are used as auxiliary matching benchmarks to ensure stable and reliable extraction of visual feature points despite drastic lighting fluctuations.
[0031] The geometric structure of the lidar point cloud is corrected for the refractive index of the point cloud coordinates in combination with the atmospheric transmittance attenuation function. In foggy weather, the laser beam is refracted by atmospheric particles, resulting in deviations in the measured distance. The deviation increases with the distance. The attenuation function is used to calculate the refractive index correction values at different distances, and the three-dimensional coordinates of each point cloud are calibrated. For example, the coordinates of a building corner point that was originally displayed as 100 meters away are closer to its actual position after correction. Through this correction, the three-dimensional spatial topological relationship is reconstructed, so that the geometric structure of the laser point cloud is precisely aligned in space with the ground features in the visual image and the position information of satellite positioning. Finally, a standardized feature vector is output after environmental error compensation, providing high-quality feature input for subsequent fusion positioning.
[0032] In S4, the process of the recursive neural network dynamically adjusting the fusion weight coefficients of each modal feature is as follows: In an urban canyon environment, when a vehicle enters an area densely populated with high-rise buildings and the multipath interference probability distribution map indicates a higher percentage of strong interference, the recurrent neural network uses a gating mechanism to reduce the weight of satellite positioning signal features. For example, at an intersection surrounded by three super-high-rise buildings, the interference value in the multipath interference probability distribution map for that area reaches 0.8 (exceeding the set percentage of 0.6). The gating mechanism then reduces the weight of the satellite positioning signal from the default 0.4 to 0.1. Simultaneously, the lidar point cloud data, which effectively captures the geometric features of building facades and outputs an atmospheric transmittance attenuation function value of 0.9 (above the critical threshold of 0.7), increases its weight from 0.3 to 0.5. If the visual image's illumination stability quantification indicator fluctuates for more than a set number of times (e.g., five times), the system uses the static feature outlines in the HD map (e.g., road edges, building corners) as auxiliary matching benchmarks, and uses the matching confidence level between the visual image and the HD map as the basis for weighting. If the visual features are successfully matched with the HD map, the visual image feature weight coefficient is increased from 0.2 to 0.3, and finally the normalized scalar values (0.1, 0.3, 0.5, 0.1) are output through the gating mechanism, corresponding to the fusion weights of satellite positioning, visual image, lidar, and inertial measurement unit respectively.
[0033] When the vehicle enters a tunnel, the output of the atmospheric transmittance attenuation function drops sharply to 0.3 (below the critical threshold of 0.7). The recurrent neural network immediately increases the weight of the lidar point cloud features to 0.8. Simultaneously, due to the absence of satellite signals in the tunnel, the multipath interference probability distribution map becomes invalid, and the satellite positioning signal weight drops to 0. At this point, due to the drastic lighting variations in the tunnel, the illumination stability quantification index fluctuates continuously for more than a set number of times. The system then relies on matching the lidar point cloud with the tunnel's internal geometric features (such as the tunnel walls and lane markings) in the HD map, reducing the visual image weight to 0.1. The inertial measurement unit data maintains a weight of 0.1 because it provides reliable short-term motion trajectories. The gating mechanism dynamically generates normalized weights (0, 0.1, 0.8, 0.1) by analyzing a set of environmental physical property parameters in real time, ensuring high-precision positioning in tunnel environments.
[0034] In S4, the spatiotemporal memory capability is specifically: The recurrent neural network maintains a motion state memory matrix that records the vehicle's historical positions and corresponding sets of environmental physical property parameters in different environments. When a vehicle first enters a tunnel from an open road, the memory matrix stores the satellite positioning signal characteristics, lidar point cloud data, six-degree-of-freedom position, and a set of environmental physical property parameters from before entering the tunnel. When the vehicle approaches the tunnel again, the system detects that the similarity between the current set of environmental physical property parameters (such as illumination stability, atmospheric transmittance, and multipath interference probability) and the tunnel entrance record in the memory matrix exceeds a set threshold (such as 0.9). The system immediately uses the fusion weight coefficients (0, 0.1, 0.8, 0.1) at that historical moment as an initialization baseline, enabling the network to quickly adapt to the tunnel environment and reducing positioning errors during the transition phase.
[0035] When the vehicle fully enters the tunnel, causing the satellite signal to be interrupted, the recurrent neural network constructs a trajectory prediction model based on the characteristics of the most recent valid satellite positioning signal in the memory matrix (i.e., the positioning data at the tunnel entrance) and the motion trajectory provided by the inertial measurement unit. This model uses historical motion parameters such as acceleration and angular velocity stored in the memory matrix to predict the vehicle's likely trajectory within the tunnel. For example, if the memory matrix shows that the vehicle was traveling in a straight line at 50 km / h with a steering wheel angle of 0 degrees before entering the tunnel, the model will predict that the vehicle will continue to move in a straight line within the tunnel and fine-tune the predicted trajectory based on the real-time acceleration data updated by the inertial measurement unit to maintain positioning continuity until the vehicle exits the tunnel and regains a valid satellite signal.
[0036] During heavy rain, the output of the atmospheric transmittance attenuation function consistently falls below a critical threshold, degrading the quality of the lidar point cloud data. The recurrent neural network then uses its memory matrix to retrieve historical positioning data from similar weather conditions. It finds that when the atmospheric transmittance is below 0.5, the matching success rate between road markings in the visual image and the HD map is higher. The system adjusts its fusion strategy accordingly, using the matching results between the visual image and the HD map as the primary basis for positioning while reducing the weight of the lidar point cloud. This historically informed weighting mechanism effectively improves the system's positioning robustness in extreme weather conditions.
[0037] As the vehicle exits the tunnel and enters the signal recovery area, the recurrent neural network detects that satellite signal quality is gradually improving and the probability of multipath interference is decreasing. The system compares the current set of physical property parameters with the historical records of tunnel exits in the memory matrix. If the similarity exceeds a set threshold, the corresponding historical fusion weight coefficient is used as the initial value and dynamically adjusted based on the current real-time data. For example, the memory matrix shows that the typical weight coefficients at the tunnel exit are (0.3, 0.2, 0.4, 0.1). Based on this, the system fine-tunes the weight coefficients to (0.4, 0.1, 0.4, 0.1) based on the current actual signal quality, achieving a smooth transition from the tunnel environment to the open air.
[0038] In an urban elevated road scenario, when a vehicle reaches the overpass, the 3D topological information provided by the HD map indicates a large billboard above the current location, potentially blocking the satellite signal. The recursive neural network, combined with similar scene data stored in a memory matrix, preemptively reduces the weight of the satellite positioning signal and strengthens the matching of the LiDAR point cloud with the HD map. If the system detects billboard texture features in the visual image that are similar to those in the memory matrix, it further verifies the similarity between the current environment and historical scenes, thereby more accurately adjusting the fusion weights and improving positioning accuracy.
[0039] This spatiotemporal memory capability is also evident during long-distance driving. When a vehicle returns to an area it has previously traveled after a long journey, the recurrent neural network quickly identifies the area by comparing the current set of environmental physical property parameters with the historical records in the memory matrix, and then uses the corresponding historical fusion weight coefficients and positioning experience. For example, when a vehicle first passes through a mountainous section of road, the memory matrix records the section's complex terrain characteristics, signal conditions, and optimal fusion strategy. When the vehicle re-enters the same section, the system directly reuses this historical experience without the need for relearning, significantly improving positioning efficiency and accuracy.
[0040] In S4, the six-degree-of-freedom pose of the output carrier includes: When sustained strong illumination is detected on a highway under direct midday sunlight, cross-modal attention calculations are disabled during the 6DOF pose output of the vehicle, locking the coupled associations between the visual image features and the LiDAR point cloud features. For example, if direct sunlight causes overexposure in some areas of the visual image, feature associations that rely on cross-modal attention for adjustment may be distorted by the bright light. Locking the coupled associations maintains a fixed association logic between the edge features of road markings in the visual image and the corresponding sudden changes in ground elevation in the LiDAR point cloud, preventing feature matching errors caused by drastic changes in illumination and ensuring the stability of the pose output.
[0041] In dense fog, atmospheric scattering can cause noise in the lidar point cloud, and low visibility can blur the visual image. When the environmental physics model detects this, it disables cross-modal attention computation to maintain the inherent coupling between visual and lidar features. For example, a guardrail point cloud cluster captured by the lidar always corresponds to the guardrail's grayscale outline in the visual image. Even if individual sensor data is noisy, the coupling between the two is not dynamically adjusted, thus ensuring the basic accuracy of the pose solution.
[0042] When a vehicle enters an extremely long tunnel and the satellite signal is completely lost for a period exceeding the maximum reliable operating time of the inertial measurement unit (IMU), the topological road network in the HD map is used as a motion constraint boundary when outputting the 6DOF pose. If the IMU in the tunnel accumulates errors due to long-term calculations, causing the originally straight trajectory to begin to deviate slightly, the tunnel's topological road network in the HD map (such as the number of lanes, turning restrictions, and tunnel wall positions) will form a virtual boundary, limiting the pose calculation range to the physical space defined by the network. For example, if the tunnel is actually a two-lane, bidirectional tunnel, the topological road network will constrain the pose calculation results to within the lane width, preventing the trajectory from deviating from the actual road due to accumulated IMU errors and ensuring that the pose output always matches the actual driving path.
[0043] After each pose update, the current set of environmental physical property parameters and the six-degree-of-freedom pose data are bound and stored in a tamper-resistant cache, forming a traceable chain of positioning evidence. For example, when a vehicle passes an overpass at 3:00 PM, the cache will record environmental parameters such as the illumination stability quantification index, the atmospheric transmittance attenuation function, and the multipath interference probability distribution map, as well as the corresponding 3D coordinates and pose data, along with the fusion weight coefficients of each modal feature at that moment.
[0044] The tamper-resistant cache uses a blockchain-style storage structure indexed by timestamps, with each node linked in chronological order. For example, the node at 3:01 contains a snapshot of environmental parameters, a fusion weight set, and pose data at that moment. When the node at 3:02 is generated, a verification hash value is generated based on the carrier's motion continuity equation and the pose data of the previous node. The two nodes are linked by the hash value. If someone attempts to modify the pose data at 3:01, the hash values of all subsequent nodes will be broken due to verification failure, directly indicating that the data has been tampered with.
[0045] When the positioning system restarts due to a sudden failure or requires recovery after a positioning anomaly on a complex road, it searches the tamper-resistant cache for the nearest valid hash node. For example, after a system restart, it first reads the last node in the cache with an unbroken hash chain. This node contains the set of environmental physical property parameters, fusion weight coefficients, and six-degree-of-freedom pose at the moment before the failure. The recurrent neural network reinitializes based on this data, loads the corresponding environmental parameters, and inherits the historical fusion weight configuration, allowing the positioning process to smoothly continue from the breakpoint, avoiding positioning data loss caused by system interruptions or errors caused by reinitialization.
[0046] In long-distance transport scenarios, vehicles travel continuously for hours, and the tamper-resistant cache continuously records positioning data at every moment. If the positioning accuracy of a specific section of the journey needs to be traced later, the environmental parameters and posture data of the corresponding node can be retrieved through the timestamp index. Using hash verification between nodes to confirm that the data has not been tampered with, the positioning process for that period can be fully restored, providing traceable evidence for the reliability of the positioning results.
[0047] When the vehicle exits the tunnel and reacquires satellite signals, the tunnel pose data and environmental parameters stored in the tamper-resistant cache serve as a reference for subsequent positioning calibration. For example, the node data at the tunnel exit includes the matching characteristics of the lidar point cloud and the tunnel wall, as well as the cumulative error trend of the IMU. When combined with the newly acquired satellite signals, this data can more accurately correct the pose output, ensuring a smooth and seamless transition from tunnel to open environment.
[0048] In S1, the application of the high-precision map data in the spatial multimodal data includes: Application of high-precision maps in S2: Verifying the rationality of satellite signal multipath reflection paths.
[0049] When a vehicle reaches a neighborhood with both glass-walled buildings and concrete residential buildings, the satellite signal, in addition to the direct path, is reflected off the building surfaces, forming multipath signals. At this point, the electromagnetic wave reflection coefficient of the building materials in the area is extracted from the HD map. Glass curtain walls have a higher reflection coefficient, while concrete walls have a lower reflection coefficient. Among the separated multipath reflection paths, if a path is identified as originating from the glass curtain wall, its signal strength should match the high reflection coefficient characteristic. If a path points toward a concrete building but displays a strong reflection signal, combined with the HD map's reflection coefficient data, this path can be identified as abnormal interference. Invalid multipath signals are then eliminated, ensuring that the retained multipath reflection paths conform to the electromagnetic wave reflection characteristics of the material, providing a basis for subsequently generating an accurate multipath interference probability distribution map.
[0050] Application of high-precision maps in S3: Generating trajectory smoothness constraint functions.
[0051] When a vehicle travels on a highway section with continuous curves, the HD map stores the precise curvature data for the lane lines in that section, reflecting the smooth transition characteristics of the road design. The inertial measurement unit (IMU) infers the real-time trajectory curvature based on motion sensor data. If the calculated trajectory curvature experiences localized mutations due to road bumps or sensor noise, it may deviate from the lane curvature in the HD map. At this point, the HD map's lane curvature is coupled with the IMU's trajectory curvature to generate a trajectory smoothness constraint function. This function fine-tunes the IMU's trajectory based on the degree of deviation between the two, ensuring that the corrected trajectory curvature more closely matches the smoothness of the actual road design, avoiding trajectory distortion caused by transient noise and providing more reliable kinematic constraints for the subsequent generation of standardized feature vectors.
[0052] Application of high-precision maps in S4: spatial consistency verification and exception handling.
[0053] When a vehicle is driving on a city's main road, there are static features on both sides of the road, such as streetlights, traffic signal poles, and guardrails. The three-dimensional coordinates of these features are accurately recorded in the high-precision map. After the six-degree-of-freedom pose is output, the spatial consistency check is performed according to the following process: Extract the 3D coordinates of all static features within a set radius centered on the current position from the HD map. For example, the precise location data for 10 streetlights and three sets of guardrails within a 50-meter radius is used. The point cloud scanned by the LiDAR in real time is projected onto the space containing these 3D coordinates. The shortest distance between each point cloud and the corresponding feature model surface is calculated. For example, if the model surface coordinates of a streetlight are (x1, y1, z1), and the coordinates of a point in the projected point cloud are (x2, y2, z2), the straight-line distance between the two is included in the distance set.
[0054] If the vehicle is accurately positioned, the LiDAR point cloud should closely match the object model in the HD map, with a small standard deviation of the distance set. If the output pose deviates due to multipath interference or sensor drift, the distance between the point cloud and the object model will generally increase, and the standard deviation of the distance set will also increase. When this standard deviation exceeds a set multiple of the historical mean, the pose is considered abnormal.
[0055] When a posture anomaly is triggered, the current fusion weight coefficients in the recurrent neural network are reset to a preset safe value. For example, if the weight of the original LiDAR point cloud was lowered due to misjudgment, it will be restored to the default reasonable range after the reset, and the 6DOF posture will be recalculated. During the recalculation, the system will prioritize the topological relationship of static objects in the high-precision map, combined with the matching results of the LiDAR point cloud, to correct the posture deviation and ensure that the output posture is consistent with the actual road environment.
[0056] For example, on a certain road section, sudden strong light caused errors in visual feature extraction, and the output posture shifted toward the guardrail. The spatial consistency check found that the standard deviation of the distance set between the lidar point cloud and the guardrail model far exceeded the historical mean. After being determined to be an anomaly, the fusion weight was reset, the lidar weight returned to normal, and the recalculated posture returned to the center of the road, avoiding positioning failure due to error accumulation.
[0057] See also Figure 2 As shown, the present invention also includes a high-precision AI positioning system based on spatial multimodal data fusion, which is used to implement the above-mentioned high-precision AI positioning method based on spatial multimodal data fusion, including: Spatial data acquisition module, used to collect spatial multimodal data, including satellite positioning signals, inertial measurement unit data, visual image sequences, lidar point clouds and high-precision map data; The environmental attribute analysis module is used to generate a set of physical attribute parameters of the current environment by real-time analysis of the illumination distribution in the visual image, the atmospheric particle density distribution in the lidar point cloud, and the multipath reflection path of the satellite signal; The cross-modal alignment module is used to input the raw observation values of the satellite positioning signal, the motion trajectory of the inertial measurement unit, the key feature points of the visual image, and the geometric structure of the lidar point cloud into the correction function bound to the physical property parameters of the environment, and output a standardized feature vector after environmental error compensation; The fusion positioning output module is used to input the standardized feature vector into a recursive neural network with spatiotemporal memory capability. The recursive neural network outputs the six-degree-of-freedom pose of the carrier by dynamically adjusting the fusion weight coefficients of each modal feature.
[0058] The above is a detailed description of an embodiment of the present invention. However, the content described is only a preferred embodiment of the present invention and should not be considered to limit the scope of the present invention. All equivalent changes and improvements made within the scope of the present invention should still fall within the scope of the patent coverage of the present invention.
Claims
1. A high-precision AI positioning method based on spatial multimodal data fusion, characterized by: The following steps are involved: S1. Collect spatial multimodal data, including satellite positioning signals, inertial measurement unit data, visual image sequences, lidar point clouds, and high-precision map data; S2, generates a set of physical property parameters of the current environment by real-time analysis of the illumination distribution in the visual image, the atmospheric particle density distribution in the lidar point cloud, and the multipath reflection path of the satellite signal; S3: Input the original observation value of the satellite positioning signal, the motion trajectory of the inertial measurement unit, the key feature points of the visual image, and the geometric structure of the lidar point cloud into the correction function bound to the physical property parameters of the environment, and output the standardized feature vector after environmental error compensation; S4. Input the normalized feature vector into a recursive neural network with spatiotemporal memory capability. The recursive neural network outputs the six-degree-of-freedom pose of the carrier by dynamically adjusting the fusion weight coefficients of each modal feature.
2. The high-precision AI positioning method based on spatial multimodal data fusion according to claim 1 is characterized in that: In S2, the process of generating the physical property parameter set is as follows: Extract scene illumination intensity gradient maps frame by frame from visual image sequences, and generate illumination stability quantitative indicators by calculating the rate of change of illumination intensity gradients between adjacent frames. Identify the three-dimensional spatial density distribution of atmospheric suspended particles from the lidar point cloud and construct an atmospheric transmittance attenuation function based on the inverse relationship between the point cloud reflection intensity and the transmission distance; Separate the propagation delay difference between the direct path and the multipath reflection path from the satellite positioning signal, and combine it with the building facade geometry data in the high-precision map to generate a multipath interference probability distribution map; The illumination stability quantitative index, the atmospheric transmittance attenuation function and the multipath interference probability distribution map are combined into an environmental physical property parameter set.
3. The high-precision AI positioning method based on spatial multimodal data fusion according to claim 1 is characterized in that: In S3, the process of outputting the normalized feature vector after environmental error compensation is as follows: For the original observation value of the satellite positioning signal, call the multipath interference probability distribution map, superimpose the building reflection weight on the signal propagation path, and reconstruct the solution parameters of the pseudo-range observation equation; For the motion trajectory of the inertial measurement unit, the air resistance compensation is performed on the motion acceleration according to the atmospheric transmittance attenuation function, and the kinematic constraints in the carrier dynamics equation are reconstructed; For key feature points of visual images, the feature point extraction threshold is dynamically adjusted based on the quantitative index of illumination stability. When the illumination gradient change rate exceeds the set threshold, the static object outline in the high-precision map is used as an auxiliary matching benchmark. For the geometric structure of the lidar point cloud, the refractive index of the point cloud coordinates is corrected by combining the atmospheric transmittance attenuation function, the three-dimensional spatial topological relationship is reconstructed, and the spatial consistency alignment of multimodal features is achieved.
4. The high-precision AI positioning method based on spatial multimodal data fusion according to claim 2 is characterized in that: In S4, the process of the recursive neural network dynamically adjusting the fusion weight coefficients of each modal feature is as follows: A gating mechanism based on a set of environmental physical property parameters is established. When the output value of the atmospheric transmittance attenuation function is lower than a critical threshold, the weight coefficient of the lidar point cloud feature is increased. When the proportion of strong interference areas in the multipath interference probability distribution map exceeds the set ratio, the weight coefficient of the satellite positioning signal feature is reduced; When the illumination stability quantitative index fluctuates continuously for more than a set number of times, the matching confidence between the visual image features and the static features on the HD map is used as the basis for weight allocation; The fusion weight coefficients of each modal feature are updated in real time through the normalized scalar value output by the gating mechanism and fed back to the computing unit of the recurrent neural network.
5. The high-precision AI positioning method based on spatial multimodal data fusion according to claim 4 is characterized in that: In S4, the spatiotemporal memory capability is specifically: The recursive neural network maintains a memory matrix of the carrier's motion state, which stores the six-degree-of-freedom pose at each historical moment and the corresponding set of environmental physical property parameters. When it is detected that the similarity between the current environmental physical attribute parameter set and a record in the historical memory matrix exceeds the set threshold, the fusion weight coefficient of the historical moment is called as the initialization benchmark; When the carrier enters an area without satellite signals, such as a tunnel or indoors, a motion trajectory prediction model is constructed based on the characteristics of the most recent valid satellite positioning signal in the memory matrix to maintain positioning continuity.
6. The high-precision AI positioning method based on spatial multimodal data fusion according to claim 1, characterized in that: In S4, the six-degree-of-freedom pose of the output carrier includes: When continuous strong light or thick fog is detected, the coupling association between visual image features and lidar point cloud features is locked by disabling cross-modal attention calculation; When the satellite signal interruption duration exceeds the maximum reliable operating time limit of the inertial measurement unit, the topological road network in the high-precision map is used as the motion constraint boundary to limit the carrier's position estimation range; After each pose update, the current environment physical property parameter set is bound to the six-degree-of-freedom pose data and stored in an anti-tampering cache to form a traceable positioning evidence chain.
7. The high-precision AI positioning method based on spatial multimodal data fusion according to claim 6, characterized in that: The data structure of the anti-tampering buffer area is: A blockchain-based storage structure indexed by timestamps, where each node contains a snapshot of the environment's physical property parameter set, a fusion weight coefficient set, and a six-degree-of-freedom pose; The nodes generate verification hash values through the carrier motion continuity equation. Any tampering of node data will cause the subsequent node hash chain to break. When the positioning system is restarted or the positioning anomaly is recovered, the recursive neural network is reinitialized from the nearest valid hash node, the corresponding environmental physical property parameter set is loaded, and the historical fusion weight coefficient configuration is inherited.
8. The high-precision AI positioning method based on spatial multimodal data fusion according to claim 1, characterized in that: In S1, the application of the high-precision map data in the spatial multimodal data includes: In said S2, the electromagnetic wave reflection coefficient of the building material in the high-precision map is extracted to verify the rationality of the satellite signal multipath reflection path; In S3, the lane line curvature in the high-precision map is coupled with the trajectory curvature calculated by the inertial measurement unit to generate a trajectory smoothness constraint function; In S4, the output six-degree-of-freedom pose is spatially consistent with the topological relationship of the static objects in the high-precision map. If the verification deviation exceeds the set tolerance, the recursive neural network weight reset is triggered.
9. The high-precision AI positioning method based on spatial multimodal data fusion according to claim 8, characterized in that: The process of the spatial consistency check is as follows: Extract the 3D coordinates of all static objects within a set radius centered on the output 6DOF pose from the HD map; Project the LiDAR point cloud into the three-dimensional coordinate space and calculate the shortest distance set between the point cloud and the surface of each feature model; When the standard deviation of the distance set exceeds the set multiple of the historical mean, it is judged as a posture abnormality. When the posture abnormality is triggered, the fusion weight coefficient of the recurrent neural network at the current moment is reset to the preset safety value, and the six-degree-of-freedom posture is recalculated.
10. A high-precision AI positioning system based on spatial multimodal data fusion, used to implement the high-precision AI positioning method based on spatial multimodal data fusion according to any one of claims 1 to 9, characterized in that: include: Spatial data acquisition module, used to collect spatial multimodal data, including satellite positioning signals, inertial measurement unit data, visual image sequences, lidar point clouds and high-precision map data; The environmental attribute analysis module is used to generate a set of physical attribute parameters of the current environment by real-time analysis of the illumination distribution in the visual image, the atmospheric particle density distribution in the lidar point cloud, and the multipath reflection path of the satellite signal; The cross-modal alignment module is used to input the raw observation values of the satellite positioning signal, the motion trajectory of the inertial measurement unit, the key feature points of the visual image, and the geometric structure of the lidar point cloud into the correction function bound to the physical property parameters of the environment, and output a standardized feature vector after environmental error compensation; The fusion positioning output module is used to input the standardized feature vector into a recursive neural network with spatiotemporal memory capability. The recursive neural network outputs the six-degree-of-freedom pose of the carrier by dynamically adjusting the fusion weight coefficients of each modal feature.
Citation Information
Patent Citations
Unmanned environment sensing accurate positioning system and method
CN119986743A
Marine physical data intelligent fusion and optimization processing method
CN120012027A
Multi-modal Transform-based UWB multi-sensor fusion positioning method and system
CN120101803A
Multi-sensor cross-scene dynamic preferential fusion positioning and mapping method
CN120333448A
Receiving positioning signals at different frequencies
GB201212592D0
Cited By
Map construction method and device, computer equipment and readable storage medium
CN121383997A
A map construction method and device, computer equipment and readable storage medium
CN121383997B
Map construction method and device, computer equipment and storage medium
CN121409215A
Complex environment multimode signal transceiving and infrared remote control integration method and device
CN121485836A
Multi-element coupling positioning method and system based on mountain road surface
CN122041857A