High-precision ai positioning method and system based on spatial multi-modal data fusion
By using multimodal data fusion and recurrent neural networks to dynamically adjust the fusion weights, the problems of limited positioning stability and environmental parameter fluctuations in satellite signal rejection scenarios are solved, achieving high-precision and robust positioning results.
Patent Information
- Application Number
- CN202510998413.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-07-21
AI Technical Summary
Existing fusion solutions suffer from limited long-term positioning stability due to over-reliance on inertial navigation in satellite signal rejection scenarios, and the static configuration of fusion weights cannot respond in real time to environmental parameter fluctuations, making the accuracy prone to fluctuations during the transition phase of complex environments.
By collecting multimodal data, the physical property parameters of visual images, LiDAR point clouds and satellite signals are analyzed in real time to generate standardized feature vectors after environmental error compensation. The fusion weight coefficients are dynamically adjusted using a recurrent neural network, and cross-modal feature alignment and environmental adaptive optimization are performed in combination with high-precision maps.
It achieves high-precision positioning in satellite signal rejection scenarios, reduces the cumulative impact of inertial navigation errors, responds in real time to environmental parameter fluctuations, and improves positioning stability and accuracy during complex environment transition phases.
Smart Images

Figure CN120491127B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to the field of fusion positioning technology, and in particular to a high-precision AI positioning method and system based on spatial multi-modal data fusion. BACKGROUND
[0002] As a core capability of intelligent mobile carriers, spatial positioning technology plays a fundamental role in the fields of automatic driving, unmanned vehicle navigation, and augmented reality. With the expansion of application scenarios to urban canyons, underground spaces, and harsh weather environments, traditional single-sensor positioning solutions face challenges in both precision and reliability. Multi-modal data fusion technology, which integrates satellite signals, inertial measurements, visual perception, and laser radar data sources, has become a mainstream development direction for improving positioning performance in complex environments.
[0003] In the prior art, sensor fusion schemes based on the Kalman filter framework have formed a relatively mature technical path. This scheme uses the absolute position information provided by the Global Navigation Satellite System (GNSS) as a reference, and combines the high-frequency pose calculation of the Inertial Measurement Unit (IMU) to achieve data complementation. Some improved schemes further introduce visual odometry or laser radar point cloud matching results, and by extending the filter state quantity, multi-source observations are included in a unified solution model. This type of technology can effectively improve the continuity of positioning in open environments by establishing a statistical model of sensor errors.
[0004] Although existing fusion schemes have made some progress, their technical implementation still has inherent constraints: in terms of environmental adaptability, the system relies too much on the cumulative error compensation mechanism of inertial navigation in satellite signal denial scenarios such as tunnels and indoor environments, and the long-term positioning stability is limited by the accuracy of the motion model, making it difficult to maintain continuous and reliable positioning output; in terms of dynamic adjustment capability, the fusion weight parameters are usually statically configured based on pre-set environmental classification, and cannot respond in real time to continuous fluctuations in physical environmental parameters such as sudden changes in light and atmospheric transmittance, leading to fluctuations in positioning accuracy during the transition period in complex environments. These factors result in the need for further improvement in the positioning accuracy and robustness of existing schemes in complex scenarios. SUMMARY
[0005] The purpose of the present application is to provide a high-precision AI positioning method and system based on spatial multi-modal data fusion, which solves the following technical problems:
[0006] The existing fusion scheme relies too much on inertial navigation in satellite signal denial scenarios, limiting long-term positioning stability, and static configuration of fusion weights cannot respond in real time to fluctuations in environmental parameters, making the accuracy prone to fluctuation during the transition period in complex environments.
[0007] The purpose of the present application can be achieved by the following technical solutions:
[0008] The high-precision AI positioning method based on spatial multi-modal data fusion comprises the following steps:
[0009] S1, collecting spatial multi-modal data, including satellite positioning signals, inertial measurement unit data, visual image sequences, laser radar point clouds and high-precision map data;
[0010] S2, generating a set of physical attribute parameters of the current environment by analyzing the light distribution in the visual image, the atmospheric particle density distribution in the laser radar point cloud and the multipath reflection path of the satellite signal in real time;
[0011] S3, inputting the original observation value of the satellite positioning signal, the motion trajectory of the inertial measurement unit, the key feature points of the visual image and the geometric structure of the laser radar point cloud into the correction function bound to the environmental physical attribute parameters respectively, and outputting the standardized feature vector compensated by the environment error;
[0012] S4, inputting the standardized feature vector into a recurrent neural network with space-time memory capability, and outputting the six-degree-of-freedom pose of the carrier by dynamically adjusting the fusion weight coefficient of each modal feature.
[0013] As a further scheme of the application, in S2, the generation process of the set of physical attribute parameters is:
[0014] Extract the scene light intensity gradient graph from the visual image sequence frame by frame, generate a light stability quantitative index by calculating the change rate of the light intensity gradient of adjacent frames;
[0015] Identify the three-dimensional spatial density distribution of atmospheric suspended particles from the laser radar point cloud, and construct an atmospheric transmittance attenuation function according to the inverse relationship between the point cloud reflection intensity and the transmission distance;
[0016] Separate the propagation time delay difference of the direct path and the multipath reflection path from the satellite positioning signal, and generate a multipath interference probability distribution map combined with the building facade geometric data in the high-precision map;
[0017] Combine the light stability quantitative index, the atmospheric transmittance attenuation function and the multipath interference probability distribution map into the set of environmental physical attribute parameters.
[0018] As a further scheme of the application, in S3, the process of outputting the standardized feature vector compensated by the environment error is:
[0019] For the original observation value of the satellite positioning signal, call the multipath interference probability distribution map, superimpose the building reflection weight on the signal propagation path, and reconstruct the solution parameters of the pseudo-range observation equation;
[0020] For the motion trajectory of the inertial measurement unit, the motion acceleration is compensated for air resistance according to the atmospheric transmittance decay function, and the kinematic constraint condition in the carrier dynamics equation is reconstructed;
[0021] For the key feature points of the visual image, the feature point extraction threshold is dynamically adjusted according to the light stability quantitative index, and when the light gradient change rate exceeds the set threshold, the static ground feature profile in the high-precision map is used as an auxiliary matching reference;
[0022] For the geometric structure of the laser radar point cloud, the point cloud coordinates are refractive index corrected in combination with the atmospheric transmittance decay function, the three-dimensional space topology is reconstructed, and the spatial consistency alignment of multi-modal features is realized.
[0023] As a further scheme of the application: in the S4, the process of dynamically adjusting the fusion weight coefficient of each modal feature by the recurrent neural network is:
[0024] A gating mechanism is established with the environmental physical property parameter set as the condition, and when the output value of the atmospheric transmittance decay function is lower than the critical threshold, the weight coefficient of the laser radar point cloud feature is increased;
[0025] When the strong interference area proportion in the multipath interference probability distribution diagram exceeds the set proportion, the weight coefficient of the satellite positioning signal feature is reduced;
[0026] When the light stability quantitative index fluctuates continuously more than a set number of times, the matching confidence of the visual image feature and the static ground feature of the high-precision map is used as the weight allocation basis;
[0027] The fusion weight coefficient of each modal feature is updated in real time by the normalized scalar value output by the gating mechanism and fed back to the calculation unit of the recurrent neural network.
[0028] As a further scheme of the application: in the S4, the spatio-temporal memory capability is specifically:
[0029] The recurrent neural network internally maintains a carrier motion state memory matrix, and the memory matrix stores the six-degree-of-freedom pose at the historical time and the corresponding environmental physical property parameter set;
[0030] When it is detected that the similarity between the current environmental physical property parameter set and a record in the historical memory matrix exceeds a set threshold, the fusion weight coefficient at the historical time is called as an initialization reference;
[0031] When the carrier enters a satellite signal-free area such as a tunnel or an indoor area, a motion trajectory prediction model is constructed based on the nearest effective satellite positioning signal feature in the memory matrix to maintain positioning continuity.
[0032] As a further scheme of the application: in the S4, the output carrier six-degree-of-freedom pose includes:
[0033] When persistent strong light or thick fog is detected, the coupling correlation of visual image features and lidar point cloud features is locked by disabling cross-modal attention calculation;
[0034] When the satellite signal interruption duration exceeds the maximum reliable working time limit of the inertial measurement unit, the topological road network in the high-precision map is enabled as a motion constraint boundary to limit the carrier pose calculation range;
[0035] After each pose update, the current set of environmental physical property parameters and the six-degree-of-freedom pose data are bound and stored in the tamper-resistant cache area to form a traceable positioning evidence chain.
[0036] As a further scheme of the application: the data structure of the tamper-resistant cache area is:
[0037] The blockchain storage structure is indexed by timestamp, and each node contains a snapshot of the set of environmental physical property parameters, a set of fusion weight coefficients, and a six-degree-of-freedom pose;
[0038] The verification hash value is generated between nodes through the carrier motion continuity equation, and any node data tampering will cause the hash chain to break;
[0039] When the positioning system restarts or the positioning anomaly recovers, the recursive neural network is reinitialized from the nearest valid hash node, the corresponding set of environmental physical property parameters is loaded, and the historical fusion weight coefficient configuration is inherited.
[0040] As a further scheme of the application: in S1, the application of high-precision map data in the spatial multi-modal data includes:
[0041] In S2, the electromagnetic wave reflection coefficient of the building material in the high-precision map is extracted to verify the rationality of the satellite signal multipath reflection path;
[0042] In S3, the lane line curvature in the high-precision map is coupled with the trajectory curvature calculated by the inertial measurement unit to generate a trajectory smoothness constraint function;
[0043] In S4, the output six-degree-of-freedom pose is spatially consistent with the topological relationship of static terrain features in the high-precision map, and if the verification deviation exceeds the set tolerance, the recursive neural network weight is reset.
[0044] As a further scheme of the application: the process of spatial consistency verification is:
[0045] All static terrain three-dimensional coordinates within a radius range centered on the output six-degree-of-freedom pose are extracted from the high-precision map;
[0046] Projecting the laser radar point cloud to the three-dimensional coordinate space, calculating the shortest distance set of the point cloud and the surface of each object model;
[0047] When the standard deviation of the distance set exceeds the set multiple of the historical mean, it is determined that the pose is abnormal, and when the pose is abnormal, the fusion weight coefficient of the recurrent neural network at the current time is reset to a preset safety value, and the six-degree-of-freedom pose is recalculated.
[0048] The application also includes a high-precision AI positioning system based on spatial multi-modal data fusion, which is used to implement the high-precision AI positioning method based on spatial multi-modal data fusion described above, comprising:
[0049] The spatial data acquisition module is used to acquire spatial multi-modal data, including satellite positioning signals, inertial measurement unit data, visual image sequences, laser radar point clouds and high-precision map data.
[0050] The environmental attribute analysis module is used to generate a set of physical attribute parameters of the current environment by analyzing the light distribution in the visual image, the atmospheric particle density distribution in the laser radar point cloud and the multipath reflection path of the satellite signal in real time.
[0051] The cross-modal alignment module is used to input the original observation value of the satellite positioning signal, the motion trajectory of the inertial measurement unit, the key feature points of the visual image and the geometric structure of the laser radar point cloud into the correction function bound to the environmental physical attribute parameters respectively, and output the standardized feature vectors after the environmental error compensation.
[0052] The fusion positioning output module is used to input the standardized feature vectors into the recurrent neural network with spatio-temporal memory capability, and the recurrent neural network outputs the six-degree-of-freedom pose of the carrier by dynamically adjusting the fusion weight coefficients of each modal feature.
[0053] The application has the following beneficial effects:
[0054] The application solves the problems of environment adaptability limitation and insufficient dynamic adjustment capability in the prior art by constructing a dynamic environment physical property model, performing cross-modal feature alignment under physical constraints, and generating an environment-adaptive fusion positioning result and other technical means. By real-time analysis of multi-modal data such as visual images, laser radar point clouds, satellite signals, and the like, an environment physical property parameter set is generated to realize active perception and modeling of the environment; through a correction function bound to the environment physical property parameters, environment error compensation is performed on the modal features to realize deep alignment of the cross-modal features; through the gating mechanism of a recurrent neural network, the fusion weight coefficient is dynamically adjusted based on the environment physical parameters to realize adaptive optimization of the positioning process. The application enables the positioning system to maintain high-precision positioning in a satellite signal denial scenario, effectively reduces the influence of inertial navigation error accumulation, can respond to fluctuations in environmental parameters such as light and atmospheric transmittance in real time, improves the positioning stability during the transition of complex environments, and significantly enhances the positioning accuracy and robustness of the system in extremely complex scenarios. BRIEF DESCRIPTION OF DRAWINGS
[0055] The application will be further described below with reference to the drawings.
[0056] Figure 1 is a flowchart of the high-precision AI positioning method based on spatial multi-modal data fusion of the application;
[0057] Figure 2 is a module schematic diagram of the high-precision AI positioning system based on spatial multi-modal data fusion of the application. DETAILED DESCRIPTION
[0058] The technical solutions in the embodiments of the application will be clearly and completely described below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only part of the embodiments of the application, rather than all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of the application.
[0059] Please refer to Figure 1 The application is a high-precision AI positioning method based on spatial multi-modal data fusion, which comprises the following steps:
[0060] S1, spatial multi-modal data is collected, including satellite positioning signals, motion data output by an inertial measurement unit, visual image sequences, point cloud data scanned by a laser radar, and environment prior information in a high-precision map. These data depict the spatial state of the carrier from different dimensions, the satellite positioning signal provides a basic position reference, the inertial data reflect real-time motion acceleration and angular velocity, the visual image captures scene texture and light features, the laser point cloud presents a three-dimensional geometric structure, and the high-precision map contains accurate information of static environments such as roads and buildings.
[0061] S2, generate a scene light stability index by real-time analyzing the light intensity distribution of pixels in the visual image; identify the spatial density distribution of atmospheric aerosols from the laser radar point cloud to construct an atmospheric transmittance attenuation function; separate the propagation time delay difference of the direct path and the multipath reflection path from the satellite positioning signal, combine the geometric data of the building facade in the high-precision map to generate a multipath interference probability distribution map, and combine these indexes, functions and distribution maps to form a set of physical property parameters of the current environment.
[0062] S3, call the multipath interference probability distribution map to the original observation value of the satellite positioning signal, superimpose the building reflection weight on the signal propagation path to reconstruct the solution parameters of the pseudorange observation equation; according to the atmospheric transmittance attenuation function, the air resistance compensation is carried out on the motion acceleration of the inertial measurement unit, and the kinematic constraint condition in the carrier dynamics equation is reconstructed; according to the light stability index, the extraction threshold of the key feature points of the visual image is dynamically adjusted, and when the light gradient change rate is large, the static ground feature profile in the high-precision map is used as an auxiliary matching reference; according to the atmospheric transmittance attenuation function, the refractive index correction is carried out on the point cloud coordinates of the laser radar point cloud, and the three-dimensional space topology relationship is reconstructed, and the standardized feature vector after environmental error compensation is output.
[0063] S4, input the standardized feature vector into the recurrent neural network with space-time memory capability, which dynamically adjusts the fusion weight coefficient of each modal feature through the gating mechanism conditioned on the set of environmental physical property parameters: when the output value of the atmospheric transmittance attenuation function is low, the weight of the laser radar point cloud feature is increased, when the strong interference area accounts for a high proportion in the multipath interference probability distribution map, the weight of the satellite positioning signal feature is reduced, and when the light stability index fluctuates many times in succession, the matching confidence between the visual image feature and the static ground feature of the high-precision map is used as the weight distribution basis. At the same time, the network maintains a carrier motion state memory matrix inside, stores the six-degree-of-freedom pose and corresponding environmental parameters at the historical time, and finally outputs the six-degree-of-freedom pose of the carrier.
[0064] In S2, the generation process of the set of physical property parameters is:
[0065] The scene light intensity gradient map is extracted from the visual image sequence frame by frame. For example, during the driving process, the vehicle successively passes through the morning tree-lined road, the noon open road, and the evening backlight road section. When passing through the tree-lined road, the sunlight passing through the leaves forms a mottled light and shadow, and the boundary of the bright and dark areas in the image forms obvious light intensity gradient; after driving into the open road, the overall light tends to be uniform, and the gradient value decreases; when driving in the evening backlight, the high-contrast profile of the scene in front of the vehicle is formed due to strong light irradiation, and the gradient value changes significantly again. By calculating the change rate of the light intensity gradient of the same region in adjacent frames in these scenes, a light stability quantitative index can be generated, which can accurately reflect the dynamic change of the light condition. For example, at the moment of backlight, the gradient change rate of adjacent frames rises sharply, and the index value increases, directly reflecting the instability of the light.
[0066] The three-dimensional spatial density distribution of atmospheric suspended particles is identified from the laser radar point cloud. In light rain weather, the laser beam emitted by the laser radar scatters with atmospheric suspended particles such as raindrops, causing the reflection intensity of part of the point cloud to be abnormal. By analyzing the distribution characteristics of the reflection intensity in the point cloud data, the density difference of atmospheric suspended particles in different regions can be identified. There may be more water vapor particles on both sides of the road due to plant transpiration, and the reflection intensity of the point cloud in the corresponding region attenuates more obviously. According to the inverse relationship between the reflection intensity and the transmission distance of the point cloud, that is, the farther away from the laser radar, the greater the degree of attenuation of the reflection intensity due to particle scattering, an atmospheric transmittance attenuation function is constructed, which can quantify the transmission loss of the laser signal at different distances, providing a basis for subsequent point cloud correction.
[0067] The propagation time delay difference between the direct path and the multipath reflection path is separated from the satellite positioning signal. When the vehicle drives into a mixed street block with both low-rise buildings and super high-rise buildings, the satellite signal propagation path is complex and diverse. Part of the signal directly reaches the receiver to form a direct path, and part of the signal is reflected by the inclined surface of the low-rise building and the vertical wall surface of the super high-rise building to reach the receiver, forming multiple multipath paths. After separating the propagation time difference of these paths by signal processing technology, combined with the facade geometry data of the buildings in the region in the high-precision map, such as the height of a high-rise building is 80 meters and the wall surface inclination is 90 degrees, and the height of another low-rise building is 10 meters and the wall surface inclination is 60 degrees, the source and intensity of each reflection path can be accurately judged, and a multipath interference probability distribution map is generated. In the surrounding area of the super high-rise building, the distribution map has a significantly higher interference probability in the corresponding area than in other areas due to the multiple reflection paths and strong signal.
[0068] The light stability quantitative index, the atmospheric transmittance attenuation function, and the multipath interference probability distribution map generated above are integrated to form a parameter set that can reflect the current environmental physical state in real time, providing accurate basis for error compensation of various modal data.
[0069] In the S3, the process of outputting the normalized feature vector compensated by the environment error is:
[0070] For the original observation value of the satellite positioning signal, the multipath interference probability distribution diagram is called for correction. In the periphery of a super high-rise building with high multipath interference probability, according to the interference probability value of the area in the distribution diagram, the building reflection weight is superimposed on the signal propagation path. The area with greater contribution of the reflected signal is adjusted more significantly in the weight of the pseudorange observation value, so as to reconstruct the solution parameter of the pseudorange observation equation and reduce the positioning deviation caused by the multipath effect.
[0071] For the motion trajectory of the inertial measurement unit, the air resistance compensation is performed on the motion acceleration according to the atmospheric transmittance attenuation function. In a light rain day with low atmospheric transmittance, the air density increases, resulting in an increase in the air resistance when the vehicle is driving, and the acceleration data recorded by the inertial measurement unit will contain the influence of the additional resistance. By calculating the resistance coefficient at different driving speeds through the attenuation function, the acceleration data is compensated, and then the kinematic constraint condition in the carrier dynamics equation is reconstructed, so that the trajectory calculation is more in line with the actual motion state.
[0072] For the key feature points of the visual image, the extraction threshold is dynamically adjusted according to the light stability quantization index. When the vehicle quickly drives from a shaded road into the noon sunlight, the light intensity rises sharply, causing the light gradient change rate of adjacent frames to exceed the set threshold. At this time, the feature point extraction threshold originally suitable for the shaded road is no longer applicable, and the threshold is automatically lowered to adapt to the strong light environment. At the same time, the precise positions of static ground objects such as road edges and lamp posts in the high-precision map are enabled as auxiliary matching references, ensuring that reliable visual feature points can be extracted stably even in the case of dramatic changes in light.
[0073] For the geometric structure of the laser radar point cloud, the refractive index correction is performed on the point cloud coordinates in combination with the atmospheric transmittance attenuation function. In foggy weather, the laser beam causes measurement distance deviation due to atmospheric particle refraction, and the farther the distance, the greater the deviation. The refractive index correction value at different distances is calculated using the attenuation function, and the three-dimensional coordinates of each point cloud are calibrated. For example, the corner point of a building originally displayed at a distance of 100 meters is closer to its actual position after correction. Through this correction, the three-dimensional spatial topological relationship is reconstructed, so that the geometric structure of the laser point cloud is accurately aligned with the ground feature in the visual image and the position information of the satellite positioning in space, and finally the normalized feature vector compensated by the environment error is output, providing high-quality feature input for subsequent fusion positioning.
[0074] In the S4, the process of dynamically adjusting the fusion weight coefficient of each modality feature by the recurrent neural network is:
[0075] In an urban canyon environment, when the vehicle drives into a high-rise building dense area, the multipath interference probability distribution map shows that the strong interference area accounts for more than a certain proportion. The recurrent neural network reduces the weight coefficient of the satellite positioning signal feature through the gating mechanism. For example, there are three super high-rise buildings around a certain intersection. After the satellite signal is reflected for many times, the interference value of this area in the multipath interference probability distribution map reaches 0.8 (more than the set proportion 0.6). At this time, the weight coefficient of the satellite positioning signal is reduced from the default 0.4 to 0.1 through the gating mechanism. At the same time, the laser radar point cloud data can effectively capture the geometric features of the building facade, and the atmospheric transmittance attenuation function output value is 0.9 (higher than the critical threshold 0.7), and its weight coefficient is increased from 0.3 to 0.5. Because the visual image has a continuous fluctuation of more than a certain number of times (such as 5 times) in the light stability quantization index, the system uses the static terrain profile (such as road edge, building corner) in the high-precision map as an auxiliary matching reference, and uses the matching reliability of the visual image and the high-precision map as the weight allocation basis. If the visual feature matches the high-precision map successfully, the weight coefficient of the visual image feature is increased from 0.2 to 0.3, and finally the normalized scalar value (0.1, 0.3, 0.5, 0.1) is output through the gating mechanism, which corresponds to the fusion weight of satellite positioning, visual image, laser radar and inertial measurement unit respectively.
[0076] When the vehicle enters the tunnel, the atmospheric transmittance attenuation function output value drops to 0.3 (lower than the critical threshold 0.7), and the recurrent neural network immediately increases the weight coefficient of the laser radar point cloud feature to 0.8. At the same time, because there is no satellite signal in the tunnel, the multipath interference probability distribution map is invalid, and the weight coefficient of the satellite positioning signal is reduced to 0. The visual image is dependent on the matching of the laser radar point cloud and the geometric features (such as tunnel wall, lane line) in the high-precision map inside the tunnel because the light stability quantization index fluctuates continuously more than a certain number of times in the tunnel. The weight coefficient of the visual image is reduced to 0.1. The inertial measurement unit data can provide a short-term reliable motion trajectory, and the weight coefficient is maintained at 0.1. The gating mechanism dynamically generates the normalized weight coefficient (0, 0.1, 0.8, 0.1) by analyzing the set of environmental physical property parameters in real time, ensuring that high-precision positioning results can still be obtained in the tunnel environment.
[0077] In the S4, the spatiotemporal memory capability is specifically:
[0078] The motion state memory matrix maintained inside the recurrent neural network records the historical pose of the vehicle under different environments and the corresponding set of environmental physical attribute parameters. When the vehicle first enters a tunnel from an open road, the memory matrix stores the satellite positioning signal features before entering the tunnel, the laser radar point cloud data, the six-degree-of-freedom pose, and the set of environmental physical attribute parameters. When the vehicle approaches the tunnel again, the system detects that the similarity between the current set of environmental physical attribute parameters (such as light stability, atmospheric transmittance, and multipath interference probability) and the tunnel entrance record in the memory matrix exceeds the set threshold (such as 0.9), and immediately calls the fusion weight coefficient (0, 0.1, 0.8, 0.1) of the historical time as the initialization reference, so that the network can quickly adapt to the tunnel environment and reduce the positioning error in the transition stage.
[0079] When the vehicle completely enters the tunnel causing the satellite signal to be interrupted, the recurrent neural network constructs a motion trajectory prediction model based on the nearest valid satellite positioning signal features in the memory matrix (i.e., the positioning data at the tunnel entrance) combined with the motion trajectory provided by the inertial measurement unit. This model uses the historical motion parameters such as acceleration and angular velocity stored in the memory matrix to predict the possible trajectory of the vehicle inside the tunnel. For example, if the memory matrix shows that the vehicle was driving straight at a speed of 50 km / h before entering the tunnel, and the steering wheel angle is 0 degrees, the model will predict that the vehicle will continue to maintain a straight motion state inside the tunnel, and will fine-tune the predicted trajectory based on the real-time updated acceleration data from the inertial measurement unit, thereby maintaining the continuity of positioning until the vehicle exits the tunnel and regains valid satellite signals.
[0080] In heavy rain weather, the atmospheric transmittance decay function output value continues to be below the critical threshold, and the quality of the laser radar point cloud data decreases. At this time, the recurrent neural network retrieves the positioning data under similar historical weather conditions from the memory matrix, and finds that when the atmospheric transmittance is below 0.5, the matching success rate of road markings in visual images and high-precision maps is higher. The system adjusts the fusion strategy accordingly, taking the matching results of visual images and high-precision maps as the main positioning basis, while reducing the weight coefficient of laser radar point cloud. This weight adjustment mechanism based on historical experience effectively improves the positioning robustness of the system in extreme weather.
[0081] When the vehicle exits the tunnel and enters the signal recovery area, the recurrent neural network detects that the satellite signal quality gradually improves and the multipath interference probability decreases. The system compares the current set of environmental physical property parameters with the historical records of the tunnel exit in the memory matrix. If the similarity exceeds the set threshold, the corresponding historical fusion weight coefficient is called as the initial value, and is dynamically adjusted in combination with the current real-time data. For example, the memory matrix shows that the typical weight coefficient at the tunnel exit is (0.3, 0.2, 0.4, 0.1). Based on this, the system fine-tunes it to (0.4, 0.1, 0.4, 0.1) according to the current actual signal quality, realizing smooth transition from the tunnel environment to the open environment.
[0082] In the urban viaduct scene, when the vehicle drives onto the viaduct, the three-dimensional topological information provided by the high-definition map shows that there is a large billboard above the current position, which may cause shielding of the satellite signal. The recurrent neural network combines similar scene data stored in the memory matrix to reduce the weight coefficient of the satellite positioning signal in advance and to strengthen the matching of the laser radar point cloud and the high-definition map. If similar billboard texture features are detected in the visual image, the system further verifies the similarity between the current environment and the historical scene, so as to more accurately adjust the fusion weight and improve the positioning accuracy.
[0083] The spatio-temporal memory capability is also reflected in long-distance driving. When the vehicle returns to a region that has been passed after a long time of driving, the recurrent neural network quickly identifies the region by comparing the current set of environmental physical property parameters with the historical records in the memory matrix, and calls the corresponding historical fusion weight coefficient and positioning experience. For example, when the vehicle first passes a mountainous road section, the memory matrix records the complex terrain features, signal conditions and the best fusion strategy of the road section. When the vehicle enters the road section again, the system directly reuses the historical experience without relearning, significantly improving the positioning efficiency and accuracy.
[0084] In the S4, the six-degree-of-freedom pose of the output carrier includes:
[0085] When detecting continuous strong light on a highway under direct sunlight at noon, the cross-modal attention calculation is disabled in the process of outputting the six-degree-of-freedom pose of the carrier, and the coupling association of the visual image features and the laser radar point cloud features is locked. For example, the direct sunlight causes overexposure in some areas of the visual image. The feature association originally relying on cross-modal attention adjustment may deviate due to high light interference. After locking the coupling association, the edge features of the road markings in the visual image and the corresponding ground elevation mutation features in the laser radar point cloud maintain fixed association logic, avoiding feature matching confusion caused by drastic changes in light, and ensuring the stability of the pose output.
[0086] When encountering thick fog weather, the laser radar point cloud has some noise points due to atmospheric scattering, and the visual image also has blurred features due to low visibility. When the environment physical property model detects this situation, the cross-modal attention calculation is also disabled to maintain the inherent coupling relationship between the visual and laser radar features. For example, the laser radar captures a point cloud cluster of a guardrail, which is always associated with the gray profile of the guardrail in the visual image. Even if there is noise in the data of a single sensor, the coupling relationship between the two is not dynamically adjusted to ensure the basic accuracy of the pose solution.
[0087] When the vehicle enters an ultra-long tunnel, the satellite signal is completely interrupted, and the interruption duration exceeds the maximum reliable working time limit of the inertial measurement unit. When outputting the six-degree-of-freedom pose, the topological road network in the high-precision map is enabled as a motion constraint boundary. Assuming that the IMU in the tunnel accumulates errors due to long-time calculation, the originally straight trajectory begins to deviate slightly. At this time, the topological road network in the high-precision map (such as the number of lanes, turning restrictions, and tunnel wall positions) will form a virtual boundary, limiting the pose calculation range within the physical space defined by the road network. For example, the actual tunnel is two-way and two-lane, and the topological road network will constrain the pose calculation result to not exceed the lane width range, preventing the trajectory from deviating from the actual road due to IMU error accumulation, and ensuring that the pose output always adheres to the real driving path.
[0088] After each pose update, the current set of environment physical property parameters and six-degree-of-freedom pose data are bound and stored in the tamper-resistant cache area, forming a traceable positioning evidence chain. For example, when the vehicle passes through an overpass at 3:00 pm, the cache area will record the environmental parameters such as the light stability quantization index, atmospheric transmittance decay function, and multipath interference probability distribution at that time, as well as the corresponding three-dimensional coordinates and attitude data, along with the fusion weight coefficients of each modal feature at that time.
[0089] The tamper-resistant cache area adopts a blockchain-like storage structure indexed by time stamps, and each node is concatenated in chronological order. For example, the node at 3:01 contains the environmental parameter snapshot, fusion weight set, and pose data at that time. When the node at 3:02 is generated, it will generate a verification hash value based on the continuity equation of the carrier motion, combined with the pose data of the previous node, and the two nodes are associated through the hash value. If someone tries to modify the pose data at 3:01, the hash values of all subsequent nodes will be broken due to verification failure, directly showing the traces of data tampering.
[0090] When the positioning system restarts due to sudden failure or needs to be restored due to positioning abnormalities in complex road sections, the nearest valid hash node is searched from the tamper-resistant cache area. For example, after the system restarts, the last node of the unbroken hash chain in the cache area is read first, which contains the set of environmental physical property parameters, the fusion weight coefficient and the six-degree-of-freedom pose at the moment before the failure. The recurrent neural network is reinitialized based on these data, loads the corresponding environmental parameters, and inherits the historical fusion weight configuration, so that the positioning process is smoothly continued from the breakpoint, avoiding errors caused by positioning data loss or reinitialization due to system interruption.
[0091] In long-distance transportation scenarios, the vehicle continuously travels for several hours, and the tamper-resistant cache area continuously records the positioning-related data at each time. If subsequent tracing of the positioning accuracy of a certain section of the road is required, the environmental parameters and pose data of the corresponding node can be retrieved by timestamp indexing, and the data is confirmed not to be tampered with by means of hash verification between nodes, and the positioning process of the period is completely restored, providing traceable basis for the reliability of the positioning result.
[0092] When the vehicle exits the tunnel and reacquires the satellite signal, the tunnel-in pose data and environmental parameters stored in the tamper-resistant cache area serve as a reference for subsequent positioning calibration. For example, the node data at the tunnel exit includes the matching features of the laser radar point cloud and the tunnel wall, and the cumulative error trend of the IMU. When these data are combined with the newly acquired satellite signal, the pose output can be more accurately corrected to ensure smooth transition without jumps from the tunnel environment to the open environment.
[0093] In S1, the application of high-precision map data in the spatial multi-modal data includes:
[0094] Application of high-precision map in S2: verify the rationality of satellite signal multipath reflection path.
[0095] When the vehicle travels to a street with both glass curtain wall buildings and concrete residential buildings, satellite signals form multipath signals through reflection on the surface of the buildings in addition to the direct path. At this time, the electromagnetic wave reflection coefficients of the building materials in this area are extracted from the high-precision map, i.e., the reflection coefficient of the glass curtain wall is high, and the reflection coefficient of the concrete wall surface is low. Among the separated multipath reflection paths, if a path is identified as coming from the glass curtain wall direction, its signal strength should match the high reflection coefficient characteristic; if a path points to a concrete building but shows a strong reflection signal, combined with the reflection coefficient data of the high-precision map, it can be judged that the path is an abnormal interference, and then the invalid multipath signal is removed to ensure that the remaining multipath reflection paths meet the electromagnetic wave reflection characteristics of the materials, providing a basis for generating an accurate multipath interference probability distribution map.
[0096] Application of high-precision map in S3: generate trajectory smoothness constraint function.
[0097] When the vehicle is driving on a continuous curve road segment on the highway, the high-definition map stores the accurate curvature data of the lane lines on this road segment, reflecting the smooth transition characteristics of road design. The inertial measurement unit calculates the real-time trajectory curvature based on motion sensor data. If the calculated trajectory curvature suddenly changes due to road bumps or sensor noise, it will deviate from the curvature of the lane lines in the high-definition map. At this time, the curvature of the lane lines in the high-definition map is coupled with the trajectory curvature of the inertial measurement unit to generate a trajectory smoothness constraint function. The function will adjust the trajectory of the inertial measurement unit according to the deviation between the two, so that the corrected trajectory curvature is more consistent with the smooth trend of the actual road design, avoiding trajectory distortion caused by instantaneous noise and providing more reliable kinematic constraints for subsequent generation of standardized feature vectors.
[0098] Application of high-definition map in S4: spatial consistency check and abnormal handling.
[0099] When the vehicle is driving on a city trunk road, there are fixed streetlights, traffic signal poles, and guardrails on both sides of the road, and the three-dimensional coordinates of these static objects are accurately recorded in the high-definition map. After outputting the six-degree-of-freedom pose, the spatial consistency check is performed as follows:
[0100] Extract all static object three-dimensional coordinates within a certain radius range centered on the current pose from the high-definition map, such as the precise position data of 10 streetlights and 3 sets of guardrails within 50 meters. Project the laser radar's real-time scanned point cloud onto the space where these three-dimensional coordinates are located, and calculate the shortest distance between each point cloud and the corresponding object model surface, for example, the model surface coordinates of a certain streetlight are (x1, y1, z1), and the coordinates of a certain point in the projected point cloud are (x2, y2, z2). The straight-line distance between the two is included in the distance set.
[0101] If the vehicle positioning is accurate, the laser radar point cloud should be highly consistent with the object model in the high-definition map, and the standard deviation of the distance set is small. If the output pose deviates due to multipath interference or sensor drift, the distance between the point cloud and the object model will generally increase, and the standard deviation of the distance set will also rise. When this standard deviation exceeds a certain multiple of the historical mean, it is determined that the pose is abnormal.
[0102] When the pose is abnormal, the fusion weight coefficient of the current time in the recurrent neural network is reset to a preset safety value, for example, the weight of the laser radar point cloud is reduced due to a false judgment, and is restored to the default reasonable range after resetting. At the same time, the six-degree-of-freedom pose is recalculated. When recalculating, the system will preferentially refer to the topological relationship of the static objects in the high-definition map, combine the matching results of the laser radar point cloud, and correct the pose deviation to ensure that the output pose is consistent with the actual road environment.
[0103] For example, in a certain section, due to the sudden strong light, the visual feature extraction is wrong, the output pose deviates to the direction of the guardrail, the spatial consistency check finds that the standard deviation of the distance set of the lidar point cloud and the guardrail model is far beyond the historical average, and after determining that it is abnormal, the fusion weight is reset, the weight of the lidar is restored to normal, and the recalculated pose returns to the road center, avoiding the positioning failure caused by error accumulation.
[0104] Referring to Figure 2 The application also includes a high-precision AI positioning system based on spatial multi-modal data fusion, which is used to implement the high-precision AI positioning method based on spatial multi-modal data fusion described above, and includes:
[0105] A spatial data acquisition module is configured to acquire spatial multi-modal data, including satellite positioning signals, inertial measurement unit data, visual image sequences, lidar point clouds, and high-definition map data.
[0106] An environmental attribute analysis module is configured to generate a set of physical attribute parameters of the current environment by analyzing the light distribution in the visual image, the atmospheric particle density distribution in the lidar point cloud, and the multipath reflection path of the satellite signal in real time.
[0107] A cross-modal alignment module is configured to input the original observation value of the satellite positioning signal, the motion trajectory of the inertial measurement unit, the key feature points of the visual image, and the geometric structure of the lidar point cloud into the correction function bound to the environmental physical attribute parameters, respectively, and output the standardized feature vectors after environmental error compensation.
[0108] A fusion positioning output module is configured to input the standardized feature vectors into a recurrent neural network with spatiotemporal memory capability, and the recurrent neural network outputs the six-degree-of-freedom pose of the carrier by dynamically adjusting the fusion weight coefficients of the features of each modality.
[0109] The above describes one embodiment of the application in detail, but the content described is only a preferred embodiment of the application and cannot be considered as limiting the scope of the implementation of the application. Any equivalent changes and improvements made in accordance with the scope of the application should still be included in the patent coverage of the application.
Claims
1. A high-precision AI positioning method based on spatial multimodal data fusion, characterized in that, Includes the following steps: S1. Collect spatial multimodal data, including satellite positioning signals, inertial measurement unit data, visual image sequences, lidar point clouds, and high-precision map data; S2. By analyzing the illumination distribution in visual images, the atmospheric particle density distribution in lidar point clouds, and the multipath reflection paths of satellite signals in real time, a set of physical property parameters of the current environment is generated. S3. Input the original observation values of the satellite positioning signal, the motion trajectory of the inertial measurement unit, the key feature points of the visual image, and the geometric structure of the lidar point cloud into the correction function bound to the environmental physical attribute parameters, and output the standardized feature vector after environmental error compensation. S4. The standardized feature vector is input into a recurrent neural network with spatiotemporal memory capability. The recurrent neural network outputs the six-degree-of-freedom pose of the carrier by dynamically adjusting the fusion weight coefficients of each modality feature. In step S2, the process of generating the set of physical property parameters is as follows: The scene illumination intensity gradient map is extracted frame by frame from the visual image sequence, and the illumination stability quantification index is generated by calculating the rate of change of the illumination intensity gradient between adjacent frames. The three-dimensional spatial density distribution of atmospheric suspended particles is identified from lidar point clouds, and an atmospheric transmittance attenuation function is constructed based on the inverse relationship between point cloud reflection intensity and transmission distance. The propagation delay difference between the direct path and the multipath reflection path is separated from the satellite positioning signal, and combined with the geometric data of building facades in the high-precision map, a multipath interference probability distribution map is generated. The aforementioned light stability quantification index, atmospheric transmittance attenuation function, and multipath interference probability distribution map are combined into a set of environmental physical property parameters. In step S3, the process of outputting the standardized feature vector after environmental error compensation is as follows: For the raw observation values of satellite positioning signals, the multipath interference probability distribution map is called, and the building reflection weights are superimposed on the signal propagation path to reconstruct the solution parameters of the pseudorange observation equation; For the motion trajectory of the inertial measurement unit, the air resistance compensation of the motion acceleration is performed based on the atmospheric transmittance attenuation function, and the kinematic constraints in the carrier dynamics equation are reconstructed. For key feature points in visual images, the feature point extraction threshold is dynamically adjusted according to the illumination stability quantification index. When the illumination gradient change rate exceeds the set threshold, the static land feature outline in the high-precision map is used as an auxiliary matching benchmark. For the geometric structure of lidar point clouds, the refractive index of the point cloud coordinates is corrected by combining the atmospheric transmittance attenuation function, and the three-dimensional spatial topological relationship is reconstructed to achieve spatial consistency alignment of multimodal features.
2. The high-precision AI positioning method based on spatial multimodal data fusion according to claim 1, characterized in that, In step S4, the process by which the recurrent neural network dynamically adjusts the fusion weight coefficients of each modality feature is as follows: Establish a gating mechanism based on a set of environmental physical property parameters. When the output value of the atmospheric transmittance attenuation function is lower than the critical threshold, increase the weight coefficient of the lidar point cloud features. When the proportion of strong interference areas in the multipath interference probability distribution map exceeds a set ratio, the weighting coefficient of the satellite positioning signal features is reduced. When the quantitative index of illumination stability fluctuates continuously beyond a set number of times, the confidence level of the matching between visual image features and static ground features on high-precision maps will be used as the basis for weight allocation. The fusion weight coefficients of each modality feature are updated in real time through the normalized scalar values output by the gating mechanism and fed back to the computation unit of the recurrent neural network.
3. The high-precision AI positioning method based on spatial multimodal data fusion according to claim 2, characterized in that, In S4, the spatiotemporal memory capability specifically refers to: The recurrent neural network internally maintains a memory matrix for the motion state of the carrier. The memory matrix stores the six-degree-of-freedom pose and the corresponding set of environmental physical attribute parameters at historical moments. When the similarity between the current set of physical attribute parameters and a record in the historical memory matrix exceeds a set threshold, the fusion weight coefficient of that historical moment is used as the initialization benchmark. When the vehicle enters an area without satellite signal, a motion trajectory prediction model is constructed based on the characteristics of the nearest valid satellite positioning signal in the memory matrix to maintain positioning continuity.
4. The high-precision AI positioning method based on spatial multimodal data fusion according to claim 1, characterized in that, In step S4, the six-degree-of-freedom pose of the output carrier includes: When continuous strong light or dense fog is detected, the coupling relationship between visual image features and lidar point cloud features is locked by disabling cross-modal attention computation. When the satellite signal interruption duration exceeds the maximum reliable operating time limit of the inertial measurement unit, the topological road network in the high-precision map is used as the motion constraint boundary to limit the range of carrier pose estimation. After each pose update, the current set of physical attribute parameters of the environment is bound to the six-degree-of-freedom pose data and stored in an anti-tampering cache to form a traceable chain of positioning evidence.
5. The high-precision AI positioning method based on spatial multimodal data fusion according to claim 4, characterized in that, The data structure of the tamper-resistant cache is as follows: A blockchain-style storage structure indexed by timestamps, where each node contains a snapshot of the set of environmental physical attribute parameters, a set of fused weight coefficients, and a six-degree-of-freedom pose; The nodes generate verification hash values through the continuity equation of carrier motion. Any tampering with the data of any node will cause the hash chain of subsequent nodes to break. When the positioning system restarts or the positioning anomaly is resolved, the recurrent neural network is reinitialized from the most recent valid hash node, the corresponding set of environmental physical attribute parameters is loaded, and the historical fusion weight coefficient configuration is inherited.
6. The high-precision AI positioning method based on spatial multimodal data fusion according to claim 1, characterized in that, In step S1, the application of high-precision map data in the spatial multimodal data includes: In step S2, the electromagnetic wave reflection coefficients of building materials within the high-precision map are extracted to verify the rationality of the satellite signal multipath reflection path. In step S3, the lane curvature in the high-precision map is coupled with the trajectory curvature calculated by the inertial measurement unit to generate a trajectory smoothness constraint function. In step S4, the spatial consistency of the output six-degree-of-freedom pose with the topological relationship of static features in the high-precision map is checked. If the check deviation exceeds the set tolerance, the recursive neural network weights are reset.
7. The high-precision AI positioning method based on spatial multimodal data fusion according to claim 6, characterized in that, The process of spatial consistency verification is as follows: Extract the 3D coordinates of all static features within a radius defined by the output six-DOF pose from the high-precision map; Project the lidar point cloud onto this three-dimensional coordinate space and calculate the set of shortest distances between the point cloud and the surfaces of various landform models. When the standard deviation of the distance set exceeds a set multiple of the historical mean, it is judged as a pose abnormality. When a pose abnormality is triggered, the fusion weight coefficients in the recurrent neural network at the current moment are reset to the preset safe value, and the six-degree-of-freedom pose is recalculated.
8. A high-precision AI positioning system based on spatial multimodal data fusion, used to implement the high-precision AI positioning method based on spatial multimodal data fusion as described in any one of claims 1-7, characterized in that, include: The space data acquisition module is used to collect multimodal space data, including satellite positioning signals, inertial measurement unit data, visual image sequences, lidar point clouds, and high-precision map data. The environmental attribute analysis module is used to generate a set of physical attribute parameters of the current environment by analyzing the illumination distribution in visual images, the atmospheric particle density distribution in lidar point clouds, and the multipath reflection paths of satellite signals in real time. The cross-modal alignment module is used to input the raw observation values of satellite positioning signals, the motion trajectory of inertial measurement units, key feature points of visual images, and the geometric structure of lidar point clouds into correction functions bound to environmental physical property parameters, and output standardized feature vectors after environmental error compensation. The fusion positioning output module is used to input standardized feature vectors into a recurrent neural network with spatiotemporal memory capability. The recurrent neural network outputs the six-degree-of-freedom pose of the carrier by dynamically adjusting the fusion weight coefficients of each modality feature.
Citation Information
Patent Citations
Multi-sensor cross-scene dynamic preferential fusion positioning and mapping method
CN120333448A