A train localization method and system with multimodal fusion and semantic enhancement

By combining spatiotemporal alignment and semantic enhancement of multimodal sensor data with track structure constraints and dynamically adjusting the factor graph optimization weights, the accuracy and stability issues of the train positioning system in GNSS degradation environments were solved, and high-precision positioning of high-speed trains in complex environments was achieved.

CN121573041BActive Publication Date: 2026-04-03TONGJI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-22
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing train positioning systems suffer from reduced positioning accuracy in GNSS degradation environments, insufficient fusion of multimodal sensors, lack of dynamic adjustment mechanisms, inability to identify degradation states, failure to utilize prior track structure information, and lack of semantic understanding capabilities, leading to system failure or reduced accuracy in complex environments.

Method used

Dense semantic point clouds are constructed by spatiotemporal alignment of multimodal sensor data, sensor confidence weights are dynamically estimated, semantic enhancement and structural constraints are introduced, a comprehensive degradation scoring function is constructed, factor graph optimization weights are dynamically adjusted, and positioning optimization is performed by combining orbital geometry and semantic information.

Benefits of technology

It achieves high-precision and reliable train positioning in GNSS degradation environments, improves the system's adaptability and positioning stability in complex scenarios, avoids trajectory drift, and enhances perception redundancy and semantic expression capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121573041B_ABST
    Figure CN121573041B_ABST
Patent Text Reader

Abstract

This invention provides a train positioning method and system based on multimodal fusion and semantic enhancement, belonging to the field of rail transit technology. The method includes obtaining a dense point cloud from spatiotemporally aligned data and constructing a dense semantic point cloud; dynamically estimating the confidence weight of each type of sensor; constructing planar ensemble residuals and Manhattan structure constraint residuals to obtain lidar point cloud parameters, while simultaneously constructing visual projection constraints; constructing a comprehensive degradation scoring function to determine degradation in the current environment; when the degradation result is positive, introducing structural and motion information independent of the external environment to maintain trajectory estimation and constructing compensation constraints; introducing prior semantic constraints and large model semantic factor constraints; constructing a global optimization objective function, dynamically adjusting the weights of each modal factor to obtain the optimal estimation state, and outputting high-precision train positioning results. This invention achieves high-precision and robust trajectory estimation in extreme scenarios, thereby ensuring the continuity, safety, and intelligence of train positioning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of rail transit technology, and in particular relates to a train positioning method and system with multimodal fusion and semantic enhancement. Background Technology

[0002] With the rapid development of high-speed railway systems, the requirements for train operation safety, automation, and intelligent dispatching capabilities are increasing. High-precision train positioning, as a core supporting technology for ensuring safe train operation and dispatching, directly impacts the overall effectiveness of the vehicle-ground cooperative system. Currently, the Global Navigation Satellite System (GNSS) remains a widely used positioning method in railway scenarios. However, in GNSS degradation environments such as tunnels, underground sections, elevated structures, and complex weather conditions (e.g., rain and fog), it suffers from critical issues such as signal unavailability, increased positioning errors, and system malfunction, severely restricting the continuous positioning capabilities of high-speed trains across all scenarios.

[0003] To address this, academia and industry have successively proposed localization methods based on the fusion of multiple sensors, including inertial measurement units (IMUs), odometry, vision, and lidar. For example, the lidar-based simultaneous localization and mapping (SLAM) method can achieve high-precision localization in well-structured environments; while visual SLAM has good semantic recognition capabilities and cost advantages.

[0004] Although existing research has proposed various positioning and navigation methods for railway environments, especially fusion methods such as lidar SLAM, visual inertial navigation (VIO), IMU pre-integration, and wheel speed meter assistance, these methods still have the following key problems and shortcomings when dealing with complex scenarios such as GNSS degraded environments, structurally repetitive areas, and extreme weather:

[0005] 1. Insufficient adaptability to degraded environments. Most existing systems rely on stable input from GNSS or lidar / vision. However, in degraded environments such as tunnels, rain, fog, obstruction, or nighttime, problems such as GNSS signal loss, sparse lidar point clouds, and blurred images frequently occur, leading to positioning system failure or a sharp drop in accuracy. There is a lack of effective perception quality assessment and compensation mechanisms.

[0006] 2. Multimodal sensor fusion is insufficient and lacks a dynamic weight adjustment mechanism. Existing fusion methods mostly use fixed weights or simple weighted averages to process observation data from various sensors. They fail to dynamically adjust the data based on the current observation quality of different sensors (such as noise, outliers, and occlusion), which leads to unreliable information amplifying errors in degraded environments.

[0007] 3. Inability to effectively identify and respond to perception degradation states. Most current systems rely solely on the increase in residuals of perception factors to judge system stability, lacking a structured assessment of perception quality (such as point cloud structure entropy, image edge density, etc.). They cannot identify degradation areas in a timely manner or actively adjust perception weights, which can easily lead to mismatches and trajectory drift.

[0008] 4. Insufficient utilization of prior information about track structure. Most existing SLAM systems are general designs that do not take into account railway-specific constraints such as track centerline, turnout structure, and speed-limited sections. They lack strong geometric constraints on train movement space, which greatly reduces the system's self-stabilization capability in the absence of perception.

[0009] 5. The lack of a semantic-level multi-source reasoning mechanism means that most systems still rely primarily on geometric features and cannot incorporate higher-order semantic factors such as image semantics, tunnel labels, and traffic signals. They also lack the ability to model the linkage between "scene-behavior-trajectory," which limits the intelligent performance of the model in scenarios with incomplete or ambiguous information.

[0010] 6. Lacking structural priors and semantic understanding capabilities, traditional methods mostly optimize trajectories based solely on geometric features (such as point-to-point and point-to-surface errors), lacking explicit modeling of high-speed rail-specific structures (such as track centerlines, tunnels, and platforms) and operational semantics (such as speed limits, red lights, and switches), and thus unable to handle scenarios with structural ambiguity or semantic ambiguity.

[0011] Therefore, there is an urgent need to propose a unified factor graph optimization localization method for the complex operating environment of high-speed rail, which can realize spatiotemporal alignment of multimodal sensing information, degradation perception recognition, adaptive compensation, semantic enhancement and structural constraints, so as to ensure the high-precision, high-reliability and high-safety operation of the system under various extreme scenarios. Summary of the Invention

[0012] The purpose of this invention is to provide a train localization method with multimodal fusion and semantic enhancement, comprising the following steps:

[0013] S1: Spatiotemporal alignment of data collected by multimodal sensors is performed to obtain dense point clouds, and semantic enhancement is used to construct dense semantic point clouds to dynamically estimate the confidence weight of each type of sensor.

[0014] S2: Construct geometric residuals and Manhattan structure constraint residuals from points to planes based on dense semantic point clouds to obtain lidar point cloud constraints. At the same time, construct visual projection constraints based on visual point projection residuals and visual line projection residuals.

[0015] S3: Based on the joint semantic information of dense semantic point cloud and image, identify whether there is perception degradation in the current environment of high-speed rail, and construct a comprehensive degradation scoring function based on multiple indicators to determine degradation.

[0016] S4: After the degradation judgment result is that the degradation state has been entered, the system automatically switches to the redundant compensation factor modeling mechanism, introduces structural and motion information that does not depend on the external environment, maintains trajectory estimation in the degradation scenario, and constructs compensation constraints based on track structure geometric factors, IMU pre-integration factors and wheel and axle speed factors.

[0017] S5: Introduce prior semantic constraints and large model semantic factor constraints to ensure the rationality of the trajectory;

[0018] S6: Under the factor graph optimization framework, a global optimization objective function is constructed by integrating multiple constraints. Based on the confidence of various sensors, the weights of each modal factor in the factor graph optimization are dynamically adjusted. The optimal state estimate is obtained by iteratively solving the global optimization objective function, and finally, a high-precision train positioning result is output.

[0019] Furthermore, in S1, the multimodal sensors include lidar, millimeter-wave radar, camera, IMU, and wheel speedometer. The construction of the dense semantic point cloud includes the following steps:

[0020] S111: Align all types of sensors to the IMU reference time;

[0021] S112: Alignment of millimeter-wave radar point cloud with lidar coordinate system is achieved through extrinsic parameter transformation, and spatial completion is performed using kernel function weighted interpolation method to generate dense point cloud to enhance geometric integrity;

[0022] S113: Generate a semantic label map using a semantic segmentation network for the current frame image, and project the LiDAR points onto the image plane to achieve alignment between semantic information and three-dimensional structure;

[0023] S114: Project the dense point cloud onto the image plane through extrinsic parameter transformation, query the semantic category of the corresponding pixel in the semantic label map, and fill the corresponding 3D point with the semantic label to finally generate a dense semantic point cloud.

[0024] Furthermore, in S1, dynamically estimating the confidence levels of various sensors includes the following steps:

[0025] S121: The deviation of the observations obtained from various sensors is expressed as:

[0026] ;

[0027] in, To calculate the deviation for each observation; For at any time No. Observations from sensor-like devices; These are historical averages or expected observations, estimated from normal operating conditions.

[0028] S122: Construct a Bayesian model using a series of historical observations to estimate the observation variance of this type of sensor, expressed as:

[0029] ;

[0030] in, Indicates the first Sensor-like devices at all times The observation variance reflects its uncertainty; The expectation operator is represented, implemented through a sliding window or Bayesian posterior estimation.

[0031] S123: Observational variance obtained from estimation Calculate the confidence weights of the sensors The calculation formula is expressed as:

[0032] ;

[0033] in, For the first Sensor-like devices at all times Confidence weights To prevent division by zero of constants.

[0034] Furthermore, in S2, the process of constructing visual projection constraints includes the following steps:

[0035] S21: Constructing the visual point projection residual:

[0036] S21: Constructing the visual point projection residual:

[0037] S211: Select the set of key points for lidar scanning and the pixel observation points in the current image frame ,in, , As the key point, ; , representing the pixel coordinates of the corresponding projection of a 3D point in the image;

[0038] S212: Define the camera pose to be estimated in the current frame, i.e., the rotation matrix. With translation vector Use this pose to locate the key points in the world coordinate system. Transform to the camera coordinate system, and then project onto the image plane using the camera intrinsic parameters to obtain the predicted pixel coordinates.

[0039] S213: Define pixel projection error residual ,

[0040] in, Let be the Euclidean distance between the pixel position of a point on a 3D map after projection and its actual observation position, then the visual point projection residual is... The expression is:

[0041] ;

[0042] S22: Constructing the visual line projection residual:

[0043] S221: Let the two endpoints of the three-dimensional line segment be... and Based on the current camera pose and projection model, two points can be mapped onto the image plane to obtain their corresponding pixel coordinates. and .

[0044] S222: Define the set of line segments for image observation: ,in, These are the start and end points of the line segment in pixel space, respectively;

[0045] S223: Define the current camera pose, i.e., the rotation matrix. Translation vector The camera intrinsic parameter matrix is Project the start and end points of the 3D line segment onto the image plane to obtain the predicted pixel coordinates.

[0046] Then the visual line projection residual The expression is:

[0047] ;

[0048] in, For the index of the observed line segment; and These are the predicted pixel coordinates and the actual observed pixel coordinates of the starting point of the k-th observed line segment, respectively. and These are the predicted pixel coordinates and the actual observed pixel coordinates of the endpoint of the k-th observed line segment, respectively.

[0049] S23: Visual Projection Constraints Represented as: + .

[0050] Furthermore, S3 includes the following steps:

[0051] S31: Point cloud structure degradation identification:

[0052] S311: Using the point cloud normal entropy and structural dimension index of the dense point cloud in step 1, project the normal vectors in the dense point cloud onto a unit sphere, and construct multiple direction buckets by uniformly dividing the sphere or by predefined direction sets.

[0053] S312: Statistically analyze the number of normal vectors in each bucket and normalize them to a probability distribution. Calculate the point cloud normal entropy using the following formula:

[0054] ;

[0055] in, The normal entropy of the point cloud is used to measure the amount of information in the directional distribution.

[0056] S32: Degradation Index (DDI)

[0057] The structure dimension index (DDI) based on the eigenvalues ​​of the Hessian matrix is ​​used to evaluate the three-dimensional distribution integrity of point clouds, thereby determining whether there is spatial structure degradation in the point cloud of the current frame.

[0058] S33: Evaluating image texture quality based on edge density;

[0059] Canny edge detection is used to extract edge pixels from the image information captured by the camera, the number of edge pixels is counted, and the image edge density is calculated. ,definition This represents a threshold for determining whether an image quality meets the required standard; when If so, the image is deemed invalid;

[0060] S34: Calculate the Comprehensive Degradation Score Function (CDS), expressed as:

[0061] ;

[0062] in, and These represent the time intervals of the LiDAR and the camera, respectively. Confidence weights; These are all task-related weighting coefficients used to adjust the contribution of each modality to the degradation score, and =1; and They represent and The maximum value is used for normalization operations;

[0063] S35: Degradation is determined based on the calculated CDS (Comprehensive Degradation Score) value.

[0064] When CDS≤0.3, it is determined to be a normal scene, and the LiDAR point cloud factor, visual projection factor, compensation factor, prior semantic factor and large model semantic factor are enabled.

[0065] when If so, it is judged as moderate degradation, and the weights of visual projection and lidar point cloud factors are reduced;

[0066] When CDS If the condition is found to be severely degraded, all sensing factors are disabled, and the track + IMU + wheel speed compensation mechanism is enabled, while only the compensation constraint factor is retained.

[0067] Furthermore, S4 specifically includes the following steps:

[0068] S41: Constructing orbital structure geometry factors:

[0069] The track structure geometric factor input is a track centerline model generated based on map, B-spline, or rule-based fitting methods. , is represented as:

[0070] ;

[0071] ;

[0072] in, The distance from the current pose point to the orbit centerline Lateral deviation; The estimated position at the current moment; for The arc length corresponding to the closest position on the line;

[0073] S42: Constructing the IMU pre-integration factor;

[0074] The input to the IMU pre-integration factor provides brief position and acceleration information from the time-synchronized IMU data from step 1, and the current keyframe. With the next keyframe The state estimation is performed using the IMU, which provides high-frequency acceleration and angular velocity. By employing IMU pre-integration technology and connecting the keyframe states, the difference between the state changes observed by the IMU and the trajectory estimates is constructed as a residual. These residuals are minimized to form an inertial constraint on the trajectory, expressed as:

[0075] ;

[0076] in, The residual is constructed from the difference between the state changes observed by the IMU and the trajectory estimates.

[0077] and These are the pose and rotation matrices for the current frame and the next frame, respectively. and These represent the velocity vectors of the current frame and the next frame, respectively. and These are the position coordinates of the current frame and the next frame, respectively; These are rotation, velocity, and position predicted by IMU integration, respectively, from the IMU. It is the vector of gravitational acceleration; This refers to the keyframe interval time. Let be the covariance matrix, representing the uncertainty;

[0078] Step 4.3. Construct the wheel and axle speed factor;

[0079] The input information for the wheel axle velocity factor is: the velocity vector estimated from the current frame state. Current orbital direction unit vector That is, the tangent direction of the track curve at that point, and the scalar value of the wheel and axle velocity observed by the encoder. Define the residual function for:

[0080] ;

[0081] in, The actual speed of the train in the direction of track movement; based on the principle of consistency of speed direction, the component of the estimated speed in the track direction is aligned with the wheel speed gauge value;

[0082] S44: Graph optimization ensemble for compensation modeling:

[0083] The track geometry factor, IMU pre-integration factor, and wheel-axle velocity observation factor are uniformly incorporated into the graph optimization framework as compensation constraints between keyframe nodes, and an overall optimization objective function is constructed, expressed as:

[0084] = ;

[0085] in, This represents the set of state variables for all keyframes. The state variables for each keyframe include the pose, velocity, and IMU bias for each frame. For each type of factor, the dynamic weight coefficients are... For the first The factor in the th Residual function under class constraints , For the geometric residuals of the track structure, For IMU pre-integration residuals, The residual for observing wheel and axle speed.

[0086] Furthermore, in S5, the expression for the prior semantic constraints is:

[0087] ;

[0088] in, This is the speed limiting factor; For signal factor; Elevation factor; This is the turnout switching factor;

[0089] Speed ​​limiting factor Constrained trajectory estimation velocity Do not exceed the speed limit of the current track section Its error term is defined as:

[0090] ;

[0091] Signal factor Based on the signal status set on the track, when the corresponding signal light is red, the estimated speed for that frame is forcibly constrained. Approaching zero is defined as:

[0092] ;

[0093] in, This factor is activated when the traffic light at the current location is red.

[0094] Elevation factor The residual is constructed by comparing the deviation between the estimated position and the corresponding elevation on the trajectory curve, and is expressed as:

[0095] ;

[0096] in, To estimate the height component of the location; For the track in position The elevation value at the location;

[0097] Turnout switching factor When approaching a turnout, the track function is changed to a branch track to maintain continuity, as shown below:

[0098] ;

[0099] in, Let be the position vector of the orbital centerline in three-dimensional space; Let be the position vector of the centerline of the branch track in three-dimensional space.

[0100] Furthermore, in S5, the construction of semantic factor constraints for large models includes the following steps:

[0101] S51: Construct a semantic cue word set:

[0102] A semantic prompting word library is constructed based on the high-speed rail scenario. Each prompting word is encoded into a high-dimensional semantic vector through a large model, represented as follows:

[0103] ;

[0104] in, Indicates the first The semantic embedding vectors of each prompt word; A text encoder for large models; This indicates all the prompt words contained in the prompt word library.

[0105] S52: Image semantic embedding extraction: Extract the semantic embedding vector from each frame of the image in S1, represented as:

[0106] ;

[0107] in, For the first Frame image; This is the semantic vector of the current image frame; For large models The visual encoder;

[0108] S53: Calculate semantic residuals:

[0109] S531: Match the optimal semantic vector of the image and the prompt word using cosine similarity.

[0110] S532: Retrieve the most matching prompt word from the prompt word library. The formula is expressed as follows:

[0111] ;

[0112] S533: Define the semantic residual of the current frame as:

[0113] ;

[0114] in, The semantic residual term in the k-th frame represents the difference between the semantics of the current image and the best prompt word. Indicates and The closest prompt word vector;

[0115] S54: Semantic Factor Constraints for Large Models Represented as:

[0116] ;

[0117] in, These are the dynamic weighting coefficients for semantic factors, dynamically determined based on the image sharpness and camera confidence of the current frame. This represents the weighted semantic residual term.

[0118] Furthermore, in S6, the global optimization objective function is expressed as:

[0119] ;

[0120] in, ; The pose of each frame; Position for each frame; For frame rate; Optimize the number of keyframes within the sliding window for the factor graph; , , , and These are constraints for multiple categories, namely lidar point cloud, visual projection, compensation, prior semantics, and large model semantic factors. and These are the confidence weights for LiDAR and vision, respectively. The weights of the compensation factors; and These are the weights of prior semantics and the semantics of the large model, respectively.

[0121] The specific weights of each modal factor in the dynamically adjusted factor graph optimization are as follows: Based on the comprehensive degradation score (CDS), the environmental state is divided into three levels and mapped to a weight adjustment strategy, including:

[0122] In normal scenarios, i.e., when CDS < 0.3: LiDAR and visual factors dominate. , Compensation factor assistance: Semantic factors enabled: , ;

[0123] In mild degradation scenarios, i.e., when 0.3 ≤ CDS < 0.6: Reduce laser / visual weights: , Enhanced compensation factor: Semantic factors should be used with caution. , ;

[0124] In severely degraded scenarios, i.e., when CDS ≥ 0.6, unreliable modes are disabled: , Completely dependent on compensation factors: Semantic factors disabled: , .

[0125] A multimodal fusion and semantic enhancement positioning system for the complex environment of high-speed rail is provided. The system includes a multimodal sensor data acquisition and fusion module, a radar point cloud and visual projection error construction module, a degradation environment perception and judgment module, a positioning compensation modeling module under degradation state, a priori semantic constraint and large model semantic constraint modeling module, and a factor graph optimization and state estimation module. Each module corresponds to S1-S6 in sequence. The multimodal fusion and semantic enhancement positioning method completes S1-S6 through each module and finally outputs high-precision train positioning results.

[0126] Compared with the prior art, the beneficial effects of the present invention are mainly reflected in:

[0127] 1. The method of the present invention extends the traditional laser-IMU-vision fusion into a data input system with semantic enhancement by introducing image semantic segmentation, CLIP coding and other means. This makes the perception results not only geometrically complete, but also semantically interpretable, improving the redundancy and usability of perception, and enhancing the structural and semantic expression capabilities of the perception data.

[0128] 2. The method of this invention constructs a degradation environment recognition method based on indicators such as normal entropy, Hessian matrix eigenvalues, and image edge density, and integrates them into a unified degradation scoring mechanism. It can identify scenarios with degraded perception quality in real time, and dynamically adjust the weight distribution of multimodal factors in the optimization system accordingly, thereby improving the system's adaptive capability and realizing degradation recognition and adaptive perception scheduling in complex scenarios.

[0129] 3. The method of the present invention introduces structured prior information such as orbital geometry constraints, elevation factors, and wheel speed constraints, which can maintain stable positioning even when GNSS is unavailable or laser or visual quality is degraded, avoiding trajectory drift or estimation interruption, and improving the system's positioning continuity and robustness in perception degradation scenarios.

[0130] 4. This invention achieves joint optimization of trajectory estimation and semantic behavior understanding by transforming image semantic information into high-dimensional semantic embeddings and constructing semantic residual terms to introduce graph optimization. The system can assist in determining trajectory paths in structurally ambiguous areas such as tunnel entrances, station forks, and bridge transitions, improving the system's reliability in ambiguous scenarios.

[0131] 5. This invention proposes a Comprehensive Degradation Scoring Function (CDS), which can dynamically adjust the weights of multimodal factors based on real-time perceived quality, and realize adaptive switching and optimization control of graph optimization in three states: normal, mild degradation and severe degradation, which is beneficial to improving the overall accuracy and stability of the system. Attached Figure Description

[0132] Figure 1 This is a flowchart illustrating a multimodal fusion and semantic enhancement localization method according to the present invention. Detailed Implementation

[0133] The following will describe in more detail a train method and system for multimodal fusion and semantic enhancement localization according to the present invention, with reference to the schematic diagrams, which illustrate preferred embodiments of the present invention. It should be understood that those skilled in the art can modify the present invention described herein while still achieving the advantageous effects of the present invention. Therefore, the following description should be understood as being of general knowledge to those skilled in the art and is not intended to limit the present invention.

[0134] Example 1

[0135] This embodiment describes a train localization method based on multimodal fusion and semantic enhancement for high-speed trains in GNSS degraded environments, in order to solve at least the following problems existing in the prior art:

[0136] 1. The data from various modal sensors were not uniformly aligned in time and space, resulting in low fusion accuracy;

[0137] 2. Inability to promptly identify scenarios of sensory degradation and lack of dynamic response mechanisms;

[0138] 3. The lack of modeling and constraints on the track structure geometry leads to poor positioning stability;

[0139] 4. The compensation mechanism is too simplistic and cannot operate reliably for extended periods when a failure is detected.

[0140] 5. It cannot incorporate semantic information into trajectory reasoning and lacks contextual intelligence;

[0141] 6. The optimization framework is static and does not dynamically adjust the weights of each factor based on the observation quality.

[0142] like Figure 1 As shown, the train localization method based on multimodal fusion and semantic enhancement includes the following steps:

[0143] Step 1: Spatiotemporally align the data collected by the multimodal sensors to obtain a dense point cloud, and construct a dense semantic point cloud through semantic enhancement; dynamically estimate the confidence weight of each type of sensor.

[0144] Step 1.1: Time synchronization mechanism.

[0145] Different types of sensors, i.e., different modes, have asynchronous sampling frequencies, therefore they need to be uniformly aligned to the IMU reference time. This invention uses linear interpolation to align state variables, and the specific formula is as follows:

[0146] ;

[0147] in, This represents the calculated state value of a certain sensor at the IMU time. In time The observation value of a certain sensor at a given time. In time The observation value of a certain sensor at a given time. For the timestamp of IMU data, , These are two consecutive sampling time points of a certain sensor.

[0148] In this embodiment, the multimodal sensors include cameras, lidar, millimeter-wave radar, IMUs, and wheel speedometers.

[0149] Step 1.2: Align the millimeter-wave radar point cloud with the lidar coordinate system through extrinsic parameter transformation, and use kernel function weighted interpolation method to perform spatial completion and generate a dense point cloud to enhance geometric integrity.

[0150] While lidar boasts high precision, it is prone to sparse point clouds or structural voids in harsh environments, affecting feature matching and image optimization. Therefore, this embodiment introduces millimeter-wave radar to fill in the void areas in the lidar, improving the integrity of the point cloud and providing strong penetration and anti-interference capabilities. This invention constructs a dense point cloud and integrates millimeter-wave radar and lidar, which helps improve the positioning robustness and accuracy of this invention in complex environments (such as tunnels, rain and fog, and areas with repetitive structures).

[0151] Dense point clouds not only enhance the reliability of geometric constraints such as point-to-plane and Manhattan, but also provide high-quality input for subsequent semantic enhancement and degradation recognition steps, thus providing robust support for subsequent state estimation and optimization.

[0152] Step 1.2.1: External parameter calibration (spatial coordinate transformation).

[0153] The spatial transformation extrinsic parameter matrix of the millimeter-wave radar to lidar coordinate system is obtained through extrinsic parameter calibration, as shown in the following formula:

[0154] ;

[0155] in, Represents the extrinsic parameter matrix. To transform the coordinate system from millimeter-wave radar to lidar, a rotation matrix is ​​used. This is the translation vector from the millimeter-wave radar to the lidar coordinate system.

[0156] Using the extrinsic matrix The formula for converting millimeter-wave radar points to the lidar coordinate system is as follows:

[0157] ;

[0158] in, The coordinates of the millimeter-wave radar points are converted to the coordinates of the lidar coordinate system; These are the coordinates of the millimeter-wave radar point in the millimeter-wave radar coordinate system. as well as The explanations have been provided above and will not be repeated here.

[0159] Step 1.2.2: Dense blending after alignment.

[0160] Because lidar point clouds contain voids when traversing tunnels and areas of repetitive structures, this invention introduces millimeter-wave radar data for spatial completion and employs a kernel-weighted interpolation method to enhance the density of the point cloud, as detailed below:

[0161] ;

[0162] ;

[0163] in, These are the new points after interpolation, used to complete the LiDAR points or mix information.

[0164] It is the first The position vector of each lidar point (target point). It is the first The position vectors (nearest points) of millimeter-wave radar points. Let Euclidean distance be the distance between two points, and let represent the similarity.

[0165] This is a distance-weighted kernel function, meaning that the closer the distance, the greater the weight. The kernel bandwidth controls the weighted decay rate. r represents the relative distance ratio, used to control the degree to which the weights decay with distance.

[0166] Step 1.3: Semantic enhancement (i.e., joint modeling of image and point cloud).

[0167] This invention uses image semantic tags, aligned with corresponding LiDAR, to assign actual semantic information to three-dimensional coordinates.

[0168] A semantic label map is generated using a semantic segmentation network for the current frame image. The LiDAR points are projected onto the image plane to align semantic information with the 3D structure. The dense point cloud is then projected onto the image plane through extrinsic parameter transformation. The semantic category of the corresponding pixel is queried in the semantic label map, and the semantic label is backfilled into the corresponding 3D point to finally generate a dense semantic point cloud.

[0169] Step 1.3.1: Image semantic segmentation.

[0170] First, examine the current frame image. Run the semantic segmentation network to generate a semantic label map:

[0171] .

[0172] in, The semantic label map represents each pixel, and the output is an integer image (each pixel is a category number). The input image has a height of [value missing]. ,Width RGB three channels; These are the pixel coordinates in the image. The horizontal pixel position. Vertical pixel position, Semantic category index, total kind.

[0173] Semantic categories should cover typical structural elements and key objects in the operational scenarios of rail transit, such as:

[0174] rail, sleepers(ties), ballast, roadbed, platform, switch(turnout), catenary, pole, signal, tunnel ceiling, bridge structure, pedestrian, vegetation.

[0175] This represents each pixel pair output by the semantic segmentation network. The predicted score (unnormalized probability) of the class; The category with the highest predicted value is selected as the final label for that pixel.

[0176] Common semantic segmentation networks include DeepLabV3+ and SegFormer.

[0177] The purpose of steps 1.3.2 and 1.3.3 below is to achieve the alignment and fusion of image semantic information and point cloud data.

[0178] Step 1.3.2: Coordinate transformation (external parameters).

[0179] 3D points in the camera coordinate system The transformation formula to the lidar coordinate system is as follows:

[0180] ;

[0181] in, The transformed 3D coordinates of the camera point in the lidar coordinate system. This represents the rotation matrix that rotates the coordinate points from the camera frame to the lidar frame. This represents the extrinsic translation vector.

[0182] Step 1.3.3. Projection transformation (intrinsic parameters).

[0183] This step utilizes the camera's imaging model to convert three-dimensional spatial points into pixel coordinates on the image plane, and establishes the projection relationship from three-dimensional points to two-dimensional images based on the camera's intrinsic parameter information, providing a foundation for subsequent visual observation and association with three-dimensional structures.

[0184] Step 1.3.4: Semantic tag mapping.

[0185] Each projection point Find its semantic tags in the semantic tag graph: ; concatenate it with the original point cloud to form a semantically enhanced point cloud, i.e. .

[0186] in, It refers to structured point cloud data points with semantic attributes, i.e., dense semantic point cloud.

[0187] Pixels in the image semantic tag number, These are the coordinates of a three-dimensional point in the coordinate system of a lidar or camera. The semantic category from the image is appended to the point. This ultimately forms a structured, semantically enhanced dense point cloud, providing high-level semantic support for subsequent map construction and trajectory optimization.

[0188] All laser points were attached with image semantic tags to form a structured semantic point cloud.

[0189] The resulting dense semantic point cloud will have the following format:

[0190] ;in The number of points in the point cloud. For the first The three-dimensional spatial position of each point For the first Semantic labels for each point (such as rail, wall, ground, etc.).

[0191] Step 1.4: Uncertainty modeling.

[0192] In complex environments, some sensors may be affected by external interference (such as shading, lighting, rain, fog, etc.), leading to distorted observations. Treating all sensor data with equal weight may result in erroneous information that negatively impacts localization estimation. Therefore, it is necessary to dynamically estimate the confidence levels of various sensor observations and adjust the sensor influence weights in factor graph optimization accordingly.

[0193] Step 1.4.1: Calculate the deviation between the sensor observations and the historical average.

[0194] Step 1.4.2: Use the Bayesian method to estimate the observation variance. Using a series of historical observations, construct a Bayesian model to estimate the observation variance of this type of sensor.

[0195] Step 1.4.3: Calculate the confidence weights.

[0196] Based on the estimated observation variance Calculate the confidence weights of the sensors Used to adjust the residual influence in factor graph optimization, sensor confidence weights The calculation formula is as follows:

[0197] ;

[0198] in, For the first Sensor-like devices at all times Confidence weights To prevent small constants from being divided by zero, they are usually set to 0. to between, This indicates the uncertainty of the current estimate; the larger the value, the less reliable it is. The smaller.

[0199] The semantic enhancement fusion mechanism of multimodal perception data in this invention differs from traditional simple fusion methods. Based on laser, IMU, camera, GNSS, etc., it adds image semantic segmentation and Bayesian confidence modeling mechanism to realize the construction of dense point cloud with semantic labels and dynamic weight adjustment, providing data input with clear structure and strong semantic interpretation capability for subsequent modules.

[0200] Step 2. Construct geometric residuals from points to planes and Manhattan structure constraint residuals based on dense semantic point clouds to obtain lidar point cloud constraints; at the same time, construct visual projection constraints based on visual point projection residuals and visual line projection residuals.

[0201] Dense semantic point clouds not only comprehensively describe the geometric structure of the environment but also eliminate semantic ambiguity in feature matching through semantic labels. The continuity and smoothness brought by density make the subsequently fitted plane more accurate and provide a good foundation for point-to-plane residual calculation. Based on semantic labels, points with consistent semantics can be selected for point-to-plane residual calculation, significantly improving matching accuracy; based on the estimated normal vectors in the point cloud, normal constraints can be added to enhance geometric consistency; when the semantic recognition is of Manhattan structure scenes such as tunnels and stations, Manhattan constraints can also be introduced using the plane normal vectors fitted from the point cloud. Therefore, dense semantic point clouds are the core support for subsequent geometric constraints, semantic constraints, and structural priors.

[0202] Step 2.1: Point-to-plane matching residuals of dense semantic point clouds. Step 2.1 aims to construct point-to-plane geometric residuals based on dense semantic point clouds for pose constraints in subsequent graph optimization.

[0203] This step utilizes the dense semantic point clouds of the current frame and the reference frame to establish relationships between points using semantic information, and constructs geometric constraints based on the offset between a point and its corresponding local plane. By estimating the matching relationship between the current frame's point cloud and the reference frame, an error quantity reflecting the spatial consistency of the two frame point clouds is obtained. This error is used to constrain pose in subsequent optimization processes, thereby improving inter-frame registration accuracy and overall positioning stability.

[0204] Step 2.2: Construct Manhattan structure constraint residuals.

[0205] To further improve the constraint capability of pose estimation in environments with structural rules, this step constructs constraint terms that align the plane normal with the global coordinate axis when identifying Manhattan structural scenes (such as tunnels, stations, etc.).

[0206] By calculating the minimum angular residual between the local plane normal of the current frame and the global Manhattan axis, and incorporating it as a penalty term into the optimization, the point cloud tends to align with the Manhattan direction in the structured scene, thereby improving the accuracy and stability of pose estimation.

[0207] Step 2.3: Construct the visual point projection residual.

[0208] Step 2.3.1: Select the set of key points for LiDAR scanning and the pixel observation points in the current image frame ,in, , As the key point, ; , representing the pixel coordinates of the corresponding projection of a 3D point in the image;

[0209] Step 2.3.2: Define the camera pose to be estimated in the current frame, i.e., the rotation matrix. With translation vector Use this pose to locate the key points in the world coordinate system. Transform to the camera coordinate system, and then project onto the image plane using the camera intrinsic parameters to obtain the predicted pixel coordinates.

[0210] Step 2.3.3: Define pixel projection error residual ,

[0211] in, Let be the Euclidean distance between the pixel position of a point on a 3D map after projection and its actual observation position, then the visual point projection residual is... The expression is:

[0212] .

[0213] Step 2.4: Construct visual line projection residuals (combined constraints of trajectory lines and point lines).

[0214] Step 2.4.1: Let the two endpoints of the three-dimensional line segment be... and Based on the current camera pose and projection model, two points can be mapped onto the image plane to obtain their corresponding pixel coordinates. and .

[0215] Step 2.4.2: Define the set of line segments for image observation: ,in, These are the start and end points of the line segment in pixel space, respectively;

[0216] Step 2.4.3: Define the current camera pose, i.e., the rotation matrix. Translation vector The camera intrinsic parameter matrix is Project the start and end points of the 3D line segment onto the image plane to obtain the predicted pixel coordinates.

[0217] Then the visual line projection residual The expression is:

[0218] ;

[0219] in, For the index of the observed line segment; and These are the predicted pixel coordinates and the actual observed pixel coordinates of the starting point of the k-th observed line segment, respectively. and These are the predicted pixel coordinates and the actual observed pixel coordinates of the endpoint of the k-th observed line segment, respectively.

[0220] Step 2.5: Visual Projection Constraints Represented as: + .

[0221] Step 3. Based on the joint semantic information of dense semantic point cloud and image, identify whether there is perception degradation in the current environment of the high-speed rail, and construct a comprehensive degradation scoring function based on multiple indicators to determine degradation.

[0222] Step 3, based on the fused point cloud and image semantic information constructed in Step 1, designs a multi-index perception degradation recognition method for situations where the perception modal point cloud and image may fail in typical high-speed rail scenarios (such as tunnels, rain and fog, stations, etc.).

[0223] This step utilizes point cloud fusion information and image semantic information from dense semantic point clouds to identify whether there is perception degradation in the current environment of the high-speed rail (such as sparse laser point clouds, blurred images, rain and fog obstruction, etc.).

[0224] Once a degradation risk is detected, a subsequent compensation mechanism (i.e., step 4) will be triggered to dynamically adjust the weights of each sensing factor, ensuring that the system can maintain stable state estimation performance even in complex environments.

[0225] Identify whether perception degradation has occurred in the current environment; assess the reliability of the current state estimate; and trigger subsequent compensation mechanisms.

[0226] Step 3.1: Point Cloud Structure Degradation Identification .

[0227] Using the point cloud normal entropy and structural dimension index obtained in step 1.2, the normal vectors in the dense point cloud are projected onto a unit sphere. Multiple direction buckets are constructed by uniformly dividing the sphere or by predefined direction sets. The number of normals in each direction bucket is counted and normalized to a probability distribution, and then the point cloud normal entropy is calculated.

[0228] This metric reflects the geometric orientation consistency of point clouds and can be used to determine whether there is structural degradation in the current environment (such as orientation confusion, rain or fog obstruction, etc.). For dense point clouds, the normal vectors are calculated and then divided according to their orientation. For each direction bucket, calculate the percentage of normal vectors in each bucket. And calculate the point cloud normal entropy, as shown in the following formula:

[0229] ;

[0230] in, This represents the normal entropy of a point cloud. The normal entropy measures the amount of information contained in the distribution of the normal vector directions on the surface of a point cloud, thus reflecting the spatial distribution characteristics of the point cloud. For the first The normalized frequency of the normal in each direction bucket.

[0231] The larger the value, the more disoriented (perceptual degeneration). The smaller the value, the clearer the structure.

[0232] Step 3.2: Degradation Index (DDI).

[0233] To determine whether the current frame's point cloud exhibits spatial structure degradation (such as tunnels, openings, or empty areas), this invention proposes a method for calculating the Dimensional Degeneration Index (DDI) based on the eigenvalues ​​of the Hessian matrix, used to assess the three-dimensional distribution integrity of the point cloud. The input to this calculation also comes from a dense point cloud, aiming to determine whether the point cloud is only distributed in a one-dimensional / two-dimensional plane and lacks spatial volume information (such as tunnels or cavities). For any frame of lidar point cloud, the system calculates the point cloud gradient field within its local region, and then constructs the Hessian matrix:

[0234] ;

[0235] in, This represents the density or reflection intensity distribution function of a point cloud in space.

[0236] The Hessian matrix characterizes the second derivative information of the local geometric changes in a point cloud, and its eigenvalues... (Sorted from largest to smallest) can be used to measure the local structure distribution of point clouds.

[0237] Take its eigenvalues Based on the eigenvalues ​​of the Hessian matrix, the degradation index DDI is defined as:

[0238] ;

[0239] in, The DDI values ​​are eigenvalues ​​of the Hessian matrix, used to describe the extent of structural expansion in different dimensions. DDI values ​​range from 0 to 1, indicating the completeness of the local structural dimensions. A smaller DDI value indicates structural degradation (concentrated on lines / areas), while a larger DDI value indicates rich structure. When DDI is close to 1, meaning the three eigenvalues ​​are similar, it indicates that the point cloud locally exhibits volumetric structure, with sufficient information dimensions and a reliable structure. When DDI is close to 0, it indicates... This indicates that the point cloud is mainly concentrated in a unidirectional or two-dimensional surface structure, and the structure is degraded.

[0240] Step 3.3: Image texture quality assessment (edge ​​density).

[0241] To determine whether the current frame image has sufficient texture information to support the construction of visual residuals, this invention introduces an image sharpness evaluation method based on edge density index, which is used to automatically identify whether the image has texture degradation such as blurring, nighttime, or overexposure in strong light, and dynamically enable or disable visual factors accordingly.

[0242] The semantic image frames obtained in step 1.3 are used to evaluate image sharpness and detect whether the image is blurry, dark, or bright, which may cause the loss of texture information and affect the reliability of visual factors.

[0243] Canny edge detection was used to extract edge pixels from the semantic image, and the number of edge pixels was counted. And calculate the edge density:

[0244] ;

[0245] in, This represents the number of edge pixels after Canny edge detection. This represents the total number of pixels in the image. Image edge density reflects the richness of texture information. A lack of edges in an image indicates weak texture, such as in nighttime or blurry conditions. This represents a threshold for determining whether an image quality meets the standard. If the image is deemed invalid, the system will automatically disable visually relevant factors (such as point projection residuals and line projection residuals) to prevent low-quality images from misleading the optimization results.

[0246] Step 3.4: Calculate the confidence level of the multimodal sensor.

[0247] In complex environments, some sensors may be affected by external interference (such as occlusion, lighting, rain, fog, etc.), leading to distorted observations. Treating all sensor data with equal weight may result in erroneous information that negatively impacts localization estimation. Therefore, the system needs to dynamically estimate the confidence levels of various sensor observations and adjust the sensor influence weights in factor graph optimization accordingly.

[0248] For the calculation method of confidence scores of various sensor observations, please refer to steps 1.4.1 to 1.4.3. Based on the above confidence score calculation method, the confidence scores of the lidar and camera at time [time value missing] are calculated. Confidence weight .

[0249] Step 3.5: Comprehensive degradation score.

[0250] The Comprehensive Degeneration Score (CDS) construction and system factor switching strategy are used to uniformly quantify and judge the multimodal perception quality of the current frame. This invention further constructs a Comprehensive Degeneration Score (CDS) to drive the system to dynamically adjust weights and factor activation strategies.

[0251] Definition of scoring function:

[0252] ;

[0253] in, Indicates the time between the lidar and the camera Confidence weights These represent task-related weighting coefficients used to adjust the contribution of each modality to the degradation score; =1.

[0254] , express , The maximum value is used for normalization. It is 2.5. It is 0.3.

[0255] A higher CDS score indicates a worse overall system perception state. The system classifies CDS scores into three levels:

[0256] If CDS≤0.3, it is determined to be a normal scene, and all factors are enabled, namely LiDAR point cloud constraint factor, visual projection constraint factor, compensation constraint factor, prior semantic constraint factor and large model semantic factor.

[0257] If so, it is judged as moderate degradation, and the weights of visual projection and lidar point cloud factors are reduced;

[0258] CDS If the condition is found to be severely degraded, all sensing factors in all factors will be disabled, and the orbit + IMU + wheel speed compensation mechanism will be enabled, retaining only the compensation constraint factor, i.e., completely relying on the compensation constraint factor.

[0259] This invention proposes a real-time degradation environment identification and adaptive perception factor control mechanism. For the first time, it proposes a perception degradation discrimination method that combines point cloud structure indicators (normal entropy, Hessian feature value) and image texture indicators (edge ​​density), and integrates them into a comprehensive degradation score CDS. This CDS is used to dynamically adjust the trust level of each sensor factor to achieve active management and mode switching of perception capabilities.

[0260] Step 4: After completing the degradation determination using the comprehensive degradation scoring function, if a degradation state is entered, the system automatically switches to the redundant compensation factor modeling mechanism. This introduces structural and motion information independent of the external environment to continue supporting pose estimation optimization. Simultaneously, compensation constraints are constructed based on track structure geometric factors, IMU pre-integration factors, and wheel axle velocity factors.

[0261] In real-world railway environments, lidar and vision sensors are highly susceptible to factors such as rain and fog, darkness, strong light reflection, and simple structures, leading to problems like perception degradation, sparse point clouds, and missing image textures. To improve the system's robustness under perception degradation, this invention automatically switches to a redundant compensation factor modeling mechanism after completing the degradation determination in step 3. This introduces structural and motion information independent of the external environment to continue supporting pose estimation optimization. After entering the degradation state, the system adds the following three types of compensation factors to the factor graph optimization: track structure geometry factor: derived from track centerline modeling; IMU pre-integration factor: providing continuous pose / velocity information; wheel and axle speed factor: supplementing the observed projection of train speed.

[0262] Step 4.1: Construct the geometric model constraints of the track structure.

[0263] In high-speed rail positioning systems, to prevent lateral drift or illegal deviations (such as deviating from the track area or crossing physical boundaries) of train tracks when perception degrades, this invention introduces a spatial geometric constraint method based on a track centerline model. This method constructs track constraint factors based on prior track structure, enhancing the physical rationality and spatial consistency of trajectory estimation.

[0264] The input to the geometric model constraints of the track structure is: the track centerline model;

[0265] The model is derived from map / B-spline / rule fitting, and the current estimated state orbit centerline is modeled as a space curve. :

[0266] ;

[0267] in, This indicates the lateral deviation (error) between the current pose point and the center line of the orbit. The estimated position at the current moment is provided by the vehicle's positioning system (such as a transponder or track circuit). For the current point The arc length corresponding to the closest position online. The three-dimensional spatial parameter curve of the orbit centerline (which can be derived from a map, B-spline, etc.).

[0268] The high-speed rail is constrained to run on a physical track; strong geometric constraints are provided to prevent track drift or crossing illegal areas; and the track structure is a highly reliable positioning reference when perception is unavailable.

[0269] Step 4.2: IMU pre-integration constraint modeling.

[0270] To maintain the continuity and short-term accuracy of trajectory estimation even when sensor perception degrades (such as LiDAR and visual failure), this invention introduces an inertial factor modeling method based on IMU pre-integration. This method utilizes high-frequency acceleration and angular velocity information provided by the inertial measurement unit (IMU) to estimate state changes between keyframes and incorporates them as constraints in graph optimization. The input to the IMU pre-integration constraint modeling is short-term position and acceleration information from the time-synchronized IMU data in step 1, representing the current keyframe. With the next keyframe State estimation is performed. The IMU provides high-frequency acceleration and angular velocity; reliable over short time periods (e.g., hundreds of milliseconds). Pre-integration techniques are used to connect keyframe states. The difference between the state changes observed by the IMU and the trajectory estimates is constructed as residuals; these residuals are minimized to form inertial constraints on the trajectory.

[0271] ;

[0272] in, It represents the pose (rotation matrix) between the current frame and the next frame. This represents the velocity vector between the current frame and the next frame. This indicates the position coordinates between the current frame and the next frame.

[0273] This indicates the rotation, velocity, and position predicted by the IMU integral (from the IMU).

[0274] It is the gravitational acceleration vector. This refers to the keyframe interval time. This is the covariance matrix (representing uncertainty).

[0275] By integrating IMU data, we can obtain the keyframes. arrive State prediction values: The rotation obtained by pre-integration; The velocity change obtained through pre-integration; The positional change obtained from pre-integration.

[0276] The process does not rely on real-state estimation, but is based solely on the raw IMU observations, initial state, and gravity vector.

[0277] The IMU pre-integration factor, by providing high-frequency, short-term reliable pose and velocity change constraints, can maintain the continuity and stability of trajectory estimation even during periods when visual or laser sensing factors fail. Combined with orbital geometry factors and wheel velocity observation factors, a complementary multi-source compensation mechanism can be formed, effectively suppressing system drift and divergence. Simultaneously, the high-frequency characteristics of the IMU can provide dense state interpolation between keyframes, further enhancing the overall system's temporal consistency and positioning robustness.

[0278] Step 4.3: Wheel speed observation constraints (wheel speed meter).

[0279] In scenarios with degraded perception (such as blurred images or sparse point clouds), this invention introduces a velocity observation constraint factor based on a wheel encoder to further improve the stability of the system's velocity estimation and trajectory constraint capability.

[0280] This factor utilizes the linear velocity information measured by the train's wheel and axle encoders, combined with the track direction vector, to directly project constraints onto the estimated velocities in the system. The input information for the wheel and axle velocity observation constraints is:

[0281] The velocity vector of the current frame state estimation: ;

[0282] Current orbital direction unit vector: , is usually the tangent direction of the orbital curve at that point;

[0283] The scalar value of the wheel axle speed observed by the encoder: This indicates the train's actual speed in the direction of its movement along the track.

[0284] Based on the principle of "consistency of velocity direction", the estimated velocity component in the track direction is aligned with the wheel speed meter value.

[0285] Define the residual function as:

[0286] .

[0287] This residual measures the degree of deviation of the system's estimated velocity in the direction of track propulsion, effectively constraining velocity misestimation caused by sensing drift. The wheel and axle encoder, as a built-in sensor, is unaffected by external environmental factors such as lighting and obstruction, exhibiting long-term stability and high accuracy. Combined with the propulsion direction obtained from the track centerline tangent, it can impose physically meaningful directional constraints on the system's velocity estimation. Even when visual or laser factors fail, the encoder can still continuously provide stable velocity observations, effectively compensating for sensing deficiencies and maintaining the smoothness and continuity of trajectory estimation.

[0288] Step 4.4: Graph optimization integration of compensation modeling.

[0289] To achieve adaptive and robust localization of the system under different sensing states, this invention integrates the track geometry factor, IMU pre-integration factor, and wheel axle speed observation factor into the graph optimization framework as compensation constraint residual terms between key frame nodes.

[0290] When the primary sensing modality (laser, vision) is operating normally, the compensation factor is weakened; however, in scenarios of perception degradation, the influence of the compensation factor is automatically enhanced to maintain the stability and continuity of the overall trajectory estimation.

[0291] make This represents the set of state variables for all keyframes, including pose, velocity, IMU bias, etc., for each frame. The overall optimization objective function is constructed as follows:

[0292] = ;

[0293] in, , For the geometric residuals of the track structure, For IMU pre-integration residuals, For the observation residual of wheel and axle speed, For each type of factor, the dynamic weight coefficients are... For the first The factor in the th Residual function under class constraints.

[0294] To adapt to different perceived quality states, the system dynamically adjusts the weight allocation of various factors in factor graph optimization based on the Comprehensive Degradation Score (CDS) calculated in step 2, thereby implementing an adaptive perception fusion strategy.

[0295] CDS<0.3 (normal perception): Laser point cloud factor and visual projection factor are dominant, providing accurate geometric and texture constraints; compensation factors such as track, IMU, and wheel speed are secondary and have lower weights.

[0296] 0.3≤CDS<0.6 (mild degradation): Appropriately reduce the weight of laser and visual factors, enhance the influence of factors such as orbital structure, IMU pre-integration and wheel speed observation, and improve the system's robustness to potential degradation.

[0297] CDS≥0.6 (severe degradation): The risk of sensory factor failure is high. The system automatically removes laser and visual factors and retains only stable compensation factors such as track, IMU, and wheel speed to participate in optimization, ensuring continuous and reliable trajectory estimation and safe system operation.

[0298] Based on the compensation modeling mechanism of track geometry and elevation priors, this invention introduces structured prior factors such as track centerline constraints, turnout judgment, elevation limits, and wheel speed assistance to construct a three-dimensional geometric compensation model in perception degradation scenarios, which effectively enhances the positioning continuity and vertical stability of the system in GNSS-deficient and laser / visually blurred environments.

[0299] Step 5: In trajectory optimization, prior semantic constraints and large-scale semantic reasoning modeling are introduced to ensure the rationality of the trajectory.

[0300] In complex railway scenarios, traditional trajectory optimization methods based on geometric features and low-level observations are prone to failure in environments with degraded perception, structural ambiguity, or semantic ambiguity (such as tunnel entrances, turnouts, red lights, bridges, etc.), leading to serious deviations in trajectory estimation, such as derailment, running red lights, drifting, and speeding.

[0301] To overcome this problem, this invention proposes a trajectory optimization modeling strategy that integrates prior semantic hard constraints with large model semantic soft reasoning. This enables the system to have the ability to understand contextual logic such as "stop at a red light", "slow down in a tunnel", and "do not deviate from the track" similar to a human driver, thereby maintaining trajectory rationality and operational consistency even in scenarios with weak perception.

[0302] To ensure that the trajectory estimation results conform to high-speed rail operation rules and physical rationality, the system introduces prior semantic constraint factors in trajectory optimization. These factors, derived from railway infrastructure and operation standards, participate in factor graph optimization as hard constraints, effectively preventing "unreasonable" trajectories caused by perception degradation or matching errors. They mainly include the following four categories:

[0303] Step 5.1: Prior semantic constraint modeling, expressed by the following formula:

[0304] .

[0305] Step 5.1.1: .

[0306] The estimated velocity of the constrained trajectory must not exceed the speed limit of the current track segment. This is to prevent the creation of "fictional" high-speed trajectories during the optimization process to reduce residuals. Its error term is defined as:

[0307] .

[0308] Step 5.1.2: Signal Factor (Red Light Stop Rule). Based on the signal status set on the track, if the corresponding signal light is red (Stop), the system forcibly constrains the estimated speed for that frame. Approaching zero is defined as:

[0309] ;

[0310] in, This factor is activated when the traffic light at the current location is red, ensuring that the trajectory does not exhibit "running a red light" or other violations.

[0311] Step 5.1.3: Elevation Factor (Vertical track constraint).

[0312] Tracks have a specific elevation orientation in space, especially in areas such as bridges, tunnels, and ramps. The system estimates the position by comparing... Corresponding elevation on the track curve The deviation between them constructs the residual:

[0313] ;

[0314] in, Represents the height component of the estimated location. Indicates the position of the track The elevation value at the location. To prevent the track from drifting to a height other than the track, ensuring that the positioning result strictly conforms to the track elevation characteristics.

[0315] For example, in road sections with drastic terrain changes such as overpasses, slopes, and tunnels, perception degradation may cause the trajectory to "float" or "sink". This invention, by increasing the elevation factor, can ensure that the trajectory changes with the elevation curve.

[0316] Step 5.1.4: Turnout switching factor (Smooth transition of geometric model).

[0317] When approaching the turnout, change the track function. For branching tracks To maintain continuity, the formula is as follows:

[0318] ;

[0319] in, Let be the position vector of the orbital centerline in three-dimensional space; Let be the position vector of the centerline of the branch track in three-dimensional space.

[0320] Ensure that the track smoothly switches to the correct track branch in the turnout area to avoid track skipping or derailment.

[0321] For example, when a train enters a switch area, sparse sensing features can easily lead to track "crossing" or "abrupt changes." This invention addresses this by introducing a switch switching factor. This ensures that the trajectory smoothly transitions according to the actual track branches.

[0322] The prior semantic constraints proposed in this invention apply physical and rule-based constraints to trajectory optimization by introducing high-speed rail operation rules such as speed limits, signals, elevation, and turnouts. These constraints compensate for the shortcomings of low-level features and geometric constraints in complex or degenerate scenarios, preventing estimation results from exceeding speed limits, derailing, running red lights, or drifting, thereby ensuring that the trajectory is both reasonable and safe.

[0323] The prior semantic constraints can be obtained as follows: .

[0324] Step 5.2: Construction of semantic factors for large models.

[0325] To enhance the semantic reasoning capability of trajectory constraints, the system introduces a multimodal large model (such as CLIP) to encode image frames into high-dimensional semantic embedding vectors and match them with preset prompt words to construct semantic constraint factors.

[0326] Using this as a soft constraint, the trajectory optimization is guided to approach the state of real semantic reasoning. This enables the system to understand the reasoning ability of "scene-behavior-trajectory" and can still operate stably even when the structure is ambiguous or the trajectory is not unique.

[0327] Step 5.2.1: Constructing the semantic cue word set.

[0328] A semantic prompt word library is constructed based on the high-speed rail scenario, including terms such as "entering the tunnel," "left fork," "city road," "stopping at the station," and "exiting the tunnel." Each prompt word is encoded into a high-dimensional semantic vector using a large model.

[0329] ;

[0330] in, Indicates the first The semantic embedding vector of each prompt word For large models of text encoders (such as The text encoder's prompt dictionary contains a total of Each prompt word corresponds to A vector.

[0331] Step 5.2.2: Image semantic embedding extraction ( Extract semantic embedding vectors from each frame of the image (from step 1):

[0332] ;

[0333] in, for Frame image, This is the semantic vector of the current image frame. For large models Visual encoder, vector This represents the high-level semantic features of the current frame image.

[0334] Step 5.2.3: Calculate semantic residuals.

[0335] Will With each of the prompt words The formula for calculating similarity (such as cosine similarity) is as follows:

[0336] ;

[0337] Retrieve the most matching prompt word from the prompt word library. The formula is expressed as follows:

[0338] ;

[0339] Then the semantic residual of the current frame is defined as:

[0340] ;

[0341] in, The semantic residual term in the k-th frame represents the difference between the semantics of the current image and the best prompt word. Indicates and The closest prompt word vector.

[0342] Step 5.2.4: Add semantic factors to the factor graph.

[0343] In factor graph optimization, the semantic factor weights are dynamically adjusted based on the image confidence (derived from step 1.4).

[0344] For example, when the image edge density is low or the image is blurry, the semantic factor weight is reduced or disabled; conversely, the semantic factor is given a higher weight to enhance the trajectory constraint.

[0345] ;

[0346] in, These are the dynamic weighting coefficients of the semantic factors. The value is based on the camera confidence score in step 1.4 and the image texture quality assessment in step 3.3. This represents the weighted semantic residual term.

[0347] The aforementioned semantic constraints will be integrated as soft / hard factors into the graph optimization framework of step 6, working together with geometric constraints to optimize the trajectory. Low-level features and geometric constraints may fail in scenarios with ambiguous structures or complex semantics. Similar to a driver needing to understand the context of red lights, speed limits, and ramps, high-speed rail trajectory optimization also requires understanding operational rules and scenario semantics. This invention ensures trajectory rationality through a combination of prior semantic constraints (hard rules) and large-scale model semantic reasoning (soft constraints).

[0348] A trajectory semantic constraint method integrating operational rules and large-scale model semantics is proposed. Hard rules such as speed limits, traffic lights, elevation, and turnouts are modeled as prior semantic factors and introduced into a large-scale model for image semantic embedding and inference, constructing soft semantic factors to achieve semantic understanding and consistent scene behavior control during trajectory optimization. The method also proposes to utilize... Multimodal pre-trained models extract semantic vectors from images and construct residual terms between these vectors and prompt words as high-level semantic constraint factors in the factor graph. This endows the system with scene understanding and trajectory reasoning capabilities, solving the problem of traditional SLAM systems failing in scenarios with structural ambiguity and trajectory ambiguity.

[0349] Step 6. Under the factor graph optimization framework, a global optimization objective function is constructed by comprehensively introducing multiple constraints, including LiDAR point cloud, visual projection, compensation, prior semantics, and large model semantic factors.

[0350] The weights of each modal factor in the factor graph optimization are dynamically adjusted based on the confidence levels of various sensors.

[0351] The optimal state estimate is obtained by iteratively solving the global optimization objective function, and finally the high-precision train positioning result is output.

[0352] In degraded environments (determined in step 3), the system provides redundancy constraints using trajectory geometry, IMU, and wheel speed factors in step 4. In non-degraded scenarios, LiDAR and visual factors dominate the optimization. Both types of factors are integrated into the factor graph through dynamic weights. After the compensation modeling in step 4, the system is able to maintain trajectory estimation even in degraded scenarios.

[0353] To further improve the global consistency, long-term stability and accuracy of positioning, factor graph optimization is introduced to unify the modeling of various observation information and eliminate error accumulation through sliding window and loop closure detection mechanisms.

[0354] The set of state variables, i.e., the optimization objective, is: ;in Indicates the first Frame pose, Indicates speed, Indicates location, This represents the number of image frames. The confidence weights for each sensor are calculated based on step 1.4. The system dynamically adjusts the weights of each modality factor during factor graph optimization. The optimization objective function is updated as follows:

[0355] Will As a factor added to the GTSAM graph optimization, it is listed alongside factors such as LiDAR, vision, compensation, prior semantics, and large model semantics, and is weighted according to the confidence weights of each sensor calculated in step 1.4. The system dynamically adjusts the weights of each modality factor during factor graph optimization. The optimization objective function is updated as follows:

[0356] ;

[0357] in, , The pose of each frame, For the position of each frame, For frame rate; To optimize the number of keyframes within the sliding window of the factor graph, it is generally a fixed value, such as... . and Confidence weights from LiDAR and vision, respectively. The weight of the compensation factor is dynamically increased based on the degradation score CDS; and The weights for semantic factors depend on image sharpness and semantic matching degree. Based on the comprehensive degradation score CDS from step 2.5, the system categorizes the environmental state into three levels and maps them to a weight adjustment strategy:

[0358] I. Normal scenario, i.e., CDS < 0.3: LiDAR and visual factors dominate: , Compensation factor assistance: Semantic factors enabled: , ;

[0359] II. Mild degradation scenarios, i.e., 0.3 ≤ CDS < 0.6: Reduce laser / visual weights: , Enhanced compensation factor: Semantic factors should be used with caution. , ;

[0360] III. In severely degraded scenarios, i.e., when CDS ≥ 0.6, disable unreliable modes: , Completely dependent on compensation factors: Semantic factors disabled: , .

[0361] The final set of state variables output by the system Includes pose, velocity, and weight information for each frame:

[0362] ;

[0363] in, Indicates the first The real-time weights of each modality in the frame are used for monitoring the system's operating status and for debugging.

[0364] The globally unified sliding window graph optimization framework integrates semantic, structural, and behavioral multi-layer constraints. This invention incorporates laser inter-frame factors, visual reprojection factors, IMU pre-integration factors, GNSS original factors, orbital geometry factors, elevation factors, wheel speed factors, semantic factors, and loop closure factors into the sliding window graph optimization framework in a unified residual form, and dynamically adjusts the weights according to the degradation score to achieve adaptive fusion of multi-source information, highly robust positioning, and interpretable trajectory generation.

[0365] This invention constructs a robust data foundation through multimodal sensor fusion in step 1, identifies degradation scenarios in real time by combining indicators such as normal vector entropy and DDI in step 3, and dynamically switches to anti-degradation factors such as orbital geometry and IMU in step 4. Finally, multi-source constraints are uniformly fused in the factor graph optimization in step 6 to achieve the following accuracy improvements:

[0366] 1. To address the high-precision positioning requirements of high-speed rail in complex environments, a robust graph-optimized positioning system integrating multimodal perception, dynamic degradation identification, redundancy compensation modeling, and semantic reasoning was constructed.

[0367] 2. Multi-sensor data fusion and semantic enhancement, degraded environment identification and state assessment, trajectory compensation modeling under perception degradation, construction of prior and large model semantic factors, and unified graph optimization of multi-source factors.

[0368] 3. It adaptively adjusts factor weights according to environmental conditions, maintaining continuous, stable, and high-precision trajectory estimation even in scenarios such as GNSS failure, visual / laser degradation, and structural ambiguity, while possessing both geometric consistency and semantic understanding capabilities.

[0369] 4. Degenerate scenario: Track geometry factors provide absolute position constraints, and IMU / wheel speed suppresses drift.

[0370] 5. Non-degradable scenarios: LiDAR and visual factors dominate, semantically enhanced point clouds improve matching accuracy.

[0371] The method of this invention is particularly applicable to rail transit systems such as high-speed rail and subway. In complex environments such as GNSS signal degradation, visual blurring, and weak laser texture, it achieves high-precision and robust positioning and state estimation by fusing multi-source information such as lidar, inertial measurement unit (IMU), camera, and wheel speed meter, combined with track geometric constraints and semantic reasoning mechanisms.

[0372] Example 2

[0373] This embodiment 2 provides a train positioning system based on multimodal fusion and semantic enhancement. This system is based on the same inventive concept as the train positioning method based on multimodal fusion and semantic enhancement described above.

[0374] The train positioning system based on multimodal fusion and semantic enhancement in this embodiment includes the following modules:

[0375] The multimodal sensor data acquisition and fusion module is used to perform spatiotemporal alignment on the data acquired by multimodal sensors to obtain dense point clouds, and to construct dense semantic point clouds through semantic enhancement; and to dynamically estimate the confidence weight of each type of sensor.

[0376] The radar point cloud and visual projection error construction module is used to construct geometric residuals and Manhattan structure constraint residuals from points to planes based on dense semantic point clouds to obtain lidar point cloud constraints. At the same time, it constructs visual projection constraints based on visual point projection residuals and visual line projection residuals.

[0377] The degradation environment perception and judgment module is used to identify whether there is perception degradation in the current environment of the high-speed rail based on the joint semantic information of dense semantic point cloud and image, and to construct a comprehensive degradation scoring function based on multiple indicators to make degradation judgment.

[0378] The positioning compensation modeling module in the degradation state automatically switches to a redundant compensation factor modeling mechanism after the degradation determination result indicates that the system has entered a degradation state. This mechanism introduces structural and motion information independent of the external environment to maintain trajectory estimation in degradation scenarios. Furthermore, it constructs compensation constraints based on track structure geometric factors, IMU pre-integration factors, and wheel axle velocity factors.

[0379] The prior semantic constraints and large model semantic constraints modeling module is used to introduce prior semantic constraints and large model semantic factor constraints to ensure the rationality of the trajectory.

[0380] The Factor Graph Optimization and State Estimation module is used to construct a global optimization objective function by comprehensively introducing multiple constraints, including LiDAR point cloud, visual projection, compensation, prior semantics, and large model semantic factors, within the factor graph optimization framework.

[0381] Based on the confidence levels of various sensors, the weights of each modal factor in the factor graph optimization are dynamically adjusted.

[0382] The optimal state estimate is obtained by iteratively solving the global optimization objective function, and finally the high-precision train positioning result is output.

[0383] It should be noted that the implementation process of the functions and roles of each functional module in the train positioning system of this embodiment 2 is detailed in the implementation process of the corresponding steps of the method in the above embodiment 1, and will not be repeated here.

[0384] Example 3

[0385] This embodiment 3 describes a computer device. The computer device includes a memory and one or more processors. Executable code is stored in the memory. When the processor executes the executable code, it implements the steps of the train localization method based on multimodal fusion and semantic enhancement described in embodiment 1 above.

[0386] In this embodiment, the computer device can be any device or apparatus with data processing capabilities, and will not be described in detail here.

[0387] Example 4

[0388] This embodiment 4 describes a computer-readable storage medium storing a program that, when executed by a processor, implements the steps of the train positioning method based on multimodal fusion and semantic enhancement in embodiment 1 above.

[0389] The computer-readable storage medium can be an internal storage unit of any device or apparatus with data processing capabilities, such as a hard disk or memory, or an external storage device of any device with data processing capabilities, such as a plug-in hard disk, smart media card (SMC), SD card, flash card, etc.

[0390] The above are merely preferred embodiments of the present invention and do not constitute any limitation on the present invention. Any equivalent substitutions or modifications made by those skilled in the art to the technical solutions and content disclosed in the present invention without departing from the scope of the present invention shall be deemed to have remained within the protection scope of the present invention.

Claims

1. A train localization method based on multimodal fusion and semantic enhancement, characterized in that, Includes the following steps: S1: Spatiotemporal alignment of data collected by multimodal sensors is performed to obtain dense point clouds, and semantic enhancement is used to construct dense semantic point clouds to dynamically estimate the confidence weight of each type of sensor. S2: Construct geometric residuals and Manhattan structure constraint residuals from points to planes based on dense semantic point clouds to obtain lidar point cloud constraints. At the same time, construct visual projection constraints based on visual point projection residuals and visual line projection residuals. S3: Based on the joint semantic information of dense semantic point cloud and image, identify whether there is perception degradation in the current environment of high-speed rail, and construct a comprehensive degradation scoring function based on multiple indicators to determine degradation. S4: After the degradation judgment result is that the degradation state has been entered, the system automatically switches to the redundant compensation factor modeling mechanism, introduces structural and motion information that does not depend on the external environment, maintains trajectory estimation in the degradation scenario, and constructs compensation constraints based on track structure geometric factors, IMU pre-integration factors and wheel and axle speed factors. S5: Introduce prior semantic constraints and large model semantic factor constraints to ensure the rationality of the trajectory; S6: Under the factor graph optimization framework, a global optimization objective function is constructed by integrating multiple constraints. Based on the confidence of various sensors, the weights of each modal factor in the factor graph optimization are dynamically adjusted. The optimal state estimate is obtained by iteratively solving the global optimization objective function, and the train positioning result is finally output.

2. The train localization method with multimodal fusion and semantic enhancement according to claim 1, characterized in that, In step S1, the multimodal sensors include lidar, millimeter-wave radar, camera, IMU, and wheel speedometer. The construction of the dense semantic point cloud includes the following steps: S111: Align all types of sensors to the IMU reference time; S112: Alignment of millimeter-wave radar point cloud with lidar coordinate system is achieved through extrinsic parameter transformation, and spatial completion is performed using kernel function weighted interpolation method to generate dense point cloud to enhance geometric integrity; S113: Generate a semantic label map using a semantic segmentation network for the current frame image, and project the LiDAR points onto the image plane to achieve alignment between semantic information and three-dimensional structure; S114: Project the dense point cloud onto the image plane through extrinsic parameter transformation, query the semantic category of the corresponding pixel in the semantic label map, and fill the corresponding 3D point with the semantic label to finally generate a dense semantic point cloud.

3. The train localization method with multimodal fusion and semantic enhancement according to claim 1, characterized in that, In step S1, dynamically estimating the confidence levels of various sensors includes the following steps: S121: Calculate the deviation between sensor observations and historical averages; S122: Construct a Bayesian model based on historical observation bias to estimate observation variance; S123: Observational variance obtained from estimation Calculate the confidence weights of the sensors The calculation formula is expressed as: ; in, For the first Sensor-like devices at all times Confidence weights To prevent division by zero of constants.

4. The train localization method with multimodal fusion and semantic enhancement according to claim 1, characterized in that, In step S2, the process of constructing visual projection constraints includes the following steps: S21: Constructing the visual point projection residual: Based on 3D spatial feature points and pixel observation points on the image plane, combined with camera pose parameters, intrinsic parameter matrices, and projection functions, the 3D spatial feature points are projected onto the image plane to obtain predicted pixel coordinates. By calculating the error between the predicted pixel coordinates and the actual observed coordinates of the pixel observation points, a visual point projection residual is constructed. ; S22: Constructing the visual line projection residual: By defining a set of 3D map line segments containing endpoint coordinates and a set of image observation line segments containing pixel endpoint coordinates, and combining camera pose parameters, intrinsic matrix, and perspective projection function, the endpoints of the 3D line segments are projected onto the image plane to obtain predicted pixel coordinates. Then, the endpoint residuals are constructed by calculating the Euclidean distance between the predicted and observed pixel coordinates, ultimately forming a visual line projection residual containing the observed line segment index and corresponding coordinate information. ; S23: Visual Projection Constraints Represented as: + .

5. The train localization method with multimodal fusion and semantic enhancement according to claim 1, characterized in that, S3 includes the following steps: S31: Point cloud structure degradation identification: S311: Utilize the dense point cloud in S1 to compute the normal entropy of the point cloud. With the structural dimension index, the normal vectors in the dense point cloud are projected onto a unit sphere, and multiple directional buckets are constructed by uniformly dividing the sphere or by predefined direction sets. S312: Statistically count the number of normal vectors in each bucket and normalize them to a probability distribution, then calculate the point cloud normal entropy. S32: Degradation Index (DDI) The structure dimension index (DDI) based on the eigenvalues ​​of the Hessian matrix is ​​used to evaluate the three-dimensional distribution integrity of point clouds, thereby determining whether there is spatial structure degradation in the point cloud of the current frame. S33: Evaluating image texture quality based on edge density; Canny edge detection is used to extract edge pixels from the image information captured by the camera, the number of edge pixels is counted, and the image edge density is calculated. ,definition This represents a threshold for determining whether an image quality meets the required standard; when If so, the image is deemed invalid; S34: Calculate the Comprehensive Degradation Score Function (CDS), expressed as: ; in, and These represent the time intervals of the LiDAR and the camera, respectively. Confidence weights; These are all task-related weighting coefficients used to adjust the contribution of each modality to the degradation score, and =1; and They represent and The maximum value is used for normalization operations; S35: Degradation is determined based on the calculated CDS (Comprehensive Degradation Score) value. When CDS≤0.3, it is determined to be a normal scene, and the LiDAR point cloud factor, visual projection factor, compensation factor, prior semantic factor and large model semantic factor are enabled. when If so, it is judged as moderate degradation, and the weights of visual projection and lidar point cloud factors are reduced; When CDS If the condition is found to be severely degraded, all sensing factors are disabled, and the track + IMU + wheel speed compensation mechanism is enabled, while only the compensation constraint factor is retained.

6. The train localization method with multimodal fusion and semantic enhancement according to claim 1, characterized in that, S4 specifically includes the following steps: S41: Constructing orbital structure geometry factors: The track structure geometric factor input is a track centerline model generated based on map, B-spline, or rule-based fitting methods. , represented as: ; ; in, The distance from the current pose point to the orbit centerline Lateral deviation; The estimated position at the current moment; for The arc length corresponding to the closest position on the line; S42: Introduce IMU pre-integration factor constraints and construct the IMU pre-integration factor; S43: Construct the wheel axle speed factor; The input information for the wheel axle velocity factor is: the velocity vector estimated from the current frame state. Current orbital direction unit vector That is, the tangent direction of the track curve at that point, and the scalar value of the wheel and axle velocity observed by the encoder. Define the residual function for: ; in, The actual speed of the train in the direction of track movement; based on the principle of consistency of speed direction, the component of the estimated speed in the track direction is aligned with the wheel speed gauge value; S44: Graph optimization ensemble for compensation modeling: The track geometry factor, IMU pre-integration factor, and wheel-axle velocity observation factor are uniformly incorporated into the graph optimization framework as compensation constraints between keyframe nodes, and an overall optimization objective function is constructed, expressed as: = ; in, This represents the set of state variables for all keyframes. The state variables for each keyframe include the pose, velocity, and IMU bias for each frame. For each type of factor, the dynamic weight coefficients are... For the first The factor in the th Residual function under class constraints.

7. The train localization method with multimodal fusion and semantic enhancement according to claim 1, characterized in that, In S5, the expression for the prior semantic constraint is: ; in, This is the speed limiting factor; For signal factor; Elevation factor; This is the turnout switching factor; Speed ​​limiting factor Constrained trajectory estimation velocity Do not exceed the speed limit of the current track section Its error term is defined as: ; Signal factor Based on the signal status set on the track, when the corresponding signal light is red, the image frame of that red light is forcibly constrained to ensure the estimated speed corresponding to that image frame. Approaching zero is defined as: ; in, This factor is activated when the traffic light at the current location is red. Elevation factor The residual is constructed by comparing the deviation between the estimated position and the corresponding elevation on the trajectory curve, and is expressed as: ; in, To estimate the height component of the location; For the track in position The elevation value at the location; Turnout switching factor When approaching a turnout, the track function is changed to a branch track to maintain continuity, as shown below: ; in, Let be the position vector of the orbital centerline in three-dimensional space; Let be the position vector of the centerline of the branch track in three-dimensional space.

8. The train localization method with multimodal fusion and semantic enhancement according to claim 1, characterized in that, In S5, the construction of semantic factor constraints for the large model includes the following steps: S51: Construct a semantic cue word set: A semantic prompting word library is constructed based on the high-speed rail scenario. Each prompting word is encoded into a high-dimensional semantic vector through a large model, represented as follows: ; in, Indicates the first The semantic embedding vectors of each prompt word; A text encoder for large models; This indicates all the prompt words contained in the prompt word library; S52: Use a large model to encode the image and the semantic vector of the prompt word to complete the image semantic embedding extraction; S53: Calculate semantic residuals: S531: Match the optimal semantic vector of the image and the prompt word using cosine similarity; S532: Retrieve the most matching prompt word from the prompt word library. , represented as: ; in, Let be the similarity between the semantic vector of the k-th frame image and the semantic embedding vector of the j-th cue word; S533: Define the semantic residual of the current frame as: ; in, Indicates the first The frame semantic residual term represents the difference between the current image semantics and the best prompt word. Indicates and The closest prompt word vector; S54: Semantic Factor Constraints for Large Models Represented as: ; in, These are the dynamic weighting coefficients for semantic factors, dynamically determined based on the image sharpness and camera confidence of the current frame. This represents the weighted semantic residual term.

9. The train localization method with multimodal fusion and semantic enhancement according to claim 5, characterized in that, In step S6, the global optimization objective function is expressed as: ; in, ; The pose of each frame; Position for each frame; For frame rate; Optimize the number of keyframes within the sliding window for the factor graph; , , , and These are constraints for multiple categories, namely lidar point cloud, visual projection, compensation, prior semantics, and large model semantic factors. and These are the confidence weights for LiDAR and vision, respectively. The weights of the compensation factors; and These are the weights of prior semantics and the semantics of the large model, respectively. The specific weights of each modal factor in the dynamically adjusted factor graph optimization are as follows: Based on the comprehensive degradation score (CDS), the environmental state is divided into three levels and mapped to a weight adjustment strategy, including: In normal scenarios, i.e., when CDS < 0.3: LiDAR and visual factors dominate. , Compensation factor assistance: Semantic factors enabled: , ; In mild degradation scenarios, i.e., when 0.3 ≤ CDS < 0.6: Reduce laser / visual weights: , Enhanced compensation factor: Semantic factors should be used with caution. , ; In severely degraded scenarios, i.e., when CDS ≥ 0.6, unreliable modes are disabled: , Completely dependent on compensation factors: Semantic factors disabled: , .

10. A train positioning system with multimodal fusion and semantic enhancement, used in the train positioning method with multimodal fusion and semantic enhancement as described in any one of claims 1-9, characterized in that, include: The multimodal sensor data acquisition and fusion module, radar point cloud and visual projection error construction module, degradation environment perception and judgment module, positioning compensation modeling module under degradation state, prior semantic constraint and large model semantic constraint modeling module, and factor graph optimization and state estimation module correspond one-to-one with S1-S6. The multimodal fusion and semantic enhancement positioning method completes S1-S6 through each module, and finally outputs high-precision train positioning results.

Citation Information

Patent Citations

  • Running train mileage positioning method and device based on scene reconstruction

    CN119963810A

  • Beidou PPP-RTK train full-scene positioning method, system and terminal

    CN119986744A