Dynamic optimization control system for driving path of rail trolley based on AI vision

Through the AI ​​vision-based dynamic optimization control system for the rail trolley's driving path, combined with the attention mechanism and fast path search method, the problems of blind spots and slow response in path selection of traditional methods in dynamic environments are solved, and the stable operation and intelligent decision-making of the rail trolley in complex environments are achieved.

CN120704344AActive Publication Date: 2025-09-26JIANGSU FLYING SHUTTLE INTELLIGENT CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511056254.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-09-26
Estimated Expiration
2045-07-30

AI Technical Summary

Technical Problem

Traditional rail vehicle path control methods are difficult to cope with dynamic environmental changes, such as temporary obstacle interference, sudden changes in path structure, and multiple vehicle intersections. These problems lead to blind spots in path selection and slow response, and even cause risks such as collisions and path interruptions. In addition, existing methods find it difficult to enhance local diversity while ensuring global optimization.

Method used

A dynamic optimization control system for the rail vehicle's travel path based on AI vision is adopted, combining the attention mechanism with a fast path search method. Through data acquisition, preprocessing, environmental modeling, model optimization and control execution modules, the rail vehicle's direction perception and path planning are realized. The three-motion attention mechanism and direction mask suppression unit are used to improve the ability of environmental understanding and feature extraction, and the A* algorithm is combined to search for the optimal path.

Benefits of technology

It effectively improves the directional environment understanding and path perception capabilities of the rail vehicle in complex dynamic environments, shortens the response time to dynamic obstacles, avoids local optimality, ensures the smoothness and safety of driving, and reduces trajectory tracking errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120704344A_ABST
    Figure CN120704344A_ABST
Patent Text Reader

Abstract

The invention relates to the field of rail traffic intelligent driving, and discloses an AI-vision-based dynamic optimization control system for a running path of a rail trolley, and the system comprises a data collection module which is used for obtaining a rail environment image, the motion state data of the rail trolley and auxiliary information in real time; the data preprocessing module is used for carrying out standardization processing on the acquired data; the environment modeling module is used for constructing a dynamic track environment model so as to predict the future position of the track trolley; the model optimization module is used for extracting direction perception features through a three-motion attention mechanism in a diagonal direction, a forward direction and a backward direction, generating a fusion feature graph, extracting path key points based on a rapid path search method of attention guidance, and then searching an optimal path among the key points by utilizing a graph search algorithm; and the control execution module is used for generating a control signal according to the optimal path and driving the rail trolley to run. According to the invention, stable operation and intelligent decision making of the small rail car in a complex dynamic environment are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of rail transit intelligent driving technology, and in particular to an AI vision-based dynamic optimization control system for a rail vehicle's travel path. Background Art

[0002] With the development of artificial intelligence and computer vision technologies, rail transit systems are rapidly evolving toward greater intelligence and autonomy. Rail trolleys serve as crucial automated execution units in scenarios such as industrial logistics, tunnel inspections, and intelligent transportation. The planning and control of their operating paths directly impact the overall efficiency and safety of the system.

[0003] Traditional rail vehicle path control methods often rely on static maps and rule-based control logic, making them incapable of handling dynamic environmental changes, such as temporary obstacles, sudden changes in path structure, and multiple vehicle intersections. Furthermore, existing methods often overlook the rail vehicle's ability to perceive the directional structure of the environment during travel, leading to blind spots in path selection, slow response, and even the risk of collisions and path interruptions.

[0004] In addition, in the path search algorithm, how to enhance local diversity while ensuring global optimality and avoid falling into early convergence and local optimal solutions is also a key challenge in the field of intelligent path optimization.

[0005] Therefore, there is an urgent need for an efficient path optimization and control mechanism with direction perception capability and the ability to integrate visual features to achieve stable operation and intelligent decision-making of rail vehicles in complex dynamic environments. Summary of the Invention

[0006] The present invention aims to provide a dynamic optimization control system for the travel path of a rail vehicle based on AI vision. The system comprehensively introduces an attention mechanism and a fast path search method based on attention guidance to achieve collaborative optimization in the path perception, planning, and control stages. It solves the problems that traditional methods are difficult to cope with dynamic environmental changes, such as temporary obstacle interference, sudden path structure changes, and multiple vehicle intersections. In addition, existing methods have blind spots in path selection, slow response, and even risks of collision and path interruption, thereby achieving stable operation and intelligent decision-making of rail vehicles in complex dynamic environments.

[0007] To achieve the above objectives, the following technical solutions are adopted:

[0008] A dynamic optimization control system for a rail vehicle's travel path based on AI vision, comprising:

[0009] Data acquisition module, used to obtain real-time track environment images, track vehicle motion status data, and auxiliary environmental information including lighting conditions and track markings;

[0010] The data preprocessing module standardizes the collected data and outputs the standardized sensor data;

[0011] The environment modeling module is used to build a dynamic track environment model with environment perception and behavior prediction interfaces based on the standardized number of sensors to predict the future position of the track vehicle;

[0012] The model optimization module extracts direction-aware features and generates a fused feature map through a three-motion attention mechanism in the diagonal, forward, and backward directions. It also extracts key points of the path based on an attention-guided fast path search method and then uses a graph search algorithm to search for the optimal path between the key points.

[0013] The control execution module generates a control signal according to the optimal path and drives the rail vehicle to travel.

[0014] Furthermore, the environment modeling module includes:

[0015] An environmental perception unit, configured to divide the track area into a grid matrix, each grid containing environmental attribute identifiers, including obstacles, track markings, and hazardous conditions;

[0016] Trajectory prediction unit, used to predict the position of the next K time steps based on the current position and velocity vector through the kinematic model;

[0017] The behavior prediction interface unit is used to integrate environmental and status data and output structured information, including: static obstacle set, dangerous area set, moving obstacle parameters, rail vehicle status and load factor.

[0018] Furthermore, the model optimization module includes:

[0019] The three-direction modeling unit is used to model the diagonal direction to detect track oblique trends and obstacles; model the forward direction to enhance the analysis of the path accessibility ahead; and model the backward direction to identify rear obstacles and path closure characteristics.

[0020] The direction mask suppression unit sets mask matrices for the three motion directions, including diagonal direction mask matrix, forward mask matrix and backward mask matrix, to suppress the attention scores of irrelevant directions.

[0021] Furthermore, the direction mask suppression unit specifically includes:

[0022] The diagonal mask matrix is ​​set to: 0 when the feature point positions i and j are different, and negative infinity when they are the same; the reverse mask matrix is ​​set to: 0 when the feature point position i is greater than j, otherwise negative infinity; the forward mask matrix is ​​set to: 0 when the feature point position i is less than j, otherwise negative infinity.

[0023] Furthermore, the model optimization module also includes:

[0024] The three-directional feature pair scoring unit is used to apply the corresponding mask matrix to any two feature points in the track image according to their positional relationship. The directional awareness transformation is performed on each of them through a learnable mapping matrix. The transformed feature vectors are added together and a learnable bias vector is added. The nonlinear interaction between features is enhanced through the hyperbolic tangent activation function. The mask values ​​of the corresponding directions are superimposed to generate a directional attention score.

[0025] The probability distribution calculation unit is used to calculate the attention scores of the three motion directions for each feature channel respectively, and obtain the weighted probability distribution of the three directions through Softmax function normalization.

[0026] Furthermore, the model optimization module also includes:

[0027] The feature fusion unit is used to weightedly aggregate adjacent features in corresponding directions according to the weighted probability distribution of the three directions to generate a direction-aware feature representation; a gating mechanism is used to weightedly combine the original visual features and the direction-aware features to generate a fused feature map representation; wherein the gating mechanism learns the degree of feature dependence through the Sigmoid function and adaptively adjusts the fusion ratio of the original information and the direction information.

[0028] Furthermore, the fast path search method based on attention guidance includes:

[0029] Key point extraction step: Extract high-attention areas from the fused feature map output by the three-motion attention mechanism as key points, and add the current position and target position of the track car to form a key point set;

[0030] Path graph construction steps: Use key points as nodes, connect adjacent nodes to form edges, and assign a weight to each edge. The weight is determined by the distance between nodes, the attention value at the node, and the risk value of the area passed by the edge;

[0031] Graph search step: On the constructed path graph, use the A* search algorithm to find the optimal path from the starting point to the end point.

[0032] Furthermore, the step of extracting key points specifically includes:

[0033] Perform channel compression on the fused feature map, compressing it into a single-channel two-dimensional attention map by taking the maximum value or average value;

[0034] Based on threshold segmentation, the positions greater than or equal to τ in the attention map are marked as candidate key points;

[0035] Perform non-maximum suppression on candidate key points to obtain a set of sparse key points;

[0036] The current position and target position of the car are also added to the key point set, and finally the key points of the track car are extracted.

[0037] Furthermore, the control execution module includes:

[0038] A path smoothing unit is used to smooth the optimal path obtained by searching using cubic spline interpolation or B-spline curve to obtain a smooth path;

[0039] The speed planning unit is used to plan the driving speed according to the curvature of each point on the smooth path and the corresponding attention value;

[0040] a control signal generating unit for generating steering and acceleration control instructions based on the smoothed path and the planned driving speed;

[0041] The real-time feedback unit is used to collect the deviation between the actual position of the rail vehicle and the nearest point on the currently planned path. When the trigger condition is met, the re-planning instruction is triggered.

[0042] Furthermore, the data preprocessing module performs the following operations: image denoising, brightness equalization, and edge enhancement operations.

[0043] Compared with the prior art, the present invention achieves the following beneficial effects:

[0044] 1. This invention is based on the Three-Motion Attention Mechanism (TMA): it simulates the human visual system's ability to selectively focus on dynamic features in different directions, constructs attention perception models for three motion directions: diagonal, forward, and backward, and guides the rail vehicle to focus on the "forward path," "potential return trajectory," and "track oblique structure" at the visual level, effectively improving the ability to understand the directional environment and extract features, and solving the directional perception blind spot problem of traditional methods.

[0045] 2. This paper proposes a directional mask suppression mechanism: a mask matrix is ​​introduced for multi-directional attention, by setting the attention score of the irrelevant direction to ,In the Softmax operation, invalid directions are suppressed, so that attention is more focused on the track area that is consistent with or related to the direction of vehicle movement, improving the effectiveness and selectivity of path perception.

[0046] 3. The present invention proposes a three-directional feature pair scoring method: constructing a directional feature scoring function, performing directional scoring on track feature points in any pair of images, capturing their spatial correlation in the current direction, and providing basic directional score support for subsequent probability calculation and feature fusion.

[0047] 4. The present invention proposes to use a gating mechanism to achieve feature fusion: the original visual features and the directional enhancement features are weightedly combined, an adaptive gating structure is introduced, and the degree of dependence on the original information and directional information in the learning area is learned through the Sigmoid function to achieve a more refined feature fusion strategy and enhance the environmental adaptability of the vehicle's path judgment.

[0048] 5. The present invention adopts attention-guided fast path search (AGFPS), extracts key points based on high-attention areas to construct a sparse path graph, combines the A* algorithm with a directional heuristic function, shortens the time consumption of long path planning, significantly improves the response speed to dynamic obstacles, and avoids local optimality, thus solving the dynamic obstacle response delay problem of traditional methods.

[0049] 6. The present invention combines dynamic replanning with multi-factor adaptive strategies to ensure driving stability and safety. It integrates path curvature, attention risk value and load factor to dynamically adjust the speed, uses pure tracking steering PID speed tracking to reduce trajectory tracking errors, and achieves high-precision and stable control under complex working conditions.

[0050] It should be understood that the contents described in the summary of the invention are not intended to limit the key or important features of the embodiments of the present invention, nor are they intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] The above and other features, advantages and aspects of the embodiments of the present invention will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. The accompanying drawings are provided for a better understanding of the present invention and do not constitute a limitation of the present invention. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, among which:

[0052] Figure 1 This is a module schematic diagram of a dynamic optimization control system for a rail vehicle's travel path based on AI vision according to an embodiment of the present invention;

[0053] Figure 2 This is a flow chart of a model optimization module of a rail trolley driving path dynamic optimization control system based on AI vision in an embodiment of the present invention. DETAILED DESCRIPTION

[0054] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0055] In this document, the term "and / or" simply describes a relationship between related objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " in this document generally indicates that the related objects are in an "or" relationship.

[0056] Figure 1 This is a module diagram of a dynamic optimization control system for a rail vehicle's travel path based on AI vision according to an embodiment of the present invention. Figure 1 As shown, a dynamic optimization control system 100 for a rail vehicle's travel path based on AI vision includes:

[0057] The data acquisition module 110 is used to obtain track environment images, vehicle motion status data and auxiliary information in real time;

[0058] In the rail trolley operation system, data collection is the basic link to achieve dynamic path optimization. This system forms a multi-angle, all-round visual perception network by deploying high-definition wide-angle cameras at key nodes along the track (such as curves, intersections, areas with frequent obstacles, etc.), and integrating forward and downward dual camera modules on the trolley body. Preferably, high-definition wide-angle cameras are deployed at track curves / intersections / slope change points with a spacing of ≤5 meters, the downward module is 0.8-1.2 meters vertically from the track, and the forward module has an inclination angle of 15°±3°; visual positioning uses ORB-SLAM3 fused with IMU data, and the positioning error is ≤2cm. The following information is collected in real time by the high-definition wide-angle camera and the dual camera module:

[0059] (1) Track environment images: including track structure, obstacles, traffic conditions (such as water accumulation, damage), etc.

[0060] (2) Car motion state data: including the car's position on the track (using visual positioning fusion IMU), real-time speed, acceleration, etc.;

[0061] (3) Auxiliary information: such as lighting changes, interference from people or other vehicles.

[0062] These raw data are pre-processed by local edge processing modules (such as Jetson or RK3588 platforms) and then uploaded to the central control unit, providing a high-quality, low-latency data source for subsequent processing.

[0063] The data preprocessing module 120 performs standardization on the collected data and constructs a dynamic track environment model with environment perception and behavior prediction interfaces based on the standardized data to predict the future position of the track vehicle;

[0064] The data preprocessing module 120 first standardizes the collected original images and sensor data through the image preprocessing module, including image denoising (such as based on wavelet transform or Gaussian filtering), brightness equalization, edge enhancement and other operations to improve the accuracy of environmental feature recognition.

[0065] Through the above-mentioned data cleaning, standardization and preliminary feature extraction (such as image denoising, brightness equalization, and edge enhancement), the module outputs standardized sensor data.

[0066] Environmental modeling module 130: Constructs a dynamic track environment model with environmental perception and behavior prediction interfaces based on preprocessed data;

[0067] The environmental modeling module 130 system integrates the real-time position information and velocity vector data of the rail trolley to construct a dynamic rail environment model. The model has the following functions: environmental perception: identifying obstacles, road markings, temporary working condition changes, etc. in the track; synchronous modeling of the rail trolley state: synchronously superimposing the motion trajectory of the rail trolley in the environment to predict its future position; behavior prediction interface: providing input to the model optimization module 140, including dynamic information such as obstacle location, movement trend, and rail trolley load changes.

[0068] The dynamic track environment model specifically includes:

[0069] An environment sensing unit 131 is configured to divide the track area into a grid matrix, where each grid contains an environmental attribute identifier, including obstacles, track markings, and hazardous conditions;

[0070] Specifically, the track area is divided into an M×N grid matrix, and the environmental attributes within each grid are identified by the target detection algorithm to generate the environmental state matrix ,in:

[0071] When there is no obstacle in the grid (i', j'), that is, in the safe area (no mark), =0;

[0072] When there is a static obstacle in grid (i', j'), =1;

[0073] When there is a track mark (such as a stop line) at grid (i', j'), =2;

[0074] When there is a temporary dangerous condition (water accumulation or oil pollution) in the grid (i', j'), =3.

[0075] Grid coordinate transformation satisfies: ,

[0076] Among them, the grid coordinate (i', j') represents the grid index after the track area is divided, i' is the row index, and j' is the column index; (x, y) is the world coordinate system coordinate; ( , ) is the origin of the track area; r is the grid resolution, and the grid resolution r satisfies 0.05m≤r≤0.2m. Target detection is implemented using a lightweight convolutional neural network.

[0077] The trajectory prediction unit 132 is used to predict the future trajectory of the rail vehicle, including: calculating the position of the rail vehicle in the next K time steps based on the current position and velocity vector of the rail vehicle through a kinematic model;

[0078] Specifically, based on the current position of the rail trolley ( , ) and the velocity vector ( , ) Calculate the position of the next K time steps through the kinematic model, including:

[0079] ,k=1,2,…,K, the prediction step number K satisfies 3≤K≤5, and the prediction period Δt is consistent with the system control period.

[0080] The behavior prediction interface unit 133 is used to integrate environmental and state data and output structured information to subsequent modules (model optimization module 140 and control execution module 150), including: static obstacle set, dangerous area set, mobile obstacle parameters, rail vehicle state and load factor.

[0081] Static obstacle coordinate set ;

[0082] Danger zone coordinate set ;

[0083] Mobile obstacle parameter set , where the trend vector ;

[0084] Specifically, the moving obstacle trend vector is calculated by the inter-frame difference method: , , t is the current control cycle, t-1 is the previous control cycle, ( ) represents the world coordinate position of the moving obstacle in the current control cycle t (obtained by grid coordinate mapping conversion), ( ) represents the world coordinate position of the moving obstacle in the previous control cycle t - 1 (obtained through grid coordinate mapping conversion). This differential calculation reflects the displacement of the moving obstacle within a control cycle Δt (optionally 0.1s).

[0085] The current state of the car D = {current position (i', j'), predicted path {( , }};

[0086] Load Factor , is the load of the current control cycle, is the load of the previous control cycle t -1.

[0087] Among them, the output of this module (environmental state matrix and behavior prediction interface data) are input to the model optimization module 140 to determine the risk value of the edge, the current state of the car (starting point) and the target position (end point), and to adjust the direction mask matrix The load factor λ output by the prediction interface unit is input to the control execution module 150 to participate in speed planning.

[0088] The model optimization module 140 extracts direction-aware features and generates a fused feature map through a three-motion attention mechanism in the diagonal, forward, and backward directions, extracts path key points based on an attention-guided fast path search method, and then uses a graph search algorithm to search for the optimal path between the key points;

[0089] The model optimization module 140 is mainly used to implement model training and optimization.

[0090] To achieve dynamic optimization of the rail vehicle's travel path, the model optimization module 140 integrates the three-motion attention mechanism and the attention-guided fast path search method. Through the synergy of feature enhancement and path search, it builds a model system with environmental perception and intelligent decision-making capabilities. It is divided into the following two core links:

[0091] (1) Introducing the Three-Motion Attention Mechanism

[0092] To enhance the rail vehicle's ability to perceive key path information in complex environments, this system introduces a triple motion attention mechanism (TMA). This mechanism simulates the human visual system's perception of dynamic features in different directions, guiding the model to focus on information such as the "forward path," "potential return trajectory," and "track oblique structure," thereby enhancing the directionality of track environment perception and improving feature discrimination capabilities.

[0093] Figure 2 This is a flow chart of a model optimization module of a dynamic optimization control system for a rail vehicle travel path based on AI vision according to an embodiment of the present invention. Figure 2 As shown, the model optimization module 140 includes:

[0094] The three-direction modeling unit 141 is used to model the diagonal direction to detect track oblique trends and obstacles; model the forward direction to enhance the analysis of the forward path feasibility; and model the backward direction to identify rear obstacles and path closure characteristics.

[0095] Specifically, the process of modeling three motion directions by the three-motion direction modeling unit 141 is as follows:

[0096] Diagonal direction (d=1): Focuses on the oblique direction of the track or obstacle, suitable for detecting obstacles such as slopes, curves, or oblique intrusions;

[0097] Forward direction (d=2): simulates the normal driving direction of the car (from upper left to lower right), used to enhance the forward path connectivity analysis;

[0098] Backward direction (d=3): Focuses on the retreat path or abnormal trajectory to improve the ability to identify rear obstacles or path closure features.

[0099] The direction mask suppression unit 142 sets mask matrices for the three motion directions respectively, including a diagonal direction mask matrix, a forward mask matrix, and a backward mask matrix, to suppress the attention scores of irrelevant directions.

[0100] Introducing directional mask suppression mechanism: by introducing mask matrix Attention suppression is performed on irrelevant directions, and the suppression item is set to , so that the score in this direction tends to 0 in the Softmax operation:

[0101]

[0102] Where i, j are the feature point indices at different locations in the track image, used to represent the position pairs in the grid map under the visual perception of the track vehicle); : Diagonal mask matrix, used to shield the invalid attention of the feature point itself; : Inverse mask matrix, used to retain feature information from before the current position (historical trajectory); : Forward mask matrix, used to retain feature information from after the current position (estimated path); : Indicates that the direction is valid and attention is retained; : Indicates that the direction is invalid and attention is suppressed to 0; : Represents the mask matrix corresponding to the motion direction (can be 、 or ), which is used to force the attention weights of irrelevant directions to be zero in Softmax.

[0103] In the direction mask suppression unit 142, the output of the behavior prediction interface unit is used to adjust the direction mask matrix The main use here is to adjust the mask using obstacle information (static and moving) so that the direction mask in the obstacle area is suppressed (set to ) or enhanced (set to 0). Specifically, the mask is adjusted using the obstacle information (S, H, Y) The specific logic is:

[0104] For any position pair (i, j), if the grid (i', j') corresponding to position i or position j belongs to the static obstacle set S or the dangerous area set H, then Set to -∞ for all directions d (or a specific direction);

[0105] For the current position (i', j') of the moving obstacle Y, the position i or j is also masked to the position of this grid. Set to -∞. In addition, according to the trend vector of the moving obstacle Predict the grid coordinates to which it may move in the short future (e.g., the next control cycle or path search estimation time) , ),in , , Represents the prediction duration. These predicted grid coordinates are also considered as temporary danger areas, and the position mask corresponding to these predicted grids at position i or j is set to -∞.

[0106] When applying obstacle information to adjust the direction mask, it is necessary to map the feature map position indexes i and j to their corresponding physical space areas. Using the known mapping relationship between the feature map space coordinate system and the physical world coordinate system, calculate the physical coordinate range or representative point coordinates corresponding to positions i and j ( , Then, according to the grid coordinate conversion rules (see the environment perception unit 131), these physical coordinates are converted into corresponding grid coordinates ( , Finally, check whether these grid coordinates belong to the static obstacle set S or the danger zone set H or the mobile obstacle set Y.

[0107] By setting mask matrices for the three motion directions, namely the diagonal direction mask matrix, the forward mask matrix, and the backward mask matrix, the information of irrelevant directions in the track image is shielded, so that the attention mechanism can focus more on the direction related to the current motion of the track car.

[0108] A three-directional feature pair scoring unit 143 is used to apply the corresponding mask matrix to any two feature points in the track image according to their positional relationship, perform direction-aware transformation on each of them using a learnable mapping matrix, add the transformed feature vectors and add a learnable bias vector, enhance the nonlinear interaction between features using a hyperbolic tangent activation function, and superimpose the mask values ​​of the corresponding directions to generate a directional attention score;

[0109] For any pair of feature points in the track image and , its attention direction correlation score is:

[0110] in, : represents the directional attention score of the i-th and j-th positions in the track image, which is used to measure whether the track car should focus on the traffic information in this direction on the track. : The visual feature vectors at the i-th and j-th positions of the track image, representing the spatial point state perceived by the track car; : A learnable direction-aware mapping matrix is ​​used to transform the direction encoding of the start and end points respectively; : A learnable bias vector used to adjust the nonlinear expression capability after feature fusion; : Scaling factor, which is an empirical parameter with a value range of 0.1 ≤ q ≤ 1.0, and a default value of q=0.5 to prevent The nonlinear function tends to saturate when the input value is too large, which is used to improve the stability of gradient propagation; : Hyperbolic tangent activation function, used to enhance the nonlinear interaction between feature pairs; : The value of the corresponding position of the mask matrix is ​​used to control whether to retain the attention score of the pair of features; : A vector (or scalar) of all 1s used to mask The weighting is applied to the score in an explicit form, thereby completing the attention masking of irrelevant directions.

[0111] This formula is used to directional score the degree of visual association between any two points in the track image and is the key to generating attention weights.

[0112] The probability distribution calculation unit 144 is used to calculate the attention scores of the three motion directions for each feature channel respectively, obtain the attention probability distribution of the three directions through the Softmax function normalization, and set the attention weight of the invalid direction to zero.

[0113] Specifically, the calculation process of the probability distribution calculation unit 144 includes the following:

[0114] The scores calculated for the three movement directions , needs to be standardized into probability form:

[0115]

[0116] : On the mth feature channel, the i-th and j-th feature points are in the direction The attention score on the direction is used to measure whether the rail vehicle should give priority to the traffic characteristics in this direction; : Movement direction type index, d=1 indicates oblique (diagonal) direction, d=2 indicates forward direction, d=3 indicates backward direction; : The learnable mapping matrix of the mth feature dimension in direction d is used to transform the source point With the target point Channel characteristics, simulating directional correlation; : The bias vector (length C) of the mth channel in direction d, used to enhance the directional sensitivity of feature fusion; : Scaling factor to avoid The activation function experiences gradient vanishing when the input amplitude is too large, thus improving learning stability; : The mask matrix in the dth direction, used to suppress illegal or invalid attention values ​​in this direction (for example, no passage exists) before the Softmax operation. It is a scalar value calculated based on the position pair (i, j) and direction d (mainly adjusted based on the output of the environment model); : A vector of all 1s of dimension m, used to match the mask Element-by-element multiplication is performed to force the attention value of invalid directions to zero.

[0117] The attention scores of the three motion directions are input into Softmax to calculate the dependency probability:

[0118]

[0119] : represents the weighted probability of aggregating information from direction d for feature channel m when fusing the features at position j. This value represents the strong or weak assessment of the possibility of the rail car passing in this direction, and its value range is ; : exponential function, used to convert the original score into a positive value and amplify the score difference to enhance the contrast of direction selection; : used to normalize all direction scores so that the results are expressed in probability form; d represents the direction index currently being calculated. Indicates the direction variable index traversed in the Softmax denominator, which is used to calculate the sum of the scores of all directions.

[0120] Among them, the diagonal direction probability represents the attention to the oblique structure, the forward direction probability represents the attention to the front path, and the backward direction probability represents the attention to the backward trajectory.

[0121] The feature fusion unit 145 is used to weightedly aggregate adjacent features of the corresponding directions according to the attention probability distribution of the three directions to generate a direction-aware feature representation; a gating mechanism is used to weightedly combine the original visual features and the direction-aware features to generate a fused feature map representation; wherein the gating mechanism learns the degree of feature dependence through the Sigmoid function and adaptively adjusts the fusion ratio of the original information and the direction information.

[0122] Specifically, for each feature point, its adjacent features from all directions are weighted and aggregated according to the directional attention probability to generate a direction-aware representation. :

[0123]

[0124] : Indicates the feature point from the dth direction Direction-aware features are used to capture the structural and semantic information of specific directions; : represents the directional perception representation after the fusion of the j-th track feature point, which integrates the weighted information in three directions and helps the rail vehicle to make a comprehensive judgment on the "upstream and downstream path structure" and "diagonal channel potential" at the current node.

[0125] Final The integration of structural information from three directions of this point helps the rail vehicle to dynamically judge the upstream and downstream and oblique channel structures of the path.

[0126] Furthermore, a gating mechanism is used to achieve a weighted combination of the original features and the directional fusion features:

[0127]

[0128] : The fusion feature map representation integrates the original visual features and directional structure features as the input of the track vehicle path decision, improving its dynamic response ability to the complex track environment. : The original local feature representation in the track image is generated by the early feature extraction network, which mainly reflects the basic visual features of the track point where the track car is located. : A three-motion direction feature representation that integrates forward, backward, and diagonal direction structural information is used to enhance the rail vehicle's perception of the upstream and downstream track environment and channel structure. : Directional feature mapping matrix in the gating mechanism, used to The feature channels of the dataset are weighted transformed to adapt to the fusion scene. : The original feature map matrix in the gating mechanism, used to A mapping transformation is performed to align and fuse with the direction-enhanced features. : Gating bias term, which controls the offset of the gate output and adjusts the feature activation threshold during the fusion process. : Sigmoid activation function, compressing the feature weighted results to The interval, as the gating vector of the fusion weight, is used to adaptively control the fusion ratio of the original features and the direction-enhanced features. : Fuse gated vectors to automatically learn which areas in the track scene rely more on original visual perception and which areas rely more on direction-enhanced perception. : Hadamard product (element-wise multiplication) is used to gate and weight feature channels. : The weighting factor of the directional feature, which is opposite to the gate vector and is used to adjust the directional perception feature The degree of participation; F is the gate vector, that is .

[0129] By introducing the three-motion attention mechanism, the system is able to explicitly model the structural directionality of the track environment at the visual level, providing a more discriminative and direction-aware feature basis for the subsequent path planning stage.

[0130] 2. Attention-guided Fast Path Search (AGFPS) to search for the optimal path

[0131] Fusion feature map output by the three-motion attention mechanism ,

[0132] Input: fused feature map (The size is H×W×C, but it is usually averaged or maximized in the channel dimension to obtain a two-dimensional attention map with a size of H×W), as well as the behavior prediction interface data output by the dynamic track environment model (including obstacle location, current state of the track car, target location, etc.).

[0133] The process of searching for the optimal path includes the following steps:

[0134] S1: Key point extraction step: extract high attention areas from the fusion feature map output by the three-motion attention mechanism as key points, and add the current position and target position of the track car to form a key point set;

[0135] In step S1, the system performs high-confidence path node extraction, specifically including:

[0136] S11: Perform channel compression on the fused feature map and compress it into a single-channel two-dimensional attention map by taking the maximum value or average value;

[0137] Specifically, the three-dimensional feature map (H×W×C) is compressed into a two-dimensional attention heat map (H×W). Preferably, channel dimension maximum pooling is used: .

[0138] S12: Based on threshold segmentation, the positions greater than or equal to τ in the attention map are marked as candidate key points;

[0139] Set a dynamic threshold τ (usually 0.5-0.7) and mark the positions greater than or equal to τ in the attention map as candidate key points, thereby screening out high attention areas:

[0140] .in, is the compressed two-dimensional attention map The pixel coordinates on .

[0141] S13: Perform non-maximum suppression (NMS) based on spatial distance on the candidate key points to obtain a set of sparse key points;

[0142] Non-maximum suppression (NMS): To avoid the adjacent points being too dense, non-maximum suppression is performed on the candidate key points. The specific operation is: set a suppression radius (unit: pixel), and the grid resolution r and the desired physical spacing ∆min are used to determine the suppression radius of NMS For each candidate point, In the neighborhood, only The point with the largest value suppresses other points. The sparse key point set obtained in this way , u=1,2,…,N; u is the index of the key point, and the spacing in physical space is at least about ∆min.

[0143] S14: The current position (starting point) and target position (end point) of the rail trolley are also added to the key point set, and the key points of the rail trolley are finally extracted.

[0144] In step S14, the current position of the rail vehicle is and target location Force the addition of key point sets to achieve key point enhancement.

[0145] Add the starting point and the end point: Add the current position (starting point) and the target position (end point) of the track car to the key point set to obtain the key point set (including the starting and ending points).

[0146] S15: Key point feature index mapping: For key point set Each key point in , using the known physical coordinates ( ) and the mapping relationship between the feature map space coordinate system (this relationship is determined by the structure or calibration parameters of the feature extraction network), and calculate its The feature map of The corresponding position index on (or coordinates( , )). Store the index Used to subsequently access the feature vector corresponding to the key point .

[0147] S2: Path graph construction step: take key points as nodes, connect adjacent nodes to form edges, and assign a weight to each edge. The weight is determined by the distance between nodes, the attention value at the node, and the risk value of the area passed by the edge;

[0148] Construct a weighted directed graph G=(V,E) as the basis for path search:

[0149] Node Set: ;

[0150] Edge generation rule: for each node , connecting all nodes in its spatial neighborhood The connection radius R is dynamically adjusted according to the complexity of the environment (3-8m).

[0151] The above process means that for each pair of nodes ( , ), if the Euclidean distance between them If the radius is smaller than the preset radius R (e.g. R = 5 meters), an edge is established between the two nodes. That is, for each node , connect nodes whose Euclidean distance is less than radius R (v≠u), forming an edge .

[0152] For each key point ,Pick The value at this position is ,Right now , (x', y') is the key point Position coordinates on the two-dimensional attention map. Among them, the key points Especially in The pixel coordinates (x', y') on the image are defined. The physical position (world coordinates) corresponding to this key point needs to be obtained through coordinate transformation. Assume There is a mapping relationship between the pixel coordinate system of the original perception image / grid map (for example, determined by the camera's internal and external parameters or the generation parameters of the grid map). This mapping relationship is part of the system calibration. For simplicity, the scheme can assume that this mapping relationship is known. The key point The physical coordinates ( ) can be obtained by looking up a table or calculating.

[0153] Feature map coordinates The corresponding attention value: .

[0154] Edge weight calculation:

[0155] in, :Every pair of nodes and The Euclidean distance between :node Normalized attention value at (obtained from the 2D attention map); :node The normalized attention value at ℓ (obtained from the 2D attention map). : represents an edge The maximum value of the risk level corresponding to all the grids passed (obtained by discretization), specifically: the risk level is based on the environmental state matrix Mapping: If =1 (static obstacle), then the risk level = (e.g. 1.0); if >=3 (temporary dangerous working condition), then the risk level = (e.g. 0.6); others ( =0 or 2), then the risk level = (e.g. 0.0)), and then take the maximum value of all grid risks after discretization of the entire edge as α, β, γ: weight coefficients, satisfying α+β+γ=1. The specific values ​​can be adjusted according to the actual scenario, for example, α=0.5, β=0.3, γ=0.2. Weight The smaller it is, the better the edge is (because of short distance, high attention value and low risk).

[0156] In addition, if the edge Direction (from arrive If the angle between the vector (of the vehicle) and the current velocity direction of the vehicle is less than 45 degrees, the weight is discounted: .

[0157] S3: Graph search step: On the constructed path graph, use the A* search algorithm to find the optimal path from the starting point to the end point.

[0158] Specifically: Use the A* algorithm to search for the optimal path (i.e. the path with the least cost) from the starting point to the end point in the constructed path graph.

[0159] (1) Cost function: , where: g(n) is the actual cost from the starting point start to the node n (that is, the sum of the weights of all edges on the path), , w(e) is the weight of edge e, Path(start→n) is the path from the starting point start to the node n. h(n) is the heuristic function, . dist(n,end) is the Euclidean distance from node n to end point end. is the average attention value of the grid in the 3×3 neighborhood around node n, from the fusion attention map φ is the attention weight coefficient, ranging from [0.1, 0.3], with 0.2 being the preferred value.

[0160] (2) Direction-guided node expansion: (2.1) Expansion priority: Arrange the nodes in the open set OpenSet in ascending order of f(n), and prioritize expanding the node with the smallest f(n). (2.2) Attention-driven direction bias: For the neighbor v of the current node n, calculate the direction angle from the current node n to each neighbor node v. (Relative to the current velocity direction of the current node n ): , in is a vector pointing from n to v. Weight adjustment condition: If Aligned with the high attention direction, that is, the neighbor v satisfies ≥ (Attention threshold =0.6), then temporarily adjust the edge weight: , weight adjustment coefficient ξ=0.5, after adjustment, use Participate in the calculation of g(n). The weight of the high attention area is reduced, and its priority is increased. (2.3) Invalid direction suppression: If v is located in the mask matrix The marked inhibition region ( =-∞), then the neighbor is skipped directly and not added to the open set.

[0161] (3) Search termination conditions: Success: reaching the target node (the current node is the end point); Failure: the open set is empty or the maximum number of iterations (for example, 1000) is exceeded.

[0162] (4) Path backtracking: trace the parent node from the end point end to the starting point s.

[0163] (5) Generate the optimal path: trace the parent node pointer from the end point back to the starting point to generate a series of ordered key point sequences: .

[0164] The control execution module 150 generates a control signal according to the optimal path and drives the rail vehicle to travel.

[0165] After completing the path planning, the system needs to generate a control signal based on the optimal path, and then drive the rail vehicle to move accurately according to the control signal. Specifically, the control execution module 150 includes:

[0166] The path smoothing unit 151 is used to smooth the searched optimal path using cubic spline interpolation or B-spline curves to obtain a smooth path. Since the searched path is a broken line composed of a series of key points, the path needs to be smoothed to ensure the smooth travel of the track vehicle. Here, cubic spline interpolation (or B-spline curve) is used to fit the broken line into a smooth curve to obtain a smooth path C(s), where s is the arc length parameter. The control point spacing is adaptive to the path curvature to meet the maximum curvature constraint. .

[0167] Speed ​​planning unit 152: On the smooth path C(s), the speed is planned based on the curvature of each point on the smooth path and the corresponding attention value (safety level). Specifically, the following steps are included:

[0168] Calculate the curvature κ(s) of each point on the smooth path C(s) (which can be obtained by differentiation) and add curvature constraints: , where δ is the load sensitivity coefficient; is the upper limit of the dynamic curvature corresponding to the arc length parameter s, which decreases when the load λ increases.

[0169] Calculate the travel speed at this point based on the curvature :

[0170] in: : Maximum speed set by the system; : The maximum curvature allowed for the rail trolley (if the curvature exceeds this, the trolley may overturn). is a preset small value (such as 0.001) to avoid division by zero; : The attention value at this point (by nearest neighbor or bilinear interpolation from Figure 1); λ is the load factor; is the sensitivity coefficient of the load change on the speed, and its value range is [0, 1]. =0 means ignoring the effect of load changes; =1 means the load change has the maximum effect on the speed (i.e. the speed is scaled by a factor of 1 / λ); 0 < < 1 means that the load change has an impact on the speed between the two. This value can be set according to the system characteristics, for example = 0.5.

[0171] Additionally, consider the risk of obstacles in dynamic environments: if there are moving obstacles near the point, reduce the speed further.

[0172] The control signal generating unit 153 is used to generate steering and acceleration control instructions based on the smoothed path and the planned driving speed, and specifically performs the following operations:

[0173] (1) Steering control: Pure Pursuit algorithm is used. During the movement of the rail trolley, a certain forward distance from the current position of the rail trolley is selected on the path. The point is taken as the target point and the steering angle δ is calculated:

[0174] Where: L: wheelbase of rail trolley (distance between front and rear wheels); : The angle between the current heading of the track car and the vector from the track car to the target point, that is, the heading deviation angle; : Foresight distance, which is positively correlated with speed (e.g. , and is a constant, and v is the real-time speed of the rail car).

[0175] Speed ​​control: Use PID controller to track the planned speed curve v(s): ; Where a is the acceleration command.

[0176] The real-time feedback unit 154 is used to collect the deviation between the actual position of the rail vehicle and the nearest point on the currently planned path, and trigger the re-planning instruction when the trigger condition is met.

[0177] Collect the actual position of the rail trolley Distance from the planned path at the current moment The nearest point Deviation:

[0178] When any of the following conditions is met, >When a preset threshold (e.g. 0.2m) is reached or a new obstacle enters the safe distance = 0.5v + 0.3 (m) or the predicted path of the moving obstacle intersects the current path, where v is the real-time speed of the track vehicle. This triggers the following replanning instructions:

[0179] Use the latest sensor data to update the car state, obstacle information, etc. in the dynamic track environment model (environment modeling module 130), and recalculate the direction mask in the three-motion attention mechanism (model optimization module 140) based on the updated environment information Based on the updated environment model and attention mask, re-execute the path search process (AGFPS method) of the model optimization module 140 to generate a new optimal path ; Control execution module 150 based on the new path Path smoothing, velocity planning and control signal generation (steering angle δ and acceleration a) are performed again.

[0180] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the methods.

[0181] It should also be noted that, in the embodiments of the present application, relational terms such as first and second, etc. are merely used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the statement "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device comprising the elements.

[0182] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined in the embodiments of the present application may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown in the embodiments of the present application, but rather will conform to the widest scope consistent with the principles and novel features disclosed in the embodiments of the present application.

Claims

1. A dynamic optimization control system for a rail vehicle's travel path based on AI vision, characterized in that: include: Data acquisition module, used to obtain real-time track environment images, track vehicle motion status data, and auxiliary environmental information including lighting conditions and track markings; The data preprocessing module standardizes the collected data and outputs the standardized sensor data; The environment modeling module is used to build a dynamic track environment model with environment perception and behavior prediction interfaces based on the standardized number of sensors to predict the future position of the track vehicle; The model optimization module extracts direction-aware features and generates a fused feature map through a three-motion attention mechanism in the diagonal, forward, and backward directions. It also extracts key points of the path based on an attention-guided fast path search method and then uses a graph search algorithm to search for the optimal path between the key points. The control execution module generates a control signal according to the optimal path and drives the rail vehicle to travel.

2. The AI ​​vision-based track trolley driving path dynamic optimization control system according to claim 1 is characterized in that: The environment modeling module includes: An environmental perception unit, configured to divide the track area into a grid matrix, each grid containing environmental attribute identifiers, including obstacles, track markings, and hazardous conditions; Trajectory prediction unit, used to predict the position of the next K time steps based on the current position and velocity vector through the kinematic model; The behavior prediction interface unit is used to integrate environmental and status data and output structured information, including: static obstacle set, dangerous area set, moving obstacle parameters, rail vehicle status and load factor.

3. The AI ​​vision-based track trolley driving path dynamic optimization control system according to claim 1 is characterized in that: The model optimization module includes: The three-direction modeling unit is used to model the diagonal direction to detect track oblique trends and obstacles; model the forward direction to enhance the analysis of the path accessibility ahead; and model the backward direction to identify rear obstacles and path closure characteristics. The direction mask suppression unit sets mask matrices for the three motion directions, including diagonal direction mask matrix, forward mask matrix and backward mask matrix, to suppress the attention scores of irrelevant directions.

4. The AI ​​vision-based track trolley driving path dynamic optimization control system according to claim 3 is characterized in that: The direction mask suppression unit specifically includes: The diagonal mask matrix is ​​set to: 0 when the feature point positions i and j are different, and negative infinity when they are the same; the reverse mask matrix is ​​set to: 0 when the feature point position i is greater than j, otherwise negative infinity; the forward mask matrix is ​​set to: 0 when the feature point position i is less than j, otherwise negative infinity.

5. The AI ​​vision-based track trolley driving path dynamic optimization control system according to claim 4 is characterized in that: The model optimization module also includes: The three-directional feature pair scoring unit is used to apply the corresponding mask matrix to any two feature points in the track image according to their positional relationship. The directional awareness transformation is performed on each of them through a learnable mapping matrix. The transformed feature vectors are added together and a learnable bias vector is added. The nonlinear interaction between features is enhanced through the hyperbolic tangent activation function. The mask values ​​of the corresponding directions are superimposed to generate a directional attention score. The probability distribution calculation unit is used to calculate the attention scores of the three motion directions for each feature channel respectively, and obtain the weighted probability distribution of the three directions through Softmax function normalization.

6. The AI ​​vision-based track trolley driving path dynamic optimization control system according to claim 5 is characterized in that: The model optimization module also includes: The feature fusion unit is used to weightedly aggregate adjacent features in corresponding directions according to the weighted probability distribution of the three directions to generate a direction-aware feature representation; a gating mechanism is used to weightedly combine the original visual features and the direction-aware features to generate a fused feature map representation; wherein the gating mechanism learns the degree of feature dependence through the Sigmoid function and adaptively adjusts the fusion ratio of the original information and the direction information.

7. The AI ​​vision-based track trolley driving path dynamic optimization control system according to claim 1 is characterized in that: The fast path search method based on attention guidance includes: Key point extraction step: Extract high-attention areas from the fused feature map output by the three-motion attention mechanism as key points, and add the current position and target position of the track car to form a key point set; Path graph construction steps: Use key points as nodes, connect adjacent nodes to form edges, and assign a weight to each edge. The weight is determined by the distance between nodes, the attention value at the node, and the risk value of the area passed by the edge; Graph search step: On the constructed path graph, use the A* search algorithm to find the optimal path from the starting point to the end point.

8. The AI ​​vision-based track trolley driving path dynamic optimization control system according to claim 7 is characterized in that: in, The step of extracting key points specifically includes: Perform channel compression on the fused feature map, compressing it into a single-channel two-dimensional attention map by taking the maximum value or average value; Based on threshold segmentation, the positions greater than or equal to τ in the attention map are marked as candidate key points; Perform non-maximum suppression on candidate key points to obtain a set of sparse key points; The current position and target position of the car are also added to the key point set, and finally the key points of the track car are extracted.

9. The AI ​​vision-based track trolley driving path dynamic optimization control system according to claim 1 is characterized in that: The control execution module includes: A path smoothing unit is used to smooth the optimal path obtained by searching using cubic spline interpolation or B-spline curve to obtain a smooth path; The speed planning unit is used to plan the driving speed according to the curvature of each point on the smooth path and the corresponding attention value; a control signal generating unit for generating steering and acceleration control instructions based on the smoothed path and the planned driving speed; The real-time feedback unit is used to collect the deviation between the actual position of the rail vehicle and the nearest point on the currently planned path. When the trigger condition is met, the re-planning instruction is triggered.

10. The AI ​​vision-based track trolley driving path dynamic optimization control system according to claim 2 is characterized in that: The data preprocessing module performs the following operations: image denoising, brightness equalization, and edge enhancement.

Citation Information

Patent Citations

  • Autonomous robot path planning method based on deep reinforcement learning

    CN119105512A

  • Method and system for establishing and planning intelligent path of car in scenic spot based on AI algorithm

    CN119809067A

  • Multi-robot path planning method and system under attention mechanism

    CN120274775A

  • Stable movement global path planning method for indoor mobile robot

    WO2023155371A1