A Reinforcement Learning-Based Control Method and System for Electrified Railway Maintenance Robots
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING YANLING JIAYE INTELLIGENT TECH CO LTD
- Filing Date
- 2026-03-10
- Publication Date
- 2026-05-26
Smart Images

Figure CN121798642B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of reinforcement learning and intelligent control, and more specifically, to a control method and system for an electrified railway maintenance robot based on reinforcement learning. Background Technology
[0002] With the continuous increase in the operational mileage of electrified railways, maintenance robots have been widely used. By controlling specialized maintenance robots, key components such as tracks, overhead contact lines, and power equipment can be inspected and maintained. Currently, the control of maintenance robots typically involves constructing a simplified simulation training environment, loading partial fault scenario data to complete strategy training, and then combining a small amount of interactive data from real maintenance scenarios to locally optimize the strategy. During maintenance operations, the pre-trained strategy is directly invoked to generate a sequence of robot actions based on the current task, and local parameters are updated for the strategy model after the actions are executed. In practical applications, the simulation environment of this method does not match the complex physical characteristics of real railway scenarios sufficiently. When faced with complex external disturbances in real scenarios, the accuracy of action control is difficult to guarantee. The dynamic correlation between training data and real maintenance scenarios is limited, the strategy's adaptability to sudden faults and scenario fluctuations is weak, and there is a lack of effective scenario adaptability verification before robot actions are executed, which can easily lead to safety hazards and make it difficult to continuously meet the complex and ever-changing railway maintenance needs. Summary of the Invention
[0003] This invention provides a control method and system for an electrified railway maintenance robot based on reinforcement learning.
[0004] In a first aspect, embodiments of the present invention provide a control method for an electrified railway maintenance robot based on reinforcement learning. The method includes: constructing a multi-physics digital twin training space with safety boundary constraints based on the physical structural attributes of railway sections, high-voltage electromagnetic effects, and airflow vibration effects, using high-voltage electric field safety thresholds and mechanical collision safety thresholds as constraints; loading random fault disturbance attributes to simulate disturbances in the multi-physics digital twin training space; generating a virtual-real coupled training data set containing safety boundary association information; processing the virtual-real coupled training data set based on a meta-reinforcement learning training strategy; combining interactive data from real railway maintenance scenarios to perform safety transfer adaptation learning; generating a multi-scenario adapted safety control strategy cluster adapted to multiple risk scenarios; and adjusting the current data based on the real railway maintenance scenario. The inspection task and real-time scene characteristics are used to match the corresponding target control strategy from the multi-scene adapted safety control strategy cluster to generate the robot's initial action execution sequence. The real-time collected railway maintenance status data is mapped to the multi-physics digital twin training space for action pre-play risk verification, generating action verification adjustment instructions. Based on the action verification adjustment instructions, the robot's initial action execution sequence is optimized, and safe execution actions are output to drive the maintenance robot to complete the railway maintenance operation. The safe execution action data and real scene characteristics are fed back to the multi-physics digital twin training space and meta-reinforcement learning training rules to update the safety boundary parameters of the multi-physics digital twin training space and the multi-scene adapted safety control strategy cluster, respectively, to complete the dual-track iterative update of virtual and real fusion.
[0005] Secondly, embodiments of the present invention provide a computer system, including: a memory storing a computer program; and a processor for loading the computer program to implement the reinforcement learning-based control method for an electrified railway maintenance robot as described above.
[0006] This invention provides a precise simulation foundation for strategy training by constructing a multi-physics field-constrained twin training space that fits the complex physical characteristics of electrified railways. It generates virtual-real coupled training data that carries the safety correlation logic of real scenarios, improving the adaptability of training data to actual scenarios. The control strategy generated by combining meta-reinforcement learning and real-scene transfer adaptation can accurately cope with various complex risks in railway scenarios. The initial action sequence generated by matching the current task with the real-time scenario is more in line with actual needs. The safety execution actions optimized by twin space pre-rehearsal verification effectively reduce safety hazards in the maintenance process. The dual-track iteration mechanism of synchronously updating the twin space boundary and control strategy continuously improves the dynamic adaptability of the system, enhancing the control accuracy and execution safety of the maintenance robot in electrified railway scenarios. Attached Figure Description
[0007] Figure 1This is a flowchart of a control method for an electrified railway maintenance robot based on reinforcement learning, provided by an embodiment of the present invention.
[0008] Figure 2 This is a schematic diagram of the composition of a computer system provided in an embodiment of the present invention. Detailed Implementation
[0009] Please see Figure 1 This is a flowchart of a reinforcement learning-based control method for an electrified railway maintenance robot, provided by the present invention. The method can be executed by a computer system and includes the following steps:
[0010] Step S100: Based on the physical structure attributes of the railway section, the high-voltage electromagnetic effect attributes, and the airflow vibration effect attributes, and with the high-voltage electric field safety threshold and the mechanical collision safety threshold as constraints, a multi-physics digital twin training space with safety boundary constraints is constructed. Random fault disturbance attributes are loaded to simulate the disturbance of the multi-physics digital twin training space, and a virtual-real coupled training data set containing safety boundary association information is generated.
[0011] The physical structural attributes of railway sections encompass the actual physical characteristics of railway track laying methods, bridge structures, and tunnel morphology, defining the basic physical scope for the activities of maintenance robots. High-voltage electromagnetic effects involve the distribution and intensity of high-voltage electric and magnetic fields around the railway. Since the electronic components of maintenance robots are susceptible to electromagnetic interference, this attribute significantly impacts their normal operation. Airflow vibration effects refer to the vibration characteristics caused by airflow generated by train movement and natural environmental airflow; this vibration can lead to loosening of railway equipment components and structural damage. The high-voltage electric field safety threshold is the upper limit of electric field strength set based on the safety tolerance of humans and equipment, while the mechanical collision safety threshold is the upper limit of collision force stipulated to avoid damage to railway facilities and personnel. The multi-physics digital twin training space uses digital twin technology to reproduce the multi-physics interactions in the actual railway environment in a virtual space, enabling maintenance robots to train in a virtual environment. The virtual-real coupled training dataset integrates training data from the virtual space with data from the real railway scenario, including safety boundary information, making the training data more realistic.
[0012] As one implementation method, in step S100, based on the physical structural properties of the railway section, the properties of high-voltage electromagnetic effects, and the properties of airflow vibration effects, a multi-physics digital twin training space with safety boundary constraints is constructed, using the high-voltage electric field safety threshold and the mechanical collision safety threshold as constraints. Specifically, this may include the following steps S110~S160:
[0013] Step S110: Track the energy transfer trajectory of all nodes in the physical structure of the railway section, mark the energy input and output ports of each node, record the energy flow rate and direction of each port simultaneously, and analyze the influence of the electromagnetic field on the energy state of the structural nodes in combination with the global influence range of the high-voltage electromagnetic field to obtain a set of structural energy state trajectories under the action of the electromagnetic field, covering all main and branch structures of the railway section.
[0014] The physical structure of a railway section consists of numerous nodes with complex energy transfer relationships. To track the energy transfer trajectory across the entire network, a sensor network can be used to monitor the energy of each node in real time. Sensors can accurately measure the energy flow rate and direction at energy input and output ports and transmit the data to a data processing center. At the data processing center, an energy transfer model is established to analyze and integrate the energy data from each node, thereby clarifying the energy transfer path within the physical structure. Considering the overall influence range of high-voltage electromagnetic fields, electromagnetic simulation software can be used to simulate the distribution of the electromagnetic field. Combining the simulation results with the node energy data allows for the analysis of the impact of the electromagnetic field on the energy state of each node.
[0015] Step S120: Based on the set of structural energy state trajectories under the action of electromagnetic field, combined with the global influence range of airflow vibration, analyze the coupling effect of electromagnetic energy and vibration energy and the comprehensive intensity change node by node to obtain the node state evolution path network under the multi-physics coupling action, and update the state record of each node synchronously. The change trajectory records the state evolution of the physical structure node to obtain the global multi-physics coupling state trajectory set.
[0016] After obtaining the set of structural energy state trajectories under the influence of electromagnetic fields, the coupling effect between electromagnetic energy and vibration energy is analyzed by combining the global influence range of airflow vibration. The global influence range of airflow vibration can be determined by combining wind tunnel tests and field monitoring. After determining the influence range of airflow vibration, multiphysics coupling analysis software is used to perform coupling calculations on the electromagnetic energy and vibration energy of each node. This software simulates the interaction process between electromagnetic energy and vibration energy based on the physical characteristics of the node and its environment, calculating the coupling effect and the overall intensity change. Through node-by-node analysis, a node state evolution path network under multiphysics coupling is constructed. In this process, the state record of each node is updated synchronously, and the state change trajectory of the node is recorded using a data storage system, ultimately obtaining a set of global multiphysics coupling state trajectories.
[0017] Step S130: Associate the set of coupled state trajectories of the global multi-physics field with the high-voltage electric field safety threshold, record the change characteristics of the node's comprehensive energy state as it approaches the threshold node node by node, obtain a continuous threshold approach node state change path network, synchronously analyze the correlation between the state change characteristics of each node and the electric field strength, record the trajectory response when the electric field strength changes using trajectory deflection data, and output the set of node state trajectories under electric field safety constraints.
[0018] Associating the set of coupled state trajectories across a multiphysics field with a high-voltage electric field safety threshold can be achieved using a data matching algorithm. This algorithm compares the comprehensive energy state of each node with the high-voltage electric field safety threshold. When the comprehensive energy state of a node approaches the threshold, a feature recording program is initiated. This program records the node's energy change rate, fluctuations, and other characteristics in real time, constructing a continuous network of state change paths for nodes approaching the threshold. Simultaneously, data analysis tools are used to analyze the correlation between the state change characteristics of each node and the electric field strength. For example, regression analysis is used to determine the mathematical relationship between state change characteristics and electric field strength. Trajectory deflection data is recorded by monitoring the deviation of the node's energy state trajectory as the electric field strength changes. These data are then integrated to output a set of node state trajectories under electric field safety constraints.
[0019] Step S140: Based on the set of node state trajectories under electric field safety constraints, combined with the mechanical collision safety threshold, record the node state anomaly patterns when the structural deformation approaches the threshold node by node, obtain a continuous deformation approaching node state anomaly path network, and simultaneously analyze the correlation between the state anomaly patterns of each node and the degree of structural deformation. The mutation mode records the trajectory features when the degree of structural deformation reaches the threshold, and solidifies them into a set of node state trajectories under collision safety constraints.
[0020] Based on the set of node state trajectories under electric field safety constraints, and combined with mechanical collision safety thresholds, structural deformation is monitored. Strain sensors can be used to measure the structural deformation of each node in real time. When the structural deformation approaches the mechanical collision safety threshold, the system automatically records the abnormal state patterns of the nodes, such as sudden changes in node vibration frequency or drastic fluctuations in energy state. By recording these abnormal patterns, a continuous network of abnormal deformation paths approaching node state is constructed. Simultaneously, the correlation between the abnormal state patterns of each node and the degree of structural deformation is analyzed to determine the quantitative relationship between the two. Sudden change patterns are recorded when the degree of structural deformation reaches the threshold, showing the characteristics of the node's energy state trajectory. This information is then solidified to obtain the set of node state trajectories under collision safety constraints.
[0021] Step S150: Integrate and analyze the set of node state trajectories under electric field safety constraints and the set of node state trajectories under collision safety constraints. Identify the node state evolution law under the combined action of electric field and collision safety constraints for each node, obtain the node state evolution path network under the coupled action of dual safety constraints, and simultaneously verify the correlation between the identified evolution law and the intensity of dual constraint action. The evolution law is used to describe the characteristics of the region under the combined action of dual constraints and integrate them into a set of node state trajectories coupled by dual safety constraints.
[0022] By fusing the sets of node state trajectories under electric field safety constraints and collision safety constraints, a data fusion algorithm can be employed to match and integrate the data from the two sets, eliminating redundancy and conflicts. Based on the fused data, the evolutionary patterns of node states under the combined effects of electric field and collision safety constraints are identified node by node. Machine learning algorithms, such as neural networks, can be used to train and analyze the node state data to obtain the evolutionary patterns. Simultaneously verifying the correlation between the identified evolutionary patterns and the intensity of the dual constraints, the changes in node state evolution can be observed by varying the intensity of the dual constraints, establishing a correlation model between the two. Integrating this information yields the set of node state trajectories coupled with dual safety constraints.
[0023] Step S160: Based on the set of state trajectories of coupled nodes with dual security constraints, construct the node mapping and state evolution logic of the digital twin space to ensure that the virtual state evolution of nodes in the twin space can reflect the multi-physics field coupled state and security constraints of real physical nodes, and obtain a multi-physics field digital twin training space with security boundary constraints.
[0024] This paper constructs a node mapping and state evolution logic for a digital twin space based on a set of state trajectories of coupled nodes with dual security constraints. Node mapping is achieved by establishing a one-to-one correspondence between real physical nodes and virtual nodes in the digital twin space. Node identification technology can be used to assign a unique identifier to each real physical node and create a corresponding virtual node in the digital twin space. The state evolution logic establishes a model of the change of virtual node states with time and multiphysics interactions based on the data in the set of state trajectories of coupled nodes with dual security constraints. Through simulation algorithms, the state evolution process of virtual nodes is simulated in the digital twin space to ensure that it accurately reflects the multiphysics coupling state and security constraints of real physical nodes. This construction yields a multiphysics digital twin training space with security boundary constraints.
[0025] As one implementation method, in step S100, random fault disturbance attributes are loaded to simulate the disturbance in the multiphysics digital twin training space, generating a virtual-real coupled training data set containing safety boundary correlation information. Specifically, this may include the following steps S170~S1120:
[0026] Step S170: Determine the simulated triggering point of the random fault disturbance, map it to the corresponding node in the multiphysics digital twin training space, and deduce the transmission path of the fault state based on the physical association logic of the nodes in the twin space to obtain the transmission path network of the fault in the twin space. The transmission path of each fault triggering point covers the related nodes affected by it, and obtain the fault-space node transmission trajectory set.
[0027] When determining the simulated trigger points for random fault disturbances, some nodes can be randomly selected as trigger points based on historical fault data and fault probability models of the railway section. These trigger points are then mapped to corresponding nodes in the multiphysics digital twin training space, which can be achieved through node identifier matching. Based on the physical association logic of nodes within the twin space, such as electrical connections and mechanical transmission relationships, graph theory algorithms are used to deduce the transmission path of the fault state. Graph theory algorithms treat nodes as vertices in a graph and the relationships between nodes as edges, using a search algorithm to find the path for the fault to propagate from the trigger point to other nodes. After deduction, a fault transmission path network is obtained within the twin space. The transmission path of each fault trigger point covers the related nodes affected by it. Integrating this path information yields a set of fault-space node transmission trajectories.
[0028] Step S180: Simulate the energy transfer logic of the fault-space node transmission trajectory set and the multiphysics digital twin training space. Simulate the fault transmission process in the space time by time period, synchronously record the spatial state change data of each time period, and synchronize the state change data of all time periods with the fault transmission trajectory to output the fault disturbance virtual state trajectory sequence.
[0029] When simulating the energy transfer logic of the fault-space node propagation trajectory set and the multiphysics digital twin training space, a fault propagation model and an energy transfer model can be established and coupled. Based on the coupled model, simulation software is used to simulate the fault propagation process in space time-by-time, updating the state of each node according to the fault propagation path and energy transfer rules. During the simulation, the spatial state change data of each time period is recorded synchronously, and the state data of the nodes can be collected in real time using a data acquisition system. Ensuring that the state change data of all time periods are synchronized with the fault propagation trajectory, these data are integrated to output a sequence of virtual state trajectories of fault disturbance.
[0030] Step S190: Compare the virtual state trajectory sequence of fault disturbance with the historical fault state trajectory of the real railway section, record the trajectory deviation points and deviation degree of virtual and real states node by node, and obtain a continuous deviation node change path network. The deviation degree of each deviation point is significantly related to the fault type and intensity. The deviation points cover the key stages of fault evolution, and obtain the set of virtual and real state deviation trajectories.
[0031] When comparing the virtual state trajectory sequence of fault disturbances with the historical fault state trajectories of real railway sections, a trajectory matching algorithm can be used. This algorithm compares the virtual and real state trajectories point by point to identify deviation points. For each deviation point, the degree of deviation between the virtual and real states is calculated, which can be achieved by calculating the difference or similarity between the two. By recording these deviation points and their degrees, a continuous network of deviation node change paths is constructed. Analyzing the correlation between the degree of deviation at each deviation point and the fault type and intensity, a correlation model between the two can be established using statistical analysis methods. The deviation points cover the key stages of fault evolution; integrating this information yields a set of virtual and real state deviation trajectories.
[0032] Step S1100: Calibrate the fault disturbance virtual state trajectory sequence, align the deviation points of the virtual and real state deviation trajectory set, adjust the evolution trajectory of the virtual state node by node, and simultaneously verify the overlap between the adjusted virtual state and the real state trajectory. The adjusted virtual state and the real state trajectory overlap and are solidified into a virtual-real fusion state trajectory sequence.
[0033] When calibrating the virtual state trajectory sequence under fault disturbance, the evolution trajectory of the virtual state is adjusted based on the deviation points of the virtual-real state deviation trajectory set. Interpolation or filtering algorithms can be used to correct the virtual state data. During the adjustment process, operations are performed node by node to ensure that the virtual state of each node is closer to the real state. The overlap between the adjusted virtual state and the real state trajectory is verified simultaneously, which can be evaluated by calculating a similarity index between the two. When the adjusted virtual state and the real state trajectory overlap, these data are solidified to obtain the virtual-real fused state trajectory sequence.
[0034] Step S1110: Associate the virtual-real fusion state trajectory sequence with the safety boundary constraints of the multi-physics digital twin training space, record the trajectory deflection amplitude when the state approaches the boundary node node by node, and obtain a continuous boundary approach node interaction path network. The deflection amplitude corresponds one-to-one with the boundary distance. The trajectory data of all boundary approach nodes are fully covered and integrated into a state-boundary associated trajectory set.
[0035] When associating the virtual-real fusion state trajectory sequence with the safety boundary constraints of the multiphysics digital twin training space, the state data of each node is compared with the safety boundary. When a node's state approaches the boundary, its trajectory deflection magnitude is recorded. The deflection magnitude can be determined by calculating the angle or distance change rate between the node's state trajectory and the safety boundary. By recording node by node, a continuous network of interaction paths between boundary-approaching nodes is obtained. Analyzing the correspondence between deflection magnitude and boundary distance allows the establishment of a function model between the two, ensuring complete coverage of trajectory data for all boundary-approaching nodes. Integrating this information yields a set of state-boundary associated trajectories.
[0036] Step S1120: Label the state-boundary associated trajectory set and the virtual-real fusion state trajectory sequence. Label the corresponding boundary association information at each state node to obtain a continuous labeled state trajectory network. The label information is consistent with the boundary constraints and covers all state evolution periods, generating a virtual-real coupled training data set containing safe boundary association information.
[0037] When labeling the state-boundary associated trajectory set and the virtual-real fusion state trajectory sequence, the corresponding boundary association information, such as the distance to the safety boundary and the trajectory deflection magnitude, is marked at each state node. Data annotation tools can be used to add this information to the data of the state nodes. Through such annotation, a continuous labeled state trajectory network is obtained. Ensuring that the labeled information is consistent with the boundary constraints and covers all state evolution periods, this labeled information is integrated to generate a virtual-real coupled training dataset containing safety boundary association information.
[0038] Step S200: Process the virtual-real coupled training data set based on the meta-reinforcement learning training strategy, combine it with the interactive data of real railway maintenance scenarios to perform safety transfer adaptation learning, and generate a multi-scenario adapted safety control strategy cluster that is adapted to multiple risk scenarios.
[0039] As one implementation method, step S200, processing the virtual-real coupled training data set based on the meta-reinforcement learning training strategy, may specifically include the following steps S210~S260:
[0040] Step S210: Track the policy update time sequence of meta-reinforcement learning, associate the state change time sequence of the virtual-real coupled training data set, record the correspondence between the state change time period and the policy update time period for each time period, obtain a continuous time period interaction path network, and synchronize the correspondence of each time period with the state change logic to obtain a policy-state time period associated trajectory set.
[0041] Tracking the policy update sequence in meta-reinforcement learning can be achieved by recording the timestamps of policy updates in the meta-reinforcement learning algorithm. By associating the state change sequence of the virtual-real coupled training dataset, the time information of state changes is matched with the policy update time. The correspondence between state change periods and policy update periods is recorded for each time period, and these correspondences can be stored using data tables or databases. Through this recording, a continuous time-segment interaction path network is obtained. The correspondence of each time period is synchronized with the state change logic, ensuring that policy updates can respond promptly to state changes. Integrating this information yields a set of policy-state time-segment associated trajectories.
[0042] Step S220: Trigger the policy for each state change period in the policy-state time period associated trajectory set, synchronously record the state response trajectory after the policy is triggered, obtain a continuous trigger-response path network, the state response after each policy is triggered matches the policy objective, all response trajectories cover all stages of state change, and output the policy trigger-state response trajectory set.
[0043] After obtaining the set of policy-state time-period associated trajectories, the corresponding policy is triggered based on the policy corresponding to each state change time period recorded in the set. Based on time and state information, the policy to be triggered for each state change time period is accurately identified, and a trigger command is sent to the system. After the policy is triggered, a corresponding state response is generated. These state responses are monitored and recorded in real time using a sensor network and a data acquisition system. The sensor network is distributed across various key nodes of the system, enabling precise measurement of changes in state variables. The data acquisition system is responsible for integrating and storing the data collected by the sensors. Through processing and analysis of this data, a continuous trigger-response path network is constructed. During the construction process, it is ensured that the state response after each policy trigger matches the policy objective. This can be achieved by setting an objective function and evaluation indicators to quantitatively evaluate the state response and determine whether it meets the policy objective. All response trajectories must cover all stages of state change. These response trajectories are organized according to time sequence and state change logic, outputting a policy trigger-state response trajectory set.
[0044] Step S230: Calculate the policy adjustment gradient of the policy trigger-state response trajectory set, the policy optimization logic of the association element reinforcement learning, record the evolution path of gradient adjustment in time period, obtain a continuous gradient adjustment path network, synchronously calibrate the correspondence between gradient adjustment magnitude and state response deviation, match the gradient adjustment logic with the policy optimization objective, and obtain the policy gradient update trajectory set.
[0045] When calculating the policy adjustment gradient of the policy trigger-state response trajectory set, gradient calculation algorithms, such as gradient descent-based optimization algorithms, can be used. Based on the difference between the state response and the policy objective, the direction and magnitude of the policy adjustment need to be calculated. This is then linked to the policy optimization logic of meta-reinforcement learning, which typically involves iterative optimization of the policy to improve its performance. The gradient calculation results are combined with the optimization logic to determine the gradient adjustment method and frequency. The evolution path of gradient adjustment is recorded time-by-time, using time series analysis methods to record and analyze the gradient adjustment data for each time period. By establishing a time series model of gradient adjustment, a continuous gradient adjustment path network is obtained. The correspondence between the gradient adjustment magnitude and the state response deviation is simultaneously calibrated; this functional relationship can be established through experiments and data analysis. The gradient adjustment logic is ensured to match the policy optimization objective by setting an optimization objective function and constraints to optimize the gradient adjustment process. Integrating this information yields the policy gradient update trajectory set.
[0046] Step S240: Iterate the policy gradient to update the trajectory set and the global state change trajectory of the virtual-real coupled training data set. Apply the policy gradient to adjust the policy parameters in each global state cycle to obtain a continuous parameter adjustment path network. Simultaneously verify the adaptability of the adjusted policy parameters to the state changes. The parameter adjustment in each cycle covers all state nodes and is solidified into a global policy parameter optimization trajectory set.
[0047] When iterating through the policy gradient update trajectory set and the global state change trajectory of the virtual-real coupled training dataset, the policy gradient update trajectory and the global state change trajectory are matched and analyzed point by point. Within each global state cycle, the policy parameters are adjusted based on the policy gradient. This adjustment process can be implemented through a parameter update module, which calculates the adjusted policy parameters based on gradient information and the current policy parameters. During the adjustment process, a continuous parameter adjustment path network is constructed, recording the steps and results of each parameter adjustment. The adaptability of the adjusted policy parameters to state changes is simultaneously verified. The execution effect of the adjusted policy under different states can be observed through simulation experiments and actual tests. The parameter adjustment in each cycle must cover all state nodes to ensure the comprehensiveness and effectiveness of the policy. The adjusted policy parameters and related information are then solidified to obtain the global policy parameter optimized trajectory set.
[0048] Step S250: Verify the policy adaptability of the global policy parameter optimization trajectory set, solidify the adapted policy parameters node by node to obtain a continuous parameter solidification path network, and simultaneously verify the consistency between the solidified policy parameters and the state change time sequence. All verified parameters cover all state change time periods and are integrated into an iteratively optimized policy parameter set.
[0049] When verifying the policy adaptability of the global policy parameter optimization trajectory set, various verification methods can be used, such as comparative experiments and simulations. By comparing the system's performance under different policy parameters, the policy adaptability is evaluated. Adapted policy parameters are then fixed node by node, and a data storage system can be used to save the appropriate policy parameters. During the fixing process, a continuous parameter fixing path network is constructed, recording the process and results of parameter fixing. The consistency between the fixed policy parameters and the state change time series is verified simultaneously. Time series analysis and correlation analysis can be used to determine whether the time correspondence between parameters and state changes is reasonable, ensuring that all verified parameters cover all state change periods. These parameters are then integrated to obtain the iteratively optimized policy parameter set.
[0050] Step S260: Integrate the state change trajectories of the iteratively optimized policy parameter set and the virtual-real coupled training data set, synchronize the policy triggering sequence and the state change sequence to obtain a continuous policy triggering path network, simultaneously verify the matching of the policy triggering sequence and the state change sequence, ensure that the policy triggering path of all nodes is fully covered, and generate the meta-trained and adapted policy parameter set.
[0051] When integrating the optimized policy parameter set with the state change trajectories of the virtual-real coupled training dataset, the policy parameters and state change trajectories are correlated and matched. A data association model is established to bind the policy parameters to the corresponding state change time periods. The policy triggering sequence and state change sequence are synchronized using a time synchronization algorithm to ensure the policy is triggered at the appropriate time. During synchronization, a continuous policy triggering path network is constructed, recording the time and path of policy triggering. The matching between the policy triggering sequence and the state change sequence is verified by calculating the time difference and correlation coefficient between them, ensuring complete coverage of the policy triggering paths for all nodes. This information is then integrated to generate the meta-trained and adapted policy parameter set.
[0052] As one implementation method, in step S200, safety migration adaptation learning is performed by combining interactive data from real railway maintenance scenarios, which may specifically include the following steps S270~S2120:
[0053] Step S270: Collect interactive data trajectories from real railway maintenance scenarios, match the triggering sequence of the policy parameter set after meta-training adaptation, record the correspondence between interactive time periods and policy triggering time periods for each time period, obtain a continuous interactive time period path network, and synchronize the correspondence of each time period with the scenario interaction logic to obtain a scenario interaction-policy time period associated trajectory set.
[0054] When collecting interactive data trajectories from real railway maintenance scenarios, various sensors and data acquisition devices distributed at the railway maintenance site can be utilized. For example, position sensors can be used to record the movement trajectory of maintenance robots, pressure sensors can be used to monitor the stress on equipment, and image sensors can be used to acquire visual information from the site. The collected data is then organized and analyzed to form interactive data trajectories. These trajectories are matched with the triggering sequence of the policy parameter set after meta-training and adaptation. By comparing time information, the correspondence between interaction periods and policy triggering periods is identified. These correspondences are recorded period by period and can be stored using a database or spreadsheet. During the recording process, a continuous path network of interaction periods is constructed to ensure that the correspondence of each period is synchronized with the scene interaction logic. For example, when a maintenance robot performs a certain operation, the corresponding policy should be triggered at an appropriate time. Integrating this information yields a set of scene interaction-policy period-related trajectories.
[0055] Step S280: Trigger the strategy for each interaction period in the strategy-time period associated trajectory set of the scene interaction, synchronously deduce the scene state response after the strategy is triggered, obtain a continuous scene response path network, match the scene response after each strategy is triggered with the task objective, all response trajectories cover all stages of scene interaction, and output the strategy trigger-scene response trajectory set.
[0056] Based on the correspondence recorded in the scenario interaction-strategy time-period associated trajectory set, the strategy for each interaction period is triggered. This triggering can be achieved through a strategy execution module, which accurately identifies and executes the appropriate strategy based on time and interaction information. After strategy triggering, scenario state responses are deduced using scenario modeling and simulation techniques. Scenario modeling can be based on the physical structure and equipment characteristics of real railway maintenance scenarios to construct a virtual scenario model. Simulation technology simulates the change process of the scenario state based on the strategy execution and the initial state of the scenario. These scenario response trajectories are recorded synchronously to construct a continuous scenario response path network, ensuring that the scenario response after each strategy trigger matches the task objective. The scenario response can be evaluated by setting a task objective function and evaluation indicators. All response trajectories must cover all stages of scenario interaction. These response trajectories are then organized and output to obtain a strategy trigger-scenario response trajectory set.
[0057] Step S290: Compare the state response trajectories of the policy trigger-scene response trajectory set with the virtual-real coupled training data set, record the trajectory deviation points and deviation degrees of virtual and real states node by node, and obtain a continuous deviation node change path network. The deviation degree of each deviation point is significantly related to the scene characteristics and policy strength. The deviation points cover the key stages of fault evolution, and obtain the virtual-real response deviation trajectory set.
[0058] When comparing the set of policy-triggered scenario response trajectories with the state response trajectories of the virtual-real coupled training dataset, a trajectory comparison algorithm is employed. This algorithm compares the virtual state trajectory and the real scenario state trajectory point by point to identify trajectory deviation points. For each deviation point, the degree of deviation between the virtual and real states is calculated, which can be achieved by calculating the difference or similarity between the two. During the recording of deviation points and their degrees, a continuous deviation node change path network is constructed. The correlation between the degree of deviation at each deviation point and scenario characteristics and policy strength is analyzed. A correlation model can be established using statistical analysis methods. The deviation points should cover the key stages of fault evolution. Integrating this information yields the set of virtual-real response deviation trajectories.
[0059] Step S2100: Calibrate the set of policy parameters after training and adaptation, align the deviation points of the virtual and real response deviation trajectory set, adjust the trigger parameters of the policy node by node, and simultaneously verify the consistency between the response of the adjusted policy and the virtual state response during the scene interaction period. The adjusted policy is adapted to the scene interaction logic and solidified into the set of policy parameter trajectories after scene adaptation.
[0060] Based on the deviation point information provided by the set of virtual and real response deviation trajectories, the policy parameter set after calibration and training can be adjusted node by node using a parameter adjustment algorithm. During the adjustment process, deviation points are aligned to ensure the targeted nature of the adjustment. The consistency between the adjusted policy's response and the virtual state response during the scene interaction period is simultaneously verified through comparative experiments and simulation analysis. The response of the adjusted policy in both real and virtual scenes is observed, and the similarity between the two is calculated to ensure that the adjusted policy adapts to the scene interaction logic. By analyzing the process and requirements of scene interaction, the policy is optimized, and the adjusted policy parameters and related information are solidified to obtain the set of policy parameter trajectories after scene adaptation.
[0061] Step S2110: Verify the global adaptability of the strategy parameter trajectory set after scene adaptation. Generalize the adapted strategy parameters in each global interaction cycle to obtain a continuous generalized parameter path network. Simultaneously verify the adaptability of the generalized strategy parameters to scene state changes. All verified parameters cover all scene interaction periods and are integrated into a global scene adaptation strategy parameter set.
[0062] To verify the global adaptability of the policy parameter trajectory set after scenario adaptation, a combination of global simulation and actual testing can be used. In global simulation, the execution of the policy is simulated under different scenario and task conditions to evaluate its adaptability. Within each global interaction cycle, the generalized policy parameters can be adjusted and optimized using generalization algorithms in machine learning. During generalization, a continuous generalized parameter path network is constructed, recording the parameter changes and simultaneously verifying the adaptability of the generalized policy parameters to changes in scenario state. By monitoring changes in scenario state and the policy execution effect, the rationality of the parameters is judged. Ensuring that all validated parameters cover all scenario interaction periods, these parameters are integrated to obtain the global scenario-adapted policy parameter set.
[0063] Step S2120: Integrate the policy parameter set after global scene adaptation with the transfer learning logic of meta-reinforcement learning, synchronize the policy adaptation time series and transfer cycle to obtain a continuous transfer stage path network, simultaneously verify the matching of the policy adaptation time series and transfer cycle, fully cover the policy parameters of all transfer stages, and generate the policy parameter set after transfer adaptation.
[0064] When integrating the policy parameter set after global scene adaptation with the transfer learning logic of meta-reinforcement learning, the policy parameters are combined with the rules and methods of transfer learning. The transfer learning logic specifies how to transfer knowledge and policies learned in one environment to another, synchronizing the policy adaptation timeline with the transfer cycle. A time synchronization mechanism ensures that the policy is transferred at the appropriate time. During synchronization, a continuous transfer stage path network is constructed, recording each stage of the transfer and parameter changes. The matching between the policy adaptation timeline and the transfer cycle is verified synchronously. By calculating the time difference and performing correlation analysis, the consistency between the two is evaluated, ensuring complete coverage of policy parameters across all transfer stages. This information is then integrated to generate the policy parameter set after transfer adaptation.
[0065] As one implementation method, step S200 generates a multi-scenario adaptive security control strategy cluster that adapts to multiple risk scenarios, which may specifically include the following steps S2130~S2180:
[0066] Step S2130: Extract the state change trajectory of multiple risk scenarios, classify the evolution characteristics of different risk scenarios, match the triggering sequence of the policy parameter set after migration and adaptation, obtain a continuous scenario policy path network, the corresponding path of each scenario is synchronized with the risk evolution logic, and integrate to obtain a risk scenario-policy related trajectory set.
[0067] When extracting state change trajectories across multiple risk scenarios, historical and real-time monitoring data can be utilized. For historical data, state records under different risk scenarios can be extracted from a database. For real-time monitoring data, the current scenario's state is collected in real-time via a sensor network. These data are analyzed and processed to obtain the state change trajectories. The evolutionary characteristics of different risk scenarios can be categorized using clustering algorithms. Based on the patterns and characteristics of state changes, different risk scenarios are classified, and the triggering sequence of the policy parameter set after migration adaptation is matched. By comparing time information, the correspondence between scenarios and policy triggering is identified. During the matching process, a continuous scenario-policy path network is constructed to ensure that the corresponding path for each scenario is synchronized with the risk evolution logic. For example, in scenarios where risk gradually increases, the corresponding policy should be triggered at an appropriate time. Integrating this information yields a set of risk scenario-policy associated trajectories.
[0068] Step S2140: Match the strategy parameters of each risk scenario in the risk scenario-strategy associated trajectory set, synchronously deduce the scenario state response after the strategy is triggered, obtain a continuous risk response path network, match the scenario response after each strategy is triggered with the risk response target, all response trajectories cover all stages of risk evolution, and output the strategy trigger-risk scenario response trajectory set.
[0069] Based on the correspondence recorded in the risk scenario-strategy correlation trajectory set, the strategy parameters for each risk scenario are matched. The strategy parameter matching module accurately identifies the corresponding strategy parameters based on scenario and time information. After matching the strategy parameters, a scenario simulation model is used to deduce the scenario state response after strategy triggering. The scenario simulation model simulates the change process of the scenario state based on the physical characteristics of the risk scenario and the execution rules of the strategy. These risk response trajectories are recorded synchronously to construct a continuous risk response path network. It is ensured that the scenario response after each strategy trigger matches the risk response objective. This can be achieved by setting a risk response objective function and evaluation indicators to assess the scenario response. All response trajectories must cover all stages of risk evolution. These response trajectories are then organized and output to obtain the strategy trigger-risk scenario response trajectory set.
[0070] Step S2150: Analyze the strategy characteristics of the strategy trigger-risk scenario response trajectory set, optimize the trigger threshold of the strategy node by node to adapt to the evolution of the risk scenario, construct a continuous threshold optimization path network, simultaneously verify the correspondence between the optimized trigger threshold and the risk level, match the strategy optimization logic with the risk evolution law, and integrate to obtain the risk scenario strategy optimization trajectory set.
[0071] When analyzing the strategy characteristics of the set of strategy triggering-risk scenario response trajectories, data analysis and machine learning algorithms can be used. For example, correlation analysis can be used to find the relationship between strategy triggering and changes in scenario state, and cluster analysis can be used to classify strategies. Node-by-node optimization of the strategy trigger threshold can be achieved using optimization algorithms, adjusting the trigger threshold based on the evolution of the risk scenario and the strategy's execution effect. During the optimization process, a continuous threshold optimization path network is constructed, the threshold change process is recorded, and the correspondence between the optimized trigger threshold and the risk level is simultaneously verified. Through experiments and data analysis, a functional relationship between the two is established to ensure that the strategy optimization logic matches the risk evolution law. By studying and analyzing the risk evolution law, the optimization direction of the strategy is adjusted. Integrating this information yields the set of risk scenario strategy optimization trajectories.
[0072] Step S2160: Verify the adaptability of the risk scenario strategy optimization trajectory set, solidify the adapted strategy parameters node by node to obtain a continuous parameter solidification path network, and simultaneously verify the adaptability of the solidified strategy parameters to the risk scenario state changes. All verified parameters cover all risk evolution stages and are solidified into a global risk scenario strategy adaptation set.
[0073] When verifying the adaptability of the risk scenario strategy optimization trajectory set, various verification methods can be used, such as comparative experiments and simulations. By comparing the system's performance under different strategy parameters in risk scenarios, the adaptability of the strategies is evaluated. Adapted strategy parameters are solidified node by node. A data storage system can be used to save suitable strategy parameters. During the solidification process, a continuous parameter solidification path network is constructed, recording the process and results of parameter solidification. Simultaneously, the adaptability of the solidified strategy parameters to changes in the risk scenario state is verified. By monitoring changes in the risk scenario state and the execution effect of the strategies, the rationality of the parameters is judged, ensuring that all verified parameters cover all stages of risk evolution. These parameters are then integrated and solidified into a global risk scenario strategy adaptation set.
[0074] Step S2170: Integrate the full-domain risk scenario strategy adaptation set, covering the strategy logic of all risk scenarios, to obtain a continuous multi-scenario strategy network. Simultaneously verify the adaptability of the integrated strategy logic with all risk scenarios. The integrated strategy covers all risk types and is integrated into an initial multi-scenario adapted security control strategy cluster.
[0075] When integrating the comprehensive risk scenario strategy adaptation set, the adaptation strategies under different risk scenarios are uniformly managed and coordinated. By establishing a strategy integration model, the logic of each strategy is merged to form a complete strategy system. During the integration process, it is ensured that the strategy logic covers all risk scenarios, with corresponding response strategies for different types of risks. A continuous multi-scenario strategy network is constructed to demonstrate the correlation between different risk scenarios and strategies, simultaneously verifying the adaptability of the integrated strategy logic to all risk scenarios. Through simulation experiments and actual tests, the execution effect of the strategies under various risk scenarios is evaluated, ensuring that the integrated strategies cover all risk types. These strategies are then integrated to form an initial multi-scenario adapted security control strategy cluster.
[0076] Step S2180: Verify the global adaptability of the initial multi-scenario adapted security control strategy cluster, adjust the strategy parameters node by node to adapt to all risk scenarios, obtain a continuous strategy verification path network, and simultaneously verify the matching of the adjusted strategy parameters with all risk scenarios. All verified strategies cover all risk evolution stages, and generate a multi-scenario adapted security control strategy cluster that adapts to multiple risk scenarios.
[0077] When verifying the global adaptability of the initial multi-scenario adaptive security control strategy cluster, a combination of full-domain simulation and actual testing can be used. In full-domain simulation, various risk scenarios and task conditions are simulated to evaluate the adaptability of the strategy cluster. Strategy parameters are adjusted node by node, and the parameters are optimized and adjusted based on simulation and testing results. During the adjustment process, a continuous strategy verification path network is constructed, and the parameter adjustment process is recorded. The matching of the adjusted strategy parameters with all risk scenarios is verified simultaneously. By monitoring changes in the risk scenario state and the execution effect of the strategy, the rationality of the parameters is judged, ensuring that all verified strategies cover all risk evolution stages. These strategies are then integrated to generate a multi-scenario adaptive security control strategy cluster suitable for multiple risk scenarios.
[0078] Step S300: Based on the current maintenance task and real-time scene characteristics of the real railway maintenance scenario, match the corresponding target control strategy from the multi-scenario adaptive safety control strategy cluster to generate the robot's initial action execution sequence.
[0079] In one implementation, step S300 may specifically include the following steps S310 to S360:
[0080] Step S310: Collect the full-process operation nodes of the current maintenance task in the real railway maintenance scenario and the operation requirements of each node, and simultaneously collect the real-time status data of the scenario at all times. Associate and map the two with the multi-scenario adapted safety control strategy cluster, and output the task-scenario strategy matching candidate set. This set contains the correspondence between each strategy and the task node and scenario status.
[0081] When collecting data on the entire process of a real railway maintenance task, including all operational nodes and their requirements, detailed analysis and planning of the task are possible. For example, for a single equipment maintenance task, the operational flow can be defined, including equipment inspection, fault diagnosis, and component replacement, with specific operational requirements for each node, such as accuracy and time constraints. Real-time, all-day status data of the scenario is collected synchronously, utilizing a sensor network distributed across the maintenance site to monitor various aspects of the scenario, such as temperature, humidity, and pressure. The collected operational nodes, requirements, and real-time scenario status data are then mapped to a multi-scenario adaptive safety control strategy suite. This can be achieved by establishing a relational database, matching and recording each strategy with task nodes and scenario states, and outputting a candidate set of task-scenario strategy matching. This set contains the correspondence between each strategy and task nodes and scenario states, providing a foundation for subsequent evaluation of strategy adaptability.
[0082] Step S320: Extract the triggering conditions and execution logic of each strategy in the task-scenario strategy matching candidate set, associate the operation requirements of the corresponding task node with the constraints of the scenario state, verify the triggering feasibility and execution compliance of each strategy, and output the strategy adaptability evaluation set, which contains a quantitative description of the adaptability of each strategy.
[0083] In one implementation, step S320 may specifically include the following steps S321 to S326:
[0084] Step S321: Extract the trigger threshold and execution path of each strategy in the task-scenario strategy matching candidate set, associate the operation precision requirements of the corresponding task node with the limit constraints of the scenario state, compare the matching degree between the strategy trigger threshold and the scenario state one by one, and output the strategy trigger adaptation subset, which contains the trigger adaptation data of each strategy.
[0085] Extract the trigger threshold and execution path of each strategy from the task-scenario strategy matching candidate set. This can be done by analyzing the strategy's configuration file or code. The trigger threshold may include thresholds for parameters such as temperature, pressure, and time, while the execution path describes the steps and order of strategy execution. Associate the operational precision requirements of the corresponding task node with the extreme constraints of the scenario state, comparing the task's operational precision requirements and the scenario's extreme states with the strategy's trigger threshold. Compare the matching degree between each strategy's trigger threshold and the scenario state to determine whether the strategy can be triggered in the current scenario. For example, if the strategy's trigger threshold is a certain temperature value, and the current scenario's temperature is below that threshold, then the strategy cannot be triggered. During the comparison process, calculate a trigger fit score for each strategy and output a subset of strategy trigger fit scores, which contains the trigger fit data for each strategy.
[0086] Step S322: Extract the execution logic and operation process requirements of each strategy and task node, verify the fit between the strategy execution path and the task operation process one by one, and output a subset of strategy execution adaptability, which contains the execution adaptability data of each strategy.
[0087] The execution logic and task node operation requirements of each strategy are extracted. This can be analyzed through detailed descriptions of the strategy and the task's operation manual. The execution logic describes the specific steps and methods for strategy execution, while the task operation requirements specify the standard steps and order for completing the task. The fit between the strategy execution path and the task operation process is verified one by one to determine whether the strategy execution meets the task requirements. For example, if the task requires equipment inspection before fault diagnosis, but the strategy execution path is fault diagnosis first and then equipment inspection, then the strategy execution does not fit the task operation process. During the verification process, an execution fit score is calculated for each strategy, and a subset of strategy execution fit scores is output, containing the execution fit data for each strategy.
[0088] Step S323: Associate the policy triggering adaptation subset and the policy execution adaptation subset, calculate the overall adaptation of each policy in a unified manner, and output the overall adaptation description subset, which contains the overall adaptation value of each policy.
[0089] The algorithm combines the trigger adaptability subset and the execution adaptability subset of the associated strategy, taking both into account. A weighted sum of the trigger adaptability and execution adaptability can be calculated by setting weight coefficients to obtain the overall adaptability of each strategy. The overall adaptability of each strategy is calculated uniformly to ensure consistency and accuracy, and a subset describing the overall adaptability is output, containing the numerical value of the overall adaptability for each strategy.
[0090] Step S324: Associate the subset of comprehensive adaptability descriptions with the redundancy execution capability and resource consumption of the strategy, verify the redundancy adaptability and resource rationality of the strategy in turn, and output the additional adaptability subset of the strategy, which contains the additional adaptability data of each strategy.
[0091] The overall adaptability description subset is correlated with the redundancy execution capability and resource consumption of the strategy. This considers the reliability and resource utilization efficiency of the strategy during execution. Redundancy execution capability refers to the backup execution plan in case of failure or anomalies. Resource consumption includes the time, energy, and storage space required for strategy execution. The redundancy adaptability and resource rationality of the strategies are verified sequentially to determine whether the strategies have sufficient backup capabilities and reasonable resource consumption. During the verification process, an additional adaptability score is calculated for each strategy, and an additional adaptability subset containing the additional adaptability data for each strategy is output.
[0092] Step S325: Integrate the policy trigger adaptation subset, policy execution adaptation subset, comprehensive adaptation description subset, and policy additional adaptation subset, and uniformly sort out the complete adaptation information of each policy to output a complete adaptation description set.
[0093] The system integrates subsets of policy trigger adaptability, policy execution adaptability, comprehensive adaptability description, and additional policy adaptability, summarizing and organizing the information from each subset. It then unifies and compiles complete adaptability information for each policy, ensuring the completeness and accuracy of the information. For example, it merges trigger adaptability, execution adaptability, comprehensive adaptability, and additional adaptability information into a single data record, outputting a complete adaptability description set containing all the complete adaptability information for each policy.
[0094] Step S326: Based on the complete adaptation description set, sort the adaptation information of each strategy according to the adaptation priority of the task node, and output the strategy adaptation evaluation set, which contains the sorted complete adaptation description.
[0095] Based on a complete set of adaptation descriptions, the adaptation information for each strategy is sorted according to the adaptation priority of task nodes. The adaptation priority of task nodes can be set based on factors such as the importance and urgency of the task. For example, for a maintenance task of critical equipment, the adaptation priority of its operation nodes is higher. During the sorting process, the adaptation information of the strategies is ensured to be arranged according to priority, and a strategy adaptation evaluation set is output, which contains the sorted complete adaptation descriptions.
[0096] Step S330: Based on the full-process coverage of the task and the full-time adaptability of the scenario, select the top N strategies from the strategy adaptability evaluation set, where N>1, and output the initial selection set of target control strategies, which contains all candidate strategies with high adaptability.
[0097] Based on the core criteria of task-wide coverage and scenario-wide adaptability, strategies are selected from the strategy adaptability evaluation set. Task-wide coverage refers to the extent to which a strategy can cover all operational nodes of the maintenance task, while scenario-wide adaptability refers to the adaptability of a strategy throughout the entire scenario time period. The top N strategies are selected, with N > 1 to ensure that multiple candidate strategies with high adaptability are identified. An initial selection set of target control strategies can be output by setting a selection threshold or directly selecting the top-ranked strategies. This set contains all candidate strategies with high adaptability. "Higher" indicates adaptability greater than the preset selection threshold.
[0098] Step S340: Associate the initial set of target control strategies with the emergency operation requirements of the current maintenance task and the reserved space for sudden states in the real-time scenario, verify the emergency response capability and sudden state adaptability of each strategy, and output the target control strategy verification set, which contains strategies that have passed the emergency verification.
[0099] The initial set of target control strategies is correlated with the emergency operation requirements of the current maintenance task and the contingency reserve space of the real-time scenario. Emergency operation requirements refer to the emergency measures that the maintenance task needs to take in the event of an emergency, such as emergency shutdown and fault isolation. The contingency reserve space of the real-time scenario refers to the resources and conditions reserved in the scenario to cope with emergencies. The emergency response capability and contingency adaptability of each strategy are verified one by one to determine whether the strategy can effectively cope with emergencies. For example, it checks whether the strategy has an emergency handling plan for sudden failures and whether it can complete emergency operations within limited resources and time. A target control strategy verification set is output, which includes strategies that have passed the emergency verification, ensuring that these strategies can guarantee the safe conduct of maintenance work in the event of an emergency.
[0100] Step S350: Integrate the effective strategies in the target control strategy verification set, uniformly calibrate the triggering sequence and execution priority of each strategy, and output a target control strategy set that adapts to the current task and scenario. This set contains control strategies with clear timing and priority.
[0101] The effective strategies in the integrated target control strategy verification set are merged and organized through emergency verification. The triggering sequence and execution priority of each strategy are uniformly calibrated to ensure that the strategies can be triggered in the correct order and time. The strategies can be sorted and adjusted by setting the triggering time and priority coefficients, and the target control strategy set adapted to the current task and scenario is output. This set contains control strategies with clear timing and priority, which provides clear guidance for subsequent conversion into robot action commands.
[0102] Step S360: Map the set of target control strategies adapted to the current task and scenario to the action execution logic of the maintenance robot, transform the robot actions corresponding to each strategy one by one, sort the actions according to the task execution sequence, and output the robot's initial action execution sequence, which contains the action instructions corresponding to each task node.
[0103] In one implementation, step S360 may specifically include the following steps S361 to S366:
[0104] Step S361: Extract the execution instructions and timing requirements of each strategy in the target control strategy set adapted to the current task and scenario, associate the motion driving logic and joint range of motion of the maintenance robot, convert the strategy execution instructions into robot joint motion instructions one by one, and output a subset of robot joint motion instructions, which contains the joint motion instructions corresponding to each strategy.
[0105] The execution instructions and timing requirements of each strategy in the target control strategy set adapted to the current task and scenario are extracted and analyzed through detailed descriptions of the strategies. The execution instructions describe the specific tasks to be completed by the strategy, and the timing requirements specify the order in which the tasks are executed. The motion-driven logic and joint range of motion of the maintenance robot are correlated to understand how the robot's joints move and the limitations of their range of motion. The strategy execution instructions are converted one by one into robot joint motion instructions. Based on the robot's kinematic model and control algorithm, the strategy execution instructions are transformed into specific motion angles, velocities, and accelerations of the joints, outputting a subset of robot joint motion instructions. This subset contains the joint motion instructions corresponding to each strategy, providing a foundation for subsequent calibration of the joint motion instructions.
[0106] Step S362: Associate the subset of robot joint motion commands with the operational accuracy requirements of the current maintenance task, calibrate the motion amplitude and speed of each joint motion command one by one, and output the calibrated subset of joint motion commands, which contains the calibrated joint motion commands.
[0107] The robot joint motion command subset is associated with the operational accuracy requirements of the current maintenance task, clarifying the accuracy requirements for robot joint movement at each operational node. The motion amplitude and speed of each joint motion command are calibrated one by one, and adjustments are made to the joint motion amplitude and speed according to the operational accuracy requirements. For example, if the operation requires high joint motion accuracy, the motion speed is appropriately reduced to improve the accuracy of the motion, and a calibrated subset of joint motion commands is output, which contains the calibrated joint motion commands.
[0108] Step S363: Associate the spatial constraints of the calibrated subset of joint motion commands with the real-time scene, verify the motion path of each joint motion command to ensure it meets the scene space requirements, and output a spatially compliant subset of joint motion commands. This subset contains spatially compliant joint motion commands.
[0109] By correlating the calibrated subset of joint motion commands with the spatial constraints of the real-time scene, we can understand the limitations such as obstacles and space size in the scene. We then verify the motion path of each joint motion command to ensure it conforms to the scene's spatial requirements, determining whether the robot joints will collide with obstacles during movement. Collision detection and path planning can be performed by establishing a scene space model and a robot motion trajectory model. If the motion path of a joint motion command does not conform to the scene's spatial requirements, it is adjusted or replanned, outputting a spatially compliant subset of joint motion commands. This subset contains all spatially compliant joint motion commands.
[0110] Step S364: Integrate the spatially compliant joint motion command subset, sort the joint motion commands according to the triggering sequence and execution priority of the target control strategy set, and output the robot motion timing command set, which contains joint motion commands with clear timing and priority.
[0111] This process integrates a subset of space-compliant joint motion commands, merging and organizing all commands that meet space requirements. Joint motion commands are then sorted according to the triggering sequence and execution priority of the target control strategy set, ensuring that joint movements are executed in the correct order and with the correct priority. The commands can be sorted by setting timestamps and priority coefficients, outputting a set of robot motion timing commands that includes joint motion commands with clearly defined timing and priority.
[0112] Step S365: Associate the robot action timing instruction set with the overall collaborative logic of the maintenance robot, verify the collaborative feasibility and execution continuity of the action sequence one by one, and output a collaborative compliant action timing instruction set, which contains collaborative compliant action sequences.
[0113] By associating the robot's motion timing instruction set with the overall collaborative logic of the maintenance robot, we can understand the collaborative relationships and motion coordination requirements between the robot's various joints. We verify the collaborative feasibility and execution continuity of each motion sequence, determining whether the robot's joint movements can be executed in a coordinated and consistent manner. For example, we check whether the movements of multiple joints interfere with each other, and whether there are any stutters or discontinuities in the execution of the movements. If the motion sequence does not meet the collaboration requirements, we adjust or optimize it, outputting a collaboratively compliant motion timing instruction set, which contains collaboratively compliant motion sequences.
[0114] Step S366: Based on the set of collaborative compliant action sequence instructions, uniformly label the task node and triggering condition corresponding to each action, and output the robot's initial action execution sequence, which contains the action instructions corresponding to each task node.
[0115] Based on a set of collaborative and compliant action sequence instructions, each action is uniformly labeled with its corresponding task node and triggering conditions, clarifying the position and execution conditions of each action within the task. For example, it is labeled that a certain joint action is triggered when performing a device inspection task, with the triggering condition being the arrival of a specific time or state. This outputs the robot's initial action execution sequence, which contains the action instructions corresponding to each task node, providing a clear and accurate action guide for the operation of the maintenance robot.
[0116] Step S400: Map the real-time collected railway maintenance status data to the multi-physics digital twin training space for action pre-play risk verification, generate action verification adjustment instructions, optimize the robot's initial action execution sequence based on the action verification adjustment instructions, and output safe execution actions to drive the maintenance robot to complete the railway maintenance operation.
[0117] Real-time collected data on the actual status of railway maintenance reflects the current situation of railway maintenance scenarios, including equipment operating parameters and the physical state of the environment. Mapping this data to a multiphysics digital twin training space allows for the mapping of real data to virtual nodes within the digital twin space through the establishment of a data mapping model. Action pre-playing and risk verification are performed in the digital twin training space, simulating the robot's actions according to the initial action sequence to check for safety risks. Action verification and adjustment instructions are generated, and the robot's actions are adjusted based on the risk verification results. The initial action sequence of the robot is optimized based on the action verification and adjustment instructions, making the action sequence safer and more reliable. Safe execution actions are output to drive the maintenance robot to complete the railway maintenance operation, ensuring the smooth and safe conduct of the maintenance work.
[0118] As one implementation method, in step S400, the real-time collected actual railway maintenance status data is mapped to a multi-physics digital twin training space for action pre-simulation risk verification, which may specifically include the following steps S410~S460:
[0119] Step S410: Collect real-time railway maintenance status data, mark the evolution characteristics of the status data node by node, match the node mapping logic of the multi-physics digital twin training space, construct a continuous real-virtual node network, and the mapping path of each real status node fits the twin space logic to obtain a set of real-virtual node associated trajectories.
[0120] When collecting real-time railway maintenance status data, various sensors distributed at the railway maintenance site are used to monitor the operating status of equipment and the physical parameters of the environment in real time. For example, temperature sensors are used to measure the temperature of the equipment, and vibration sensors are used to detect the vibration of the equipment. The evolution characteristics of the status data are marked node by node, and the trend of the status data of each node over time is analyzed, such as the characteristics of data rise, fall, and fluctuation. The node mapping logic of the multi-physics digital twin training space is matched. By establishing node identifiers and mapping rules, the real status nodes are mapped to virtual nodes in the digital twin space. A continuous real-virtual node network is constructed to ensure that the mapping path of each real status node conforms to the twin space logic. For example, in the digital twin space, the state changes of virtual nodes should be consistent with the state changes of real nodes. The real-virtual node association trajectory set records the association information between real nodes and virtual nodes.
[0121] Step S420: Deduce the state evolution logic of the real-virtual node associated trajectory set and the multi-physics digital twin training space, map the real state data to the virtual space in time intervals, record the evolution trajectory of the virtual state synchronously, construct a continuous virtual state evolution network, the virtual state replicates the evolution of the real state, and generate a set of real state mapped virtual trajectories.
[0122] This study deduces the state evolution logic of the real-virtual node correlation trajectory set and the multiphysics digital twin training space, understanding how the state of virtual nodes in the digital twin space evolves with the state changes of real nodes. A state evolution model can be established to describe the relationship between the states of virtual and real nodes. Real state data is mapped to the virtual space in time intervals, and real-time collected real state data is transmitted to the corresponding virtual nodes in the digital twin space in chronological order. The evolution trajectory of the virtual state is recorded synchronously, and the state changes of the virtual nodes are recorded using a data acquisition and storage system. A continuous virtual state evolution network is constructed to ensure that the virtual state can replicate the evolution process of the real state. For example, if the temperature of a real node increases, the temperature of the virtual node should also increase accordingly, generating a set of real-state mapped virtual trajectories that records the evolution of real state data in the virtual space.
[0123] Step S430: Associate the set of virtual trajectories mapped from the real state with the robot's initial action execution sequence, rehearse the action execution process in the twin space time by time, record the virtual state changes in each action time period, obtain a continuous action rehearsal state network, the action rehearsal triggers the corresponding virtual state changes, and output the set of action rehearsal virtual state change trajectories.
[0124] By associating the set of virtual trajectories mapped from the real-world state to the robot's initial action execution sequence, the robot's action commands are linked to the node states in the virtual space. The action execution process is rehearsed time-by-time within the twin space, simulating the robot's actions in the virtual space according to the chronological order of the initial action execution sequence. The changes in the virtual state during each action time period are recorded, and the state of the virtual nodes is monitored in real time using sensors and a data acquisition system. A continuous action rehearsal state network is obtained, demonstrating the changes in the virtual state during the action rehearsal. Action rehearsal triggers corresponding virtual state changes to ensure the realism of the rehearsal. For example, if a robot action causes a change in the position of a virtual node, the corresponding virtual state should also be updated, outputting a set of virtual state change trajectories from the action rehearsal, which records the changes in the virtual state during the action rehearsal.
[0125] Step S440: Compare the set of virtual state change trajectories in the action pre-play with the safety boundary constraints of the multi-physics digital twin training space. Record the distance changes between the virtual state and the safety boundary node by node to obtain a continuous boundary distance change network. The correspondence between the state and the boundary distance is clear, and the change data of all boundary approaching nodes are fully covered to obtain the set of action pre-play boundary associated trajectories.
[0126] By comparing the set of virtual state change trajectories from the motion pre-play with the safety boundary constraints of the multiphysics digital twin training space, it is possible to check whether the virtual state is approaching or exceeding the safety boundary. The safety boundary constraints define the safe range of the virtual node states, such as upper limits for temperature and pressure. The distance change between the virtual state and the safety boundary is recorded node by node, and the distance change data is obtained by calculating the difference between the virtual state and the safety boundary. A continuous boundary distance change network is obtained, showing the trend of the distance change between the virtual state and the safety boundary. The correspondence between state and boundary distance is clear, ensuring accurate judgment of the safety of the virtual state. For example, if the virtual state approaches the safety boundary, the distance change data will show that the distance gradually decreases. The change data of all boundary-approaching nodes are fully covered. This information is integrated to obtain the motion pre-play boundary-related trajectory set, which records the association information between the virtual state and the safety boundary during the motion pre-play process.
[0127] Step S450: Analyze the risk characteristics of the action pre-play boundary associated trajectory set, evaluate the risk level of action execution node by node, obtain a continuous risk level evaluation network, the correspondence between risk level and boundary distance is clear, the evaluation of each risk level fits the boundary constraint logic, and form an action pre-play risk trajectory set.
[0128] The risk characteristics of the action pre-play boundary-related trajectory set are analyzed. Based on information such as the distance change and rate of change between the virtual state and the safety boundary, the potential risks of action execution are determined. For example, if the virtual state rapidly approaches the safety boundary, it indicates a high risk. The risk level of action execution is evaluated node by node. By setting risk level classification criteria, the risk level of nodes is divided into different levels based on risk characteristics. A continuous risk level assessment network is obtained, showing the risk level changes of each node. The correspondence between risk level and boundary distance is clear, ensuring the accuracy of risk assessment. For example, the closer to the safety boundary, the higher the risk level. The assessment of each risk level conforms to the boundary constraint logic. Risk assessment is performed according to the regulations and requirements of the safety boundary, forming a set of action pre-play risk trajectories. This set records the risk level of each node during the action pre-play process.
[0129] Step S460: Integrate the action pre-play risk trajectory set with the state evolution logic of the multi-physics digital twin training space, verify the risk level of the action pre-play in time periods, obtain a continuous risk verification network, and simultaneously verify the consistency between the risk assessment and the twin space logic. The risk verification results are consistent with the virtual space pre-play logic, and the action pre-play risk verification result set is obtained.
[0130] This method integrates the risk trajectory set of action pre-playing with the state evolution logic of a multi-physics digital twin training space, combining risk assessment results with the state evolution process of nodes in the virtual space. The risk level of the action pre-playing is verified time-by-time, and the acceptable risk of action execution in each time period is determined based on the risk level and state evolution. A continuous risk verification network is obtained, demonstrating the changes in risk level during action pre-playing. The consistency between risk assessment and twin space logic is simultaneously verified to ensure that the risk assessment results conform to the actual situation in the virtual space. For example, the risk assessment results should match the state change trends of virtual nodes. The risk verification results, aligned with the virtual space pre-playing logic, are integrated to obtain a set of action pre-play risk verification results, which records the risk verification results of the action pre-playing.
[0131] In one implementation, step S400 involves generating motion verification and adjustment instructions, optimizing the robot's initial motion execution sequence based on these instructions, and outputting safe execution actions to drive the maintenance robot to complete railway maintenance operations. Specifically, this may include the following steps S470~S4120:
[0132] Step S470: Extract the risk points from the action pre-play risk verification result set, locate the action execution path corresponding to each risk point node by node, obtain a continuous risk action path network, accurately locate the action path of each risk point, and obtain a risk action trajectory set.
[0133] Risk points are extracted from the action pre-play risk verification result set. By analyzing the risk verification results, high-risk node locations are identified. The action execution path corresponding to each risk point is located node by node. Based on the information recorded during the action pre-play, the specific action and execution path causing the risk are determined. A continuous network of risk action paths is obtained, displaying the action execution path corresponding to each risk point. Precise location of the action path for each risk point ensures accurate adjustment of risky actions. For example, if a risk point is determined to be caused by the movement of a certain joint of the robot, the risk action trajectory set records the execution path of the risky action.
[0134] Step S480: Adjust the action execution path of the risk action trajectory set, optimize the action path node by node to avoid the safety boundary of the twin space, obtain a continuous action path optimization network, and simultaneously verify the adaptability of the adjusted path to the boundary constraints. The adjusted action path avoids the boundary constraints and generates an action adjustment trajectory set.
[0135] When adjusting the execution paths of a set of risky action trajectories, a path planning algorithm can be used. This algorithm replans the execution paths of risky actions based on the safety boundary constraints of the twin space and the robot's kinematic model. The action paths are optimized node by node to ensure that the action path of each node avoids the safety boundary. During the optimization process, a continuous action path optimization network is constructed, and the steps and results of path adjustment are recorded. The adaptability of the adjusted path to the boundary constraints is verified simultaneously. Through simulation and analysis, it is determined whether the adjusted action path meets the requirements of the safety boundary. For example, it checks whether the adjusted path will cause the robot joints to collide with obstacles. The adjusted action paths avoid the boundary constraints, and this adjusted path information is integrated to generate a set of adjusted action trajectories.
[0136] Step S490: Deduce the state evolution logic of the action adjustment trajectory set and the multi-physics digital twin training space. Simulate the execution process of the adjusted action in the twin space time by time, record the virtual state changes of each action time period, obtain a continuous adjustment action pre-playing network, the action pre-playing triggers the corresponding virtual state changes, and outputs the action adjustment pre-playing state change trajectory set.
[0137] The algorithm combines the set of motion adjustment trajectories with the state evolution logic of the multiphysics digital twin training space, integrating the adjusted motion path with the state evolution rules of nodes in the virtual space. The algorithm simulates the execution process of the adjusted motion in the twin space time-by-time, mimicking the robot's actions according to the adjusted motion path and time sequence. The virtual state changes for each motion time period are recorded, and the state of virtual nodes is monitored in real time using sensors and a data acquisition system. This results in a continuous motion adjustment pre-play network, demonstrating the execution of the adjusted motion in the virtual space. Motion pre-play triggers corresponding virtual state changes, ensuring the realism of the simulation. For example, if the adjusted motion causes a change in the speed of a virtual node, the corresponding virtual state should also be updated, outputting a set of motion adjustment pre-play state change trajectories that record the state changes of the adjusted motion in the virtual space.
[0138] Step S4100: Verify the safety of the set of action adjustment pre-simulation state change trajectories, calibrate the action execution path node by node to avoid safety boundaries, obtain a continuous safe action calibration network, and simultaneously verify the consistency between the calibrated path and the boundary constraints. If the calibrated path meets the safety requirements, it is solidified into a set of safe action trajectories.
[0139] The safety of the set of pre-simulated state change trajectories for action adjustment is verified. By analyzing the virtual state change trajectories, it is determined whether the adjusted actions will lead to safety risks. Node-by-node action execution paths are calibrated to avoid safety boundaries. Based on the requirements of the safety boundaries and the changes in the virtual state, the action execution paths are further adjusted. During the calibration process, a continuous safe action calibration network is constructed, and the steps and results of path calibration are recorded. The consistency between the calibrated paths and boundary constraints is verified simultaneously to ensure that the calibrated paths fully comply with the requirements of the safety boundaries. For example, it is checked whether the calibrated paths will cause the temperature of virtual nodes to exceed the safety limit. Once the calibrated paths meet the safety requirements, this safe action path information is solidified to obtain a set of safe action trajectories.
[0140] Step S4110: Integrate the set of safe motion trajectories with the robot's initial motion execution sequence, replace the original motion path in time intervals to obtain a continuous motion sequence optimization network, and simultaneously verify the compatibility of the optimized sequence with the robot's driving logic. The optimized sequence conforms to the robot's operating logic, and the optimized robot motion execution sequence is obtained.
[0141] The system integrates a set of safe motion trajectories with the robot's initial motion execution sequence, replacing risky motion paths in the original sequence with safe motion paths from the safe motion trajectory set. Replacement is performed time-by-time to ensure the continuity of the motion sequence. During the replacement process, a continuous motion sequence optimization network is constructed, recording the steps and results of motion sequence adjustments. The compatibility of the optimized sequence with the robot's drive logic is simultaneously verified. Through simulation and experimentation, it is determined whether the optimized motion sequence can be correctly executed by the robot's drive system. For example, parameters such as velocity and acceleration of the motion sequence are checked to ensure they are within the allowable range of the robot's drive system. Once the optimized sequence conforms to the robot's operating logic, this optimized motion sequence information is integrated to obtain the optimized robot motion execution sequence.
[0142] Step S4120: Synchronize the optimized robot action execution sequence with the driving logic of the maintenance robot, output safe execution actions in time intervals to obtain a continuous action drive synchronization network, synchronously verify the matching of the action sequence and the driving logic, ensure that the action output matches the robot driving rhythm, and drive the maintenance robot to complete the railway maintenance operation.
[0143] The optimized robot motion execution sequence is synchronized with the maintenance robot's drive logic, matching the timing information and motion parameters of the motion sequence with the robot's drive system. Safe execution actions are output in time intervals, sending motion commands to the robot drive system according to the time sequence of the optimized motion execution sequence. During the output process, a continuous motion drive synchronization network is constructed to record the timing and parameters of the motion output. The matching between the motion sequence and the drive logic is verified synchronously by monitoring the robot's actual actions and the feedback information from the drive system to determine whether the motion sequence is consistent with the drive logic. The motion output matches the robot's drive rhythm, ensuring that the robot can execute actions smoothly and accurately. This drives the maintenance robot to complete railway maintenance operations, enabling the maintenance work to be carried out safely and efficiently.
[0144] Step S500: Feed the safety execution action data and real scene features back to the multi-physics digital twin training space and the meta-reinforcement learning training rules, and update the safety boundary parameters and multi-scene adapted safety control strategy clusters of the multi-physics digital twin training space respectively, to complete the dual-track iterative update of virtual-real fusion.
[0145] As one implementation method, the safety execution action data and real-world scene features are fed back to the multi-physics digital twin training space to update the safety boundary parameters of the multi-physics digital twin training space. Specifically, this may include the following steps S510~S560:
[0146] Step S510: Collect safe execution action data, record the temporal characteristics of action execution node by node, match the evolution trajectory of real scene features to obtain a continuous action scene interaction network, accurately record the scene interaction trajectory of each action, and obtain a set of action-scene state interaction trajectories.
[0147] When collecting data on safe execution actions, sensors and data acquisition devices installed on the maintenance robot are used to monitor the motion state of the robot's joints and the application of forces in real time. The temporal characteristics of each action execution are recorded node by node, including the start time, end time, and duration of the action. The evolutionary trajectory of real-world scene features is matched, and environmental sensors distributed throughout the maintenance site collect data on ambient temperature, humidity, and illumination, analyzing the trends of these data over time. The temporal characteristics of action execution are correlated with the evolutionary trajectory of real-world scene features to construct a continuous action-scene interaction network. For example, the robot's execution of a certain action at a specific ambient temperature is recorded. The scene interaction trajectory of each action is accurately recorded to ensure data accuracy; the action-scene state interaction trajectory set records the interaction information between the action and the scene state.
[0148] Step S520: Analyze the boundary impact of the action-scene state interaction trajectory set, record the effect of scene state changes on the safety boundary node by node, obtain a continuous scene boundary impact network, accurately evaluate the boundary impact of each scene change, and generate a scene-boundary associated trajectory set.
[0149] This study analyzes the boundary effects of action-scene-state interaction trajectory sets to investigate how changes in scene states affect the safety boundaries of a multiphysics digital twin training space. For example, an increase in ambient temperature may alter the safe operating temperature range of certain electronic devices. The impact of scene state changes on the safety boundary is recorded node by node. By comparing the changes in the safety boundary under different scene states, the degree of influence of scene changes on the boundary is determined. During the recording process, a continuous scene boundary influence network is constructed to demonstrate the relationship between scene state changes and safety boundary changes. The boundary impact of each scene change is accurately assessed using scientific evaluation methods and models to ensure the accuracy of the evaluation results. The scene-boundary correlation trajectory set records the correlation information between scene state changes and the safety boundary.
[0150] Step S530: Adjust the safety boundary parameters of the multiphysics digital twin training space, calibrate the boundary range node by node to adapt to the impact of scene changes, obtain a continuous boundary parameter calibration network, and simultaneously verify the adaptability of the adjusted boundary to scene changes. The adjusted boundary fits the scene state evolution, and the updated safety boundary constraint set is obtained.
[0151] When adjusting the safety boundary parameters of the multiphysics digital twin training space, the safety boundary range of each node is adjusted based on the information provided by the scene-boundary associated trajectory set. The boundary range is calibrated node by node to ensure it adapts to changes in scene state. During the adjustment process, a continuous boundary parameter calibration network is constructed, recording the steps and results of the boundary parameter adjustment. The adaptability of the adjusted boundary to scene changes is verified simultaneously. Through simulation and analysis, it is determined whether the adjusted boundary meets the requirements of scene state evolution. For example, it checks whether the adjusted temperature safety boundary can adapt to changes in ambient temperature. The adjusted boundary conforms to scene state evolution; these adjusted boundary parameter information are then integrated to obtain the updated safety boundary constraint set.
[0152] Step S540: Verify the global adaptability of the updated security boundary constraint set, verify the consistency between the boundary constraints and the twin space logic in time intervals, obtain a continuous boundary verification network, and simultaneously verify the matching of the boundary and the twin space interaction logic. The updated boundary conforms to the twin space operation logic, forming a global boundary update trajectory set.
[0153] To verify the global adaptability of the updated safety boundary constraint set, a combination of global simulation and actual testing is employed. In the global simulation, various scenarios and task conditions are modeled to evaluate the adaptability of the updated safety boundary constraint set throughout the entire twin space. The consistency between the boundary constraints and the twin space logic is verified time-by-time, checking whether the boundary constraints conform to the state evolution rules of nodes in the twin space. During the verification process, a continuous boundary verification network is constructed, recording the verification steps and results. The matching between the boundary and the twin space interaction logic is verified synchronously to ensure that the boundary settings do not affect the normal interaction between nodes in the twin space. For example, it is checked whether the boundary constraints would cause some nodes to be unable to transmit energy or information normally. The updated boundary conforms to the twin space operating logic. These verification information are integrated to form a global boundary update trajectory set, which records the adaptation and update process of the updated safety boundary throughout the entire twin space.
[0154] Step S550: Integrate the global boundary update trajectory set with the state evolution logic of the multi-physics digital twin training space, synchronize the boundary constraints and spatial interaction logic node by node to obtain a continuous spatial boundary synchronization network, synchronously verify the consistency between the boundary and the spatial logic, synchronize the boundary constraints of all nodes in the space, and update the multi-physics digital twin training space.
[0155] When integrating the global boundary update trajectory set with the state evolution logic of the multiphysics digital twin training space, the boundary update information is combined with the changes in node states over time and under the influence of multiphysics. Boundary constraints and spatial interaction logic are synchronized node by node to ensure that the boundary constraints of each node are consistent with the interaction rules between it and other nodes. During synchronization, data fusion algorithms and state evolution models are used to deeply integrate boundary constraints and spatial interaction logic. For example, for a pair of nodes with an energy transfer relationship, the boundary constraints must ensure that they do not hinder the reasonable transfer of energy between them. A continuous spatial boundary synchronization network is constructed, recording the specific steps of synchronization and the state changes of each node. The consistency between the boundary and spatial logic is verified through synchronization. By comparing and analyzing the performance of boundary constraints and spatial interaction logic at each node, it is ensured that they do not conflict. Once the boundary constraints of all nodes in the space are synchronized, the update of the multiphysics digital twin training space is completed, enabling it to more accurately reflect the actual situation of real railway sections under the influence of multiphysics and the new safety boundary requirements.
[0156] Step S560: Synchronize the updated multiphysics digital twin training space with the subsequent training logic, verify the adaptability of the spatial logic to the training requirements in time intervals, obtain a continuous spatial training synchronization network, synchronously verify the matching between the spatial update and the training task, ensure that the updated space meets the subsequent training requirements, and complete the safety boundary update of the twin space.
[0157] When synchronizing the updated multiphysics digital twin training space with subsequent training logic, it is necessary to clearly define the goals, tasks, and rules of subsequent training, and match this information with the updated spatial state and boundary constraints. The adaptability of the spatial logic to training requirements is verified time-by-time. At different training stages, it is checked whether the state evolution and boundary constraints of nodes in the space meet the requirements of the training task. For example, when training a maintenance robot to handle a certain fault scenario, it is verified whether the space can accurately simulate the multiphysics environment and safety boundaries under that fault scenario. During the verification process, a continuous spatial training synchronization network is constructed to record the verification results and changes in spatial state at each time period. The matching between the spatial update and the training task is verified synchronously, and it is evaluated whether the updated space can provide a suitable environment and conditions for training. For example, it is checked whether the safety boundaries of the space conform to the requirements for safe robot operation during training. When the updated space meets the requirements of subsequent training, the safety boundary update of the twin space is completed.
[0158] As one implementation method, the safety execution action data and real-world scene features are fed back to the meta-reinforcement learning training rules to update the multi-scene adapted safety control strategy cluster, completing the dual-track iterative update of virtual-real fusion. Specifically, this may include the following steps S570~S5120:
[0159] Step S570: Collect safe execution action data and real scene features, record the action-scene interaction features node by node, match the update logic of the meta-reinforcement learning training rules, obtain a continuous action rule interaction network, the rule adaptation of each interaction feature is accurate, and obtain the action-rule associated trajectory set.
[0160] When collecting data on safe execution actions and real-world scene characteristics, various sensors distributed throughout the railway maintenance site and on the robot are utilized. For example, accelerometers record changes in robot acceleration, and environmental sensors acquire features such as temperature and humidity. Action-scene interaction features are recorded node by node, analyzing the interaction methods and effects between actions and the scene at each node. For instance, the interaction between a robot's grasping action and factors such as the weight and position of objects in the scene at a specific location. The update logic of the meta-reinforcement learning training rules is matched, mapping the recorded action-scene interaction features to the rule update conditions according to the rules and algorithms of meta-reinforcement learning. During the matching process, a continuous action rule interaction network is constructed to demonstrate the connection between action-scene interaction features and rule updates. To ensure accurate rule adaptation for each interaction feature, precise data analysis and algorithmic calculations determine the rule update content corresponding to each interaction feature.
[0161] Step S580: Adjust the parameters of the multi-scenario adaptive security control strategy cluster, calibrate the strategy triggering conditions node by node to adapt to the action-scenario interaction features, obtain a continuous strategy parameter calibration network, and simultaneously verify the adaptability of the adjusted strategy and interaction features. The adjusted strategy fits the interaction features, and a set of strategy parameter adjustment trajectories is generated.
[0162] When adjusting the parameters of a multi-scenario adaptive safety control strategy cluster, information provided by the action-rule associated trajectory set is used. The strategy triggering conditions are calibrated node-by-node, optimizing the triggering conditions based on the action-scenario interaction characteristics at each node. For example, if a robot action in a certain scenario requires specific ambient temperature and object position conditions to be triggered, the triggering temperature threshold and position range of the strategy are adjusted accordingly. During the adjustment process, a continuous strategy parameter calibration network is constructed, recording the parameter adjustment steps and parameter changes at each node. The adaptability of the adjusted strategy to the interaction characteristics is simultaneously verified. Through simulation and actual testing, the execution effect of the adjusted strategy in actual action-scenario interactions is observed. For example, it checks whether the adjusted strategy can accurately trigger the robot's actions under suitable conditions. When the adjusted strategy matches the interaction characteristics, this adjustment information is integrated to generate a strategy parameter adjustment trajectory set, which records the detailed process and results of the strategy parameter adjustment.
[0163] Step S590: Verify the adaptability of the strategy parameter adjustment trajectory set with the updated multiphysics digital twin training space, verify the operational logic of the strategy in the twin space time by time, obtain a continuous strategy space verification network, and simultaneously verify the consistency between the strategy and the twin space logic. The adjusted strategy fits the twin space logic, forming a strategy adaptation verification trajectory set.
[0164] When verifying the adaptability of the adjusted strategy trajectory set to the updated multiphysics digital twin training space, the strategy's operation is simulated in the updated twin space. The strategy's operational logic within the twin space is verified time-by-time, checking whether the strategy can execute as expected at different time points and spatial nodes. For example, in a specific multiphysics environment, it's verified whether the strategy can correctly control the robot's movements. During verification, a continuous strategy space verification network is constructed, recording the verification results and strategy execution status for each time period. The consistency between the strategy and the twin space logic is verified synchronously, ensuring that the strategy's execution does not conflict with the state evolution and boundary constraints of nodes in the twin space. For example, it's checked whether the strategy causes the robot's movements to exceed safety boundaries. When the adjusted strategy fits the twin space logic, this verification information is integrated to form a strategy adaptation verification trajectory set, which records the strategy's adaptation status and verification process within the updated twin space.
[0165] Step S5100: Integrate the policy adaptation verification trajectory set with the multi-scenario adapted security control policy cluster, update the policy parameters node by node to obtain a continuous policy update network, and simultaneously verify the adaptability of the updated policy to multiple risk scenarios. The updated policy covers multiple risk scenarios, and the updated multi-scenario adapted security control policy cluster is obtained.
[0166] When integrating the strategy adaptation verification trajectory set with the multi-scenario adapted safety control strategy cluster, the verified strategy parameter adjustment information is incorporated into the original strategy cluster. Strategy parameters are updated node by node based on the results recorded in the strategy adaptation verification trajectory set. During the update process, a continuous strategy update network is constructed, recording the parameter update steps and parameter changes at each node. The adaptability of the updated strategy to multiple risk scenarios is verified simultaneously. Various risk scenario simulations and actual tests are used to evaluate the execution effect of the updated strategy under different risk conditions. For example, simulating electrical and mechanical faults that may occur during railway maintenance checks whether the strategy can effectively address them. When the updated strategy covers multiple risk scenarios, this updated strategy information is integrated to obtain an updated multi-scenario adapted safety control strategy cluster. This cluster can better adapt to various actual risk scenarios, improving the safety and reliability of the maintenance robot.
[0167] Step S5110: Synchronize the updated multi-scenario adaptation security control strategy cluster and the updated multi-physics digital twin training space, verify the interaction logic of strategy-space in time intervals, obtain a continuous virtual-real interaction network, synchronously verify the consistency of interaction logic, synchronize the interaction between strategy and space, and generate a set of virtual-real fusion dual-track update trajectories.
[0168] When synchronizing the updated multi-scenario adaptive security control strategy cluster with the updated multiphysics digital twin training space, the strategies in the strategy cluster are associated with the nodes and boundary constraints in the twin space. The interaction logic between the strategy and space is verified time-by-time. At different time points, it is checked whether the execution of the strategy has a reasonable impact on the state of the twin space, and whether changes in the twin space are reflected in the adjustment of the strategy. For example, when a node in the twin space undergoes a state change due to the multiphysics effect, it is verified whether the strategy can respond in a timely manner. During the verification process, a continuous virtual-real interaction network is constructed to record the interaction between the strategy and space and the state changes in each time period. The consistency of the interaction logic is verified synchronously to ensure that the interaction rules and logic between the strategy and space are coordinated and consistent. For example, it is checked whether the triggering conditions of the strategy match the security boundary constraints of the twin space. When the interaction between the strategy and space is synchronized, this interaction information is integrated to generate a virtual-real fusion dual-track update trajectory set, which records the interaction and synchronization between the strategy cluster and the twin space during the update process.
[0169] Step S5120: Based on the virtual-real fusion dual-track update trajectory set, following the iterative logic of meta-reinforcement learning, advance the collaborative update process of the twin space and policy cluster, and verify the effectiveness of this process in improving the model's generalization and adaptation capabilities, thus completing the virtual-real fusion dual-track iterative update.
[0170] Based on the virtual-real dual-track update trajectory set, and following the iterative logic of meta-reinforcement learning, the twin space and policy clusters are collaboratively updated. The iterative logic of meta-reinforcement learning involves continuously adjusting the model's parameters and policies based on new data and feedback information to improve model performance. During the collaborative update process, the state evolution model of the twin space and the parameters of the policy clusters are further optimized based on the interaction between the policies and the space recorded in the virtual-real dual-track update trajectory set. For example, if the policy's performance is found to be poor in certain scenarios, the policy parameters are adjusted; if the simulation results of the twin space deviate from the actual situation, the physical parameters and boundary constraints of the space are updated. Throughout the update process, the effectiveness of this process in improving the model's generalization and adaptability is continuously verified. Cross-validation and real-world application testing are used to evaluate the performance of the updated model in different scenarios and tasks. For example, the adaptability and control effect of the model can be tested in new railway maintenance scenarios. When the verification results show that the process can effectively improve the generalization and adaptation capabilities of the model, the dual-track iterative update of virtual and real integration is completed. This enables the reinforcement learning-based electrified railway maintenance robot control method to better cope with various complex actual situations and improve the efficiency and safety of railway maintenance.
[0171] Please see Figure 2 , Figure 2 This is a schematic diagram of a computer system provided in an embodiment of the present invention. The computer system includes at least a processor 101, a communication interface 102, and a memory 103. The processor 101, communication interface 102, and memory 103 can be connected via a bus or other means. The processor 101 (or Central Processing Unit, CPU) is the computing and control core of the computer system, capable of parsing various instructions and processing various data within the computer system. The communication interface 102 may optionally include a standard wired interface or a wireless interface (such as Wi-Fi, mobile communication interface, etc.), and can be used to send and receive data under the control of the processor 101; the communication interface 102 can also be used for data transmission and interaction within the computer system. The memory 103 is a storage device in the computer system used to store programs and data. It is understood that the memory 103 here can include the computer system's built-in memory, or it can include extended memory supported by the computer system. The memory 103 provides storage space, which stores the computer system's operating system; this invention does not limit this storage space.
[0172] In one embodiment, the processor 101 executes the reinforcement learning-based control method for electrified railway maintenance robots provided in the above embodiments of the present invention by running a computer program in the memory 103.
Claims
1. A control method for an electrified railway maintenance robot based on reinforcement learning, characterized in that, The method includes: Based on the physical structure attributes of railway sections, the attributes of high-voltage electromagnetic effects, and the attributes of airflow vibration effects, a multi-physics digital twin training space with safety boundary constraints is constructed, with high-voltage electric field safety threshold and mechanical collision safety threshold as constraints. Random fault disturbance attributes are loaded to simulate the disturbance of the multi-physics digital twin training space, generating a virtual-real coupled training data set containing safety boundary association information. The virtual-real coupled training data set is processed based on the meta-reinforcement learning training strategy, and safety transfer adaptation learning is performed by combining the interactive data of real railway maintenance scenarios to generate a multi-scenario adaptive safety control strategy cluster that is adapted to multiple risk scenarios. Based on the current maintenance task and real-time scene characteristics of the real railway maintenance scenario, the corresponding target control strategy is matched from the multi-scenario adaptive safety control strategy cluster to generate the robot's initial action execution sequence. The real-time collected railway maintenance status data is mapped to the multi-physics digital twin training space for action pre-simulation risk verification, and action verification adjustment instructions are generated. Based on the action verification adjustment instructions, the initial action execution sequence of the robot is optimized, and safe execution actions are output to drive the maintenance robot to complete the railway maintenance operation. The safety execution action data and real scene features are fed back to the multi-physics digital twin training space and the meta-reinforcement learning training rules, respectively updating the safety boundary parameters of the multi-physics digital twin training space and the multi-scene adaptive safety control strategy cluster, thus completing the dual-track iterative update of virtual-real fusion. The aforementioned construction of a multi-physics digital twin training space with safety boundary constraints, based on the physical structural properties of railway sections, the properties of high-voltage electromagnetic effects, and the properties of airflow vibration, with high-voltage electric field safety thresholds and mechanical collision safety thresholds as constraints, includes: Track the energy transfer trajectory of all nodes in the physical structure of the railway section, mark the energy input and output ports of each node, and record the energy flow rate and direction of each port simultaneously. Combine the full-domain influence range of high-voltage electromagnetic action to analyze the influence of electromagnetic field on the energy state of structural nodes, and obtain the set of structural energy state trajectories under the action of electromagnetic field, covering all main and branch structures of the railway section. Based on the set of structural energy state trajectories under the action of the electromagnetic field, and combined with the global influence range of airflow vibration, the coupling effect of electromagnetic energy and vibration energy and the comprehensive intensity change are analyzed node by node to obtain the node state evolution path network under the coupling effect of multiple physics fields. The state record of each node is updated synchronously to obtain the set of global multi-physics field coupling state trajectories. The set of coupled state trajectories of the global multi-physics field is associated with the safety threshold of the high-voltage electric field. The change characteristics of the comprehensive energy state of the node as it approaches the threshold are recorded node by node. A continuous network of state change paths of the threshold-approaching nodes is obtained. The correlation between the state change characteristics of each node and the electric field strength is analyzed synchronously, and the set of node state trajectories under the electric field safety constraint is output. Based on the set of node state trajectories under the electric field safety constraints, combined with the mechanical collision safety threshold, the abnormal node state patterns when the structural deformation approaches the threshold are recorded node by node, resulting in a continuous abnormal path network of deformed nodes approaching the threshold. The correlation between the abnormal state patterns of each node and the degree of structural deformation is analyzed simultaneously and solidified into a set of node state trajectories under collision safety constraints. By integrating and analyzing the set of node state trajectories under electric field safety constraints and the set of node state trajectories under collision safety constraints, the evolution law of node state under the combined action of electric field and collision safety constraints is identified node by node, and the node state evolution path network under the coupled action of dual safety constraints is obtained. The correlation between the identified evolution law and the intensity of dual constraint action is verified simultaneously, and integrated into a set of node state trajectories coupled by dual safety constraints. Based on the set of state trajectories of coupled nodes with dual security constraints, the node mapping and state evolution logic of the digital twin space is constructed to ensure that the virtual state evolution of nodes in the twin space can reflect the multi-physics field coupling state and security constraints of real physical nodes, thus obtaining a multi-physics field digital twin training space with security boundary constraints.
2. The method according to claim 1, characterized in that, The loading of random fault disturbance attributes simulates the disturbance in the multiphysics digital twin training space, generating a virtual-real coupled training data set containing security boundary correlation information, including: The simulated triggering points of random fault disturbances are determined and mapped to the corresponding nodes in the multiphysics digital twin training space. Based on the physical association logic of the nodes in the twin space, the transmission path of the fault state is deduced to obtain the transmission path network of the fault in the twin space. The transmission path of each fault triggering point covers the relevant nodes affected by it, and the fault-space node transmission trajectory set is obtained. The energy transfer logic of the fault-space node propagation trajectory set and the multiphysics digital twin training space is deduced. The fault propagation process is simulated in the space time period by time, and the spatial state change data of each time period is recorded synchronously. The state change data of all time periods are synchronized with the fault propagation trajectory, and the fault disturbance virtual state trajectory sequence is output. By comparing the virtual state trajectory sequence of the fault disturbance with the historical fault state trajectory of the real railway section, the trajectory deviation points and deviation degrees of the virtual and real states are recorded node by node, and a continuous deviation node change path network is constructed to obtain the set of virtual and real state deviation trajectories. The fault disturbance virtual state trajectory sequence is calibrated, the deviation points of the virtual and real state deviation trajectory set are aligned, the evolution trajectory of the virtual state is adjusted node by node, the overlap between the adjusted virtual state and the real state trajectory is verified synchronously, the adjusted virtual state and the real state trajectory overlap, and the virtual state and the real state trajectory overlap, and the virtual and real state trajectory sequence is solidified. By associating the virtual-real fusion state trajectory sequence with the safety boundary constraints of the multi-physics digital twin training space, the trajectory deflection amplitude when the state approaches the boundary is recorded node by node, resulting in a continuous boundary approach node interaction path network. The deflection amplitude corresponds one-to-one with the boundary distance, and the trajectory data of all boundary approach nodes are fully covered and integrated into a state-boundary associated trajectory set. The state-boundary associated trajectory set and the virtual-real fusion state trajectory sequence are labeled, and the corresponding boundary association information is marked at each state node to obtain a continuous labeled state trajectory network. The labeling information is consistent with the boundary constraints and covers all state evolution periods, generating a virtual-real coupled training data set containing safe boundary association information.
3. The method according to claim 1, characterized in that, The meta-reinforcement learning-based training strategy processes the virtual-real coupled training data set, including: Track the policy update sequence of meta-reinforcement learning, associate the state change sequence of the virtual-real coupled training data set, record the correspondence between the state change time period and the policy update time period in each time period, and obtain a continuous time period interaction path network. The correspondence of each time period is synchronized with the state change logic, and a policy-state time period associated trajectory set is obtained. The strategy triggers the strategy for each state change period in the strategy-state time period associated trajectory set, synchronously records the state response trajectory after the strategy is triggered, and obtains a continuous trigger-response path network. The state response after each strategy is triggered matches the strategy objective, and all response trajectories cover all stages of state change. The strategy trigger-state response trajectory set is then output. The policy adjustment gradient of the policy trigger-state response trajectory set is calculated, the policy optimization logic of the association meta-reinforcement learning is used, the evolution path of gradient adjustment is recorded time-by-time to obtain a continuous gradient adjustment path network, the correspondence between gradient adjustment magnitude and state response deviation is simultaneously calibrated, the gradient adjustment logic is matched with the policy optimization objective, and the policy gradient update trajectory set is obtained. Iterate through the policy gradient update trajectory set and the global state change trajectory of the virtual-real coupled training data set. Apply the policy gradient to adjust the policy parameters in each global state cycle to obtain a continuous parameter adjustment path network. Simultaneously verify the adaptability of the adjusted policy parameters to the state changes. The parameter adjustment in each cycle covers all state nodes and is solidified into a global policy parameter optimization trajectory set. Verify the policy adaptability of the global policy parameter optimization trajectory set, solidify the adapted policy parameters node by node to obtain a continuous parameter solidification path network, and simultaneously verify the consistency between the solidified policy parameters and the state change time sequence. All verified parameters cover all state change time periods and are integrated into an iteratively optimized policy parameter set. By integrating the state change trajectories of the iteratively optimized policy parameter set and the virtual-real coupled training data set, the policy triggering sequence and the state change sequence are synchronized to obtain a continuous policy triggering path network. The matching between the policy triggering sequence and the state change sequence is verified simultaneously, and the policy triggering path of all nodes is fully covered, generating a meta-trained and adapted policy parameter set.
4. The method according to claim 3, characterized in that, The safety transfer adaptation learning, which combines interactive data from real railway maintenance scenarios, includes: Collect interactive data trajectories from real railway maintenance scenarios, match the triggering sequence of the policy parameter set after meta-training and adaptation, record the correspondence between interactive time periods and policy triggering time periods for each time period, obtain a continuous interactive time period path network, and synchronize the correspondence of each time period with the scenario interaction logic to obtain a scenario interaction-policy time period associated trajectory set. The strategy for each interaction period in the set of scene interaction-strategy time period associated trajectory is triggered, and the scene state response after the strategy is triggered is synchronously deduced to obtain a continuous scene response path network. The scene response after each strategy is triggered is matched with the task objective. All response trajectories cover all stages of scene interaction and output the set of strategy trigger-scene response trajectories. By comparing the set of policy-triggered scene response trajectories with the state response trajectories of the virtual-real coupled training data set, the deviation points between scene response and virtual state response are recorded node by node to obtain a continuous response deviation path network. The degree of deviation of each deviation point corresponds one-to-one with the scene features. The deviation points cover all scene interaction stages, thus obtaining a set of virtual-real response deviation trajectories. The meta-trained and adapted policy parameter set is calibrated, the deviation points of the virtual-real response deviation trajectory set are aligned, the trigger parameters of the policy are adjusted node by node, a continuous parameter adjustment path network is obtained, the consistency between the adjusted policy response and the virtual state response during the scene interaction period is verified synchronously, the adjusted policy is adapted to the scene interaction logic, and solidified into the scene-adapted policy parameter trajectory set. Verify the global adaptability of the set of strategy parameter trajectories after scene adaptation. Generalize the adapted strategy parameters in each global interaction cycle to obtain a continuous generalized parameter path network. Simultaneously verify the adaptability of the generalized strategy parameters to scene state changes. All verified parameters cover all scene interaction periods and are integrated into a set of strategy parameters after global scene adaptation. By integrating the policy parameter set after global scene adaptation with the transfer learning logic of meta-reinforcement learning, the policy adaptation time series and transfer cycle are synchronized to obtain a continuous transfer stage path network. The matching of the policy adaptation time series and transfer cycle is verified simultaneously, and the policy parameters of all transfer stages are fully covered, generating a policy parameter set after transfer adaptation.
5. The method according to claim 4, characterized in that, The generated multi-scenario adaptation security control strategy cluster includes: Extract the state change trajectory of multiple risk scenarios, classify the evolution characteristics of different risk scenarios, match the triggering sequence of the policy parameter set after migration and adaptation, and obtain a continuous scenario policy path network. The corresponding path of each scenario is synchronized with the risk evolution logic, and the risk scenario-policy associated trajectory set is integrated. Match the strategy parameters of each risk scenario in the risk scenario-strategy associated trajectory set, synchronously deduce the scenario state response after the strategy is triggered, obtain a continuous risk response path network, the scenario response after each strategy is triggered matches the risk response target, all response trajectories cover all stages of risk evolution, and output the strategy trigger-risk scenario response trajectory set. The strategy characteristics of the strategy trigger-risk scenario response trajectory set are analyzed, the trigger threshold of the strategy is optimized node by node to adapt to the evolution of risk scenarios, a continuous threshold optimization path network is constructed, the correspondence between the optimized trigger threshold and the risk level is verified simultaneously, the strategy optimization logic is matched with the risk evolution law, and the risk scenario strategy optimization trajectory set is integrated. Verify the adaptability of the risk scenario strategy optimization trajectory set, solidify the adapted strategy parameters node by node to obtain a continuous parameter solidification path network, and simultaneously verify the adaptability of the solidified strategy parameters to the risk scenario state changes. All verified parameters cover all risk evolution stages and are solidified into a global risk scenario strategy adaptation set. The comprehensive risk scenario strategy adaptation set is integrated to cover the strategy logic of all risk scenarios, resulting in a continuous multi-scenario strategy network. The compatibility of the integrated strategy logic with all risk scenarios is verified simultaneously. The integrated strategy covers all risk types and is integrated into an initial multi-scenario adapted security control strategy cluster. Verify the global adaptability of the initial multi-scenario adapted security control strategy cluster, adjust the strategy parameters node by node to adapt to all risk scenarios, obtain a continuous strategy verification path network, and simultaneously verify the matching of the adjusted strategy parameters with all risk scenarios. All verified strategies cover all risk evolution stages, and generate a multi-scenario adapted security control strategy cluster that adapts to multiple risk scenarios.
6. The method as described in claim 1, characterized in that, The step involves matching the corresponding target control strategy from the multi-scenario adaptive safety control strategy cluster based on the current maintenance task and real-time scenario characteristics of a real railway maintenance scenario, and generating the robot's initial action execution sequence, including: Collect the full-process operation nodes of the current maintenance task in a real railway maintenance scenario and the operation requirements of each node, and simultaneously collect the real-time scenario status data for all time periods. Associate and map the two with the multi-scenario adapted safety control strategy cluster, and output a task-scenario strategy matching candidate set, which contains the correspondence between each strategy and task nodes and scenario status. Extract the triggering conditions and execution logic of each strategy in the task-scenario strategy matching candidate set, associate the operation requirements of the corresponding task node with the constraints of the scenario state, verify the triggering feasibility and execution compliance of each strategy, and output a strategy adaptability evaluation set, which contains a quantitative description of the adaptability of each strategy. Based on the full-process coverage of tasks and the adaptability of scenarios throughout the time period, the top N strategies are selected from the strategy adaptability evaluation set, where N>1, and the initial selection set of target control strategies is output. This set contains all candidate strategies whose adaptability is greater than the preset screening threshold. The initial set of target control strategies is associated with the emergency operation requirements of the current maintenance task and the reserved space for sudden situations in the real-time scenario. The emergency response capability and adaptability to sudden situations of each strategy are verified one by one, and the target control strategy verification set is output, which contains the strategies that have passed the emergency verification. Integrate the effective strategies in the target control strategy verification set, uniformly calibrate the triggering sequence and execution priority of each strategy, and output a target control strategy set that adapts to the current task and scenario. This set contains control strategies with clear timing and priority. The set of target control strategies adapted to the current task and scenario is mapped to the action execution logic of the maintenance robot. The robot actions corresponding to each strategy are transformed one by one, and the actions are sorted according to the task execution sequence. The initial action execution sequence of the robot is output, which contains the action instructions corresponding to each task node.
7. The method as described in claim 6, characterized in that, The extraction of the task-scenario strategy matching candidate set includes the triggering conditions and execution logic of each strategy, which are then associated with the operation requirements of the corresponding task node and the constraints of the scenario state. The triggering feasibility and execution compliance of each strategy are verified one by one, and a strategy suitability evaluation set is output, including: Extract the trigger threshold and execution path of each strategy in the task-scenario strategy matching candidate set, associate the operation precision requirements of the corresponding task node with the limit constraints of the scenario state, compare the matching degree between the strategy trigger threshold and the scenario state one by one, and output a subset of strategy trigger adaptability, which contains the trigger adaptability data of each strategy. Extract the execution logic and task node operation process requirements of each strategy, verify the fit between the strategy execution path and the task operation process one by one, and output a subset of strategy execution adaptability, which contains the execution adaptability data of each strategy. The policy trigger adaptation subset and policy execution adaptation subset are associated, and the overall adaptation of each policy is calculated in a unified manner. The overall adaptation description subset is output, which contains the overall adaptation value of each policy. The comprehensive adaptability description subset is associated with the redundancy execution capability and resource consumption of the strategy. The redundancy adaptability and resource rationality of the strategy are verified in turn, and the additional adaptability subset of the strategy is output. This subset contains the additional adaptability data of each strategy. Integrate the policy trigger adaptation subset, policy execution adaptation subset, comprehensive adaptation description subset, and policy additional adaptation subset, and uniformly sort out the complete adaptation information of each policy to output a complete set of adaptation descriptions; Based on the complete set of adaptation descriptions, the adaptation information of each strategy is sorted according to the adaptation priority of the task node, and the strategy adaptation evaluation set is output, which contains the sorted complete adaptation descriptions. The process involves mapping the set of target control strategies adapted to the current task and scenario to the action execution logic of the maintenance robot, transforming the robot actions corresponding to each strategy one by one, sorting the actions according to the task execution sequence, and outputting the initial action execution sequence of the robot, including: Extract the execution instructions and timing requirements of each strategy in the target control strategy set adapted to the current task and scenario, associate the motion driving logic and joint range of motion of the maintenance robot, convert the strategy execution instructions into robot joint motion instructions one by one, and output a subset of robot joint motion instructions, which contains the joint motion instructions corresponding to each strategy. The robot's joint motion command subset is associated with the operational accuracy requirements of the current maintenance task. The motion amplitude and speed of each joint motion command are calibrated one by one, and the calibrated joint motion command subset is output. This subset contains the calibrated joint motion commands. The spatial constraints of the real-time scene are correlated with the subset of joint motion commands after calibration. The motion path of each joint motion command is checked one by one to see if it meets the scene space requirements. The spatially compliant subset of joint motion commands is output, which contains spatially compliant joint motion commands. Integrate a subset of space-compliant joint motion commands, sort the joint motion commands according to the trigger timing and execution priority of the target control strategy set, and output a set of robot motion timing commands, which contains the timing-sorted joint motion commands; The robot action timing instruction set is associated with the overall collaborative logic of the maintenance robot. The collaborative feasibility and execution continuity of the action sequence are verified one by one, and a collaborative compliant action timing instruction set is output, which contains collaborative compliant action sequences. Based on the set of collaborative and compliant action sequence instructions, the task nodes and triggering conditions corresponding to each action are uniformly labeled, and the robot's initial action execution sequence is output. This sequence contains the correspondence between task nodes, triggering conditions and action instructions.
8. The method according to claim 1, characterized in that, The process of mapping real-time collected railway maintenance status data to the multi-physics digital twin training space for action pre-simulation risk verification includes: Real-time railway maintenance status data is collected, the evolution characteristics of the status data are marked node by node, and the node mapping logic of the multi-physics digital twin training space is matched to construct a continuous real-virtual node network. The mapping path of each real status node fits the twin space logic, and a set of real-virtual node associated trajectories is obtained. The state evolution logic of the real-virtual node associated trajectory set and the multi-physics digital twin training space is deduced. Real state data is mapped to virtual space in time intervals, and the evolution trajectory of virtual state is recorded synchronously to construct a continuous virtual state evolution network. The virtual state replicates the evolution of real state and generates a set of real state mapped virtual trajectory. The set of virtual trajectories mapped from the real state is associated with the robot's initial action execution sequence. The action execution process is pre-enacted in the twin space time by time, and the virtual state changes in each action time are recorded to obtain a continuous action pre-enactment state network. The action pre-enactment triggers the corresponding virtual state changes and outputs the set of action pre-enactment virtual state change trajectories. By comparing the set of virtual state change trajectories in the action pre-play with the safety boundary constraints of the multi-physics digital twin training space, the distance change between the virtual state and the safety boundary is recorded node by node, resulting in a continuous boundary distance change network. The correspondence between the state and the boundary distance is clear, and the change data of all boundary approaching nodes are fully covered, thus obtaining the set of action pre-play boundary associated trajectories. The risk characteristics of the action pre-play boundary associated trajectory set are analyzed, and the risk level of action execution is evaluated node by node to obtain a continuous risk level evaluation network. The correspondence between risk level and boundary distance is clear, and the evaluation of each risk level fits the boundary constraint logic, forming a set of action pre-play risk trajectories. By integrating the set of risk trajectories for action pre-play with the state evolution logic of the multi-physics digital twin training space, the risk level of action pre-play is verified time by time to obtain a continuous risk verification network. The consistency between risk assessment and twin space logic is verified simultaneously. The risk verification results are consistent with the virtual space pre-play logic, resulting in a set of action pre-play risk verification results.
9. A computer system, characterized in that, include: A memory, wherein a computer program is stored; A processor is configured to load the computer program to implement the reinforcement learning-based control method for an electrified railway maintenance robot as described in any one of claims 1-8.