Rail maintenance method, device and equipment and storage medium
Through the combination of digital twin model and reinforcement learning model, the status data of track components are obtained and maintenance strategies are decided, which solves the problem of inaccurate track maintenance in the existing technology, and efficient and accurate track repair is achieved, avoiding casualties and property losses.
Patent Information
- Application Number
- CN202510335164.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-04
AI Technical Summary
The prediction accuracy of existing track maintenance methods is not high, which can easily lead to track operation failures, resulting in casualties and property losses.
By obtaining the on-site data of the track component, using the digital twin model to extract state data, and combining the reinforcement learning model to output accurate maintenance strategies based on service life and reward function decision-making goal maintenance strategies.
It improves the efficiency and accuracy of track repair, reduces casualties and property losses, and realizes timely detection and precise maintenance of track defects.
Smart Images

Figure CN120258765A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of electronic technology, and in particular to a track maintenance method, device, equipment and storage medium. Background Art
[0002] As the speed of rail vehicles and loads increases, the demand for track maintenance increases. Tracks can refer to any type of track, such as railway tracks, high-speed rail tracks, tram tracks, subway tracks, etc. However, track maintenance is very complex and difficult. For example, the track consists of many components, and different components affect each other, such as the influence relationship between rails and sleepers. When one of the two is abnormal, the track will operate abnormally. If the abnormalities of the components of the rails are discovered in time, maintenance can be carried out in a timely manner.
[0003] At present, the more common track maintenance method is generally to use machine learning models to predict track defects and alert maintenance personnel to track defects, and then the maintenance personnel will perform maintenance manually. However, the current track prediction accuracy is not high, which can easily cause track operation failures, resulting in casualties and property losses. Therefore, how to perform efficient and accurate track maintenance is a technical problem that needs to be solved urgently. Summary of the invention
[0004] The present disclosure provides a track maintenance method, device, equipment and storage medium to achieve timely detection of track defects and acquisition of accurate maintenance strategies for the defects, thereby improving track repair efficiency and accuracy and effectively avoiding casualties and property losses.
[0005] In a first aspect, the present disclosure provides a track maintenance method, comprising:
[0006] Acquire field data of at least one track component of the track to be maintained, the field data including relevant data of each track component, the relevant data including at least one of the following: geometric parameters, component defects, service life or historical maintenance data;
[0007] The field data of each track component is input into the digital twin model respectively, and the status data of each track component is extracted through the digital twin model, and the status data includes the component category and defect data of the corresponding track component;
[0008] The status data of each track component is input into the reinforcement learning model. The reinforcement learning model decides the target maintenance strategy for each track component based on the service life of each track component and the reward function. The reward function is based on the maintenance cost and the occurrence of defects.
[0009] Output the target maintenance strategy for each track component.
[0010] In some embodiments, the geometric parameters include at least one of the following: measurement position, superelevation, first longitudinal level, second longitudinal level, first plane, second plane, gauge, and twist; the superelevation refers to the height difference between the inner and outer rail surfaces of the track in a curved section, the first longitudinal level refers to a 10-meter chord of the longitudinal level, specifically, within the range of a 10-meter chord length, the longitudinal unevenness of the track; the second longitudinal level refers to a 20-meter chord of the longitudinal level, specifically, within the range of a 20-meter chord length, the longitudinal unevenness of the track; the first plane refers to a 10-meter chord of the plane, specifically, within the range of a 10-meter chord length, the lateral position unevenness of the track; the second plane refers to a 20-meter chord of the plane, specifically, within the range of a 20-meter chord length, the lateral position unevenness of the track.
[0011] Component defects include at least one of the following: multiple track defect information, where the track defect information refers to the defect information of any track component and the date and location where the defect information exists. The defect information of the component includes the component type and / or defect type. The component type includes at least one of the following: ballast bed, fastener, rail, sleeper, turnout, and crossing. The defect type includes at least one of the following: track geometric parameter defect, track component defect. The track geometric parameter defect includes at least one of the following: abnormal superelevation, abnormal longitudinal level, abnormal plane, abnormal gauge, abnormal twist. The track component defect includes at least one of the following: abnormal ballast bed, abnormal fastener, abnormal rail, abnormal sleeper, turnout and crossing anomaly;
[0012] Historical maintenance data includes at least one of the following: the maintained track components, specific maintenance strategies, and maintenance time. The specific maintenance strategies include at least one of the following: tamping, rail grinding, ballast cleaning, sleeper replacement, rail replacement, fastener replacement, or ballast unloading.
[0013] In some embodiments, the state data of each track component is respectively input into the reinforcement learning model, including:
[0014] According to at least one maintenance strategy, determine at least one maintenance strategy that the track to be maintained can have. The maintenance strategy includes one or more maintenance strategies selected from at least one maintenance strategy. The at least one maintenance strategy includes at least one of the following: tamping, rail grinding, ballast cleaning, sleeper replacement, rail replacement, fastener replacement, or ballast unloading.
[0015] Input the state data of each track component and at least one maintenance strategy into the trained reinforcement learning model, and the reinforcement learning model selects the target maintenance strategy for each track component from at least one maintenance strategy.
[0016] In some embodiments, the reinforcement learning model is trained with the highest total maintenance benefit of each track component as the training objective. The training steps of reinforcement learning include:
[0017] Obtain multiple pieces of training data, where the training data includes the status data of the track components participating in the training;
[0018] Take at least one maintenance strategy as the state of the reinforcement learning model respectively, input each piece of training data into the reinforcement learning model, and select the first maintenance strategy for the track components corresponding to each piece of training data through the reinforcement learning model;
[0019] Take the highest total maintenance benefit obtained by the first maintenance strategy of the track components corresponding to each piece of training data as the iteration target, and iteratively train the reinforcement learning model until the reinforcement learning model meets the iteration target.
[0020] In some embodiments, the total policy benefit of the reinforcement learning model refers to the sum of the maintenance benefits of each track component;
[0021] The maintenance benefit of a track component refers to the sum of the maintenance cost of the track component and the maintenance benefit that can be obtained by executing the target maintenance strategy when a defect of the track component occurs;
[0022] The maintenance cost includes the replacement cost corresponding to the remaining service life of the track component.
[0023] In some embodiments, the maintenance benefit of the track component is obtained by calculating through a reward function, and the reward function is as follows: reward = w1*C + w2*D, where reward is the reward result of the reward function, C is the maintenance cost, D is the maintenance benefit, w1 is the weight of the maintenance cost, and w2 is the weight of the maintenance benefit.
[0024] In some embodiments, it further includes:
[0025] After each track component performs track maintenance according to the corresponding target maintenance strategy, collect the track maintenance results of each track component;
[0026] Update the on-site data of the track to be maintained according to the maintenance results of each track component.
[0027] In a second aspect, the present disclosure provides a track maintenance device, including:
[0028] A data acquisition unit for obtaining on-site data of at least one track component of the track to be maintained, where the on-site data includes relevant data of each track component, and the relevant data includes at least one of the following: geometric parameters, component defects, service life, or historical maintenance data;
[0029] A state extraction unit for inputting the on-site data of each track component into the digital twin model respectively, and extracting the state data of each track component through the digital twin model, where the state data includes the component category and defect data of the corresponding track component;
[0030] A policy acquisition unit is configured to input the status data of each track component into a reinforcement learning model respectively. The reinforcement learning model determines the target maintenance policy for each track component based on the service life of each track component and a reward function, and the reward function is determined based on the maintenance cost and the occurrence of defects.
[0031] A policy prompt unit is configured to output the target maintenance policy of each track component.
[0032] In a third aspect, the present disclosure provides a computer device, including a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps of the method in the above aspect.
[0033] In a fourth aspect, the present disclosure provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method in the above aspect are implemented.
[0034] In a fifth aspect, the present disclosure provides a computer program product, including a computer program / instructions. When the computer program is executed by a processor, the steps of the method in the above aspect are implemented.
[0035] A track maintenance method, device, equipment, and storage medium provided by the present disclosure collect the on-site data of the track to be maintained in real time, and combine with a digital twin model to analyze the usage status of each track component in a timely and effective manner, and obtain the status data of each track component. The status data of each track component can be used as the input of the reinforcement learning model, and the reinforcement learning model can determine the target maintenance policy for each track component based on the service life of each track component and a reward function, and the reward function is determined based on the maintenance cost and the occurrence of defects, so that the reinforcement learning model determines the target maintenance policy with the highest reward for the track component, and realizes the accurate selection of the target maintenance policy of each track component. To detect the defects of the track in time, and obtain the accurate maintenance policy for the defects, improve the track repair efficiency and accuracy, and effectively avoid casualties and property losses. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] The present disclosure will be described in more detail below based on embodiments and with reference to the accompanying drawings:
[0037] Figure 1 It is a schematic diagram of a track maintenance system provided by an embodiment of the present disclosure;
[0038] Figure 2 It is a flowchart of a track maintenance method provided by an embodiment of the present disclosure;
[0039] Figure 3 It is a flowchart of another track maintenance method provided by an embodiment of the present application;
[0040] Figure 4 Flowchart of a training method for a reinforcement learning model provided by an embodiment of the present application;
[0041] Figure 5 Workflow diagram of a reinforcement learning model provided by an embodiment of the present application;
[0042] Figure 6 Schematic structural diagram of an orbit maintenance device provided by an embodiment of the present application;
[0043] Figure 7 Schematic block diagram of a computer device provided by an embodiment of the present application.
[0044] In the drawings, the same components are denoted by the same reference numerals, and the drawings are not drawn to actual scale. Detailed implementation manners
[0045] In order to enable those skilled in the art of the present technology to better understand the technical solutions of the present disclosure, and to fully understand how the present disclosure applies technical means to solve technical problems and the implementation process of achieving corresponding technical effects and to implement accordingly, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all of the embodiments. The embodiments of the present disclosure and each feature in the embodiments can be combined with each other without conflict, and the formed technical solutions are all within the protection scope of the present disclosure. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present disclosure.
[0046] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0047] It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from that here.
[0048] The technical solution disclosed in the present invention can be applied to track maintenance scenarios. By collecting the on-site data of the track to be maintained in real time and combining it with the digital twin model, the use status of each track component can be analyzed in a timely and effective manner to obtain the status data of each track component. The status data of each track component can be used as the input of the reinforcement learning model, so as to use the reinforcement learning model to accurately select the target maintenance strategy for each track component. In this way, the defects of the track can be detected in time, and accurate maintenance strategies can be acquired for the defects, so as to improve the efficiency and accuracy of track repair, and effectively avoid casualties and property losses.
[0049] In related technologies, in order to achieve track repair, machine learning models are used to predict track defects and alert maintenance personnel to track defects, who then perform maintenance manually. However, the current track prediction accuracy is not high, which can easily cause track operation failures, resulting in casualties and property losses. Therefore, how to perform efficient and accurate track maintenance is a technical problem that needs to be solved urgently.
[0050] In order to solve the above technical problems, in the technical solution of this application, the field data of at least one track component of the track to be maintained is collected, and the field data can cover the geometric parameters, component defects and historical maintenance data of each track component. The data types are relatively diverse, providing comprehensive and detailed original materials for subsequent analysis. The field data is input into the digital twin model, and the field data is deeply analyzed by the digital twin model to output the status data of each track component. The defect data is sorted and integrated by the twin model, which can present the full picture of the problem more accurately and systematically, allowing the staff to quickly locate the key fault points. The reinforcement learning model can decide the target maintenance strategy for each track component based on the service life and reward function of each track component. The reward function is determined based on the maintenance cost and the occurrence of defects, so that the reinforcement learning model decides the target maintenance strategy with the highest reward for the track component, and gives the best maintenance plan that fits the current status of the track component, and achieves the highest maintenance effect with the lowest maintenance cost, so that the target maintenance strategy of each track component provides a clear and direct action guide for track maintenance personnel. By timely detecting the defects of the track and obtaining accurate maintenance strategies for the defects, the efficiency and accuracy of track repair are improved, which can effectively avoid casualties and property losses.
[0051] Figure 1 This is an example diagram of a track maintenance system provided in an embodiment of the present disclosure. The track maintenance system may include: a data acquisition device 10 , a data processing device 20 , a data storage device 30 , a first device 40 , and a second device 50 .
[0052] Among them, the data acquisition device 10 can be used to acquire data such as the geometric parameters of the track, inspection reports, and maintenance records. The data processing device 20 can perform data analysis on the data obtained by the data acquisition device 10, such as steps of data cleaning, preprocessing, mining and sharing, etc., to extract the on-site data of the track to be maintained. The on-site data can include the geometric parameters of the track components, component defects, service life, and / or historical maintenance data.
[0053] The data storage device 30 can store the on-site data of the track to be maintained. The first device 40 can read the on-site data of the track to be maintained from the data storage device 30, and input the on-site data of each track component into the digital twin model respectively. The state data of each track component can be extracted through the digital twin model. The state data includes the component category and defect data of the corresponding track component.
[0054] After that, the first device 40 can input the state data of each track component into the second device 50. In the second device 50, the state data of each track component can be input into the reinforcement learning model respectively, and the target maintenance strategy of each track component can be predicted through the reinforcement learning model. The target maintenance strategy of each track component can be output for the maintenance personnel to view.
[0055] Among them, the reinforcement learning model makes decisions on the target maintenance strategy for each track component based on the service life and reward function of each track component, and the reward function is determined based on the maintenance cost and the occurrence of defects.
[0056] Of course, each device in the above track maintenance system, such as the data acquisition device 10, the data processing device 20, the data storage device 30, and the first device 40 and the second device 50, can be independent devices respectively, or one or more devices can belong to the same device. For example, the data acquisition device 10, the data processing device 20, and the data storage device 30 can be the same device, and the first device 40 and the second device 50 can also be the same device. Another example is that the data acquisition device 10 can be one device, the data processing device 20 and the data storage device 30 can be another device, and the first device 40 and the second device 50 can also be the same device. Another example is that the data acquisition device 10, the data processing device 20, the data storage device 30, and the first device 40 and the second device 50 can be the same device. In this embodiment, there is no excessive limitation on the devices where the above devices are located and the number of devices. Figure 1 The shown device distribution method is also exemplary and does not constitute a specific limitation.
[0057] Figure 2 It is a schematic flowchart of a track maintenance method provided by an embodiment of the present disclosure. As Figure 1 shown, a track maintenance method includes:
[0058] S201. Obtain on-site data of at least one track component of the track to be maintained. The on-site data includes relevant data of each track component, and the relevant data includes at least one of the following: geometric parameters, component defects, service life, or historical maintenance data.
[0059] Among them, the track to be maintained may include at least one track component. Real-time data of railway infrastructure, such as track geometric parameters and railway component defects, can be collected through devices such as sensors, cameras, and radars. The data acquisition device can receive the implementation data of the above railway infrastructure and can also collect data such as inspection reports and maintenance records of the track. The data acquisition device can send the acquired data to the data processing device.
[0060] The data processing device can perform data cleaning, preprocessing, and analysis on the data, extract useful features and information, obtain the on-site data of the track to be maintained, and send the on-site data of the track to be maintained to the data storage device. The data storage device can store the on-site data of the track to be maintained.
[0061] Optionally, S201 may include: The electronic device sends a data reading request of the track to be maintained to the data storage device. The data storage device responds to the data reading request, queries the on-site data of the track to be maintained, and sends the queried on-site data of the track to be maintained to the electronic device. The electronic device can receive the on-site data of the track to be maintained.
[0062] Optionally, the track to be maintained can be divided into several track components (which can also be called cross-sections or segment tracks). The length of each track component can be 10 meters, and the service life of each track component can be obtained by evaluating the geometric parameters and component defects of the track component. After normalizing the service life of each track component, it is used as the input of the reinforcement learning model. That is, the service life of each track component can be represented by a numerical value in the interval [0, 1]. For example, 0 represents complete failure, 1 represents completely normal, and 0.8 represents the need for maintenance.
[0063] In addition, the service life of each track component can also participate in the calculation process of maintenance benefits. For example, the replacement cost or repair cost corresponding to the remaining service life of each track component can be determined according to the service life of each track component. For example, the service life of the maintenance component is 0.7, and the remaining service life is 1 - 0.7 = 0.3. The replacement cost and repair cost of the track component with a remaining service life of 0.3 can be 20.
[0064] S202. Input the on-site data of each track component into the digital twin model respectively, and extract the state data of each track component through the digital twin model. The state data includes the component category and defect data of the corresponding track component.
[0065] Among them, the component category can be any one of the ballast bed, fasteners, rails, sleepers, switches and crossings.
[0066] The defect data can include the current defect information of the track components, specifically referring to the current specific defect types of the track. The specific defect types can include track geometric parameter defects and / or track component defects. Track geometric parameter defects include one or more of abnormal superelevation, longitudinal level anomaly, planar anomaly, gauge anomaly, and twist anomaly. Track component defects can include one or more of ballast bed anomaly, fastener anomaly, rail anomaly, sleeper anomaly, switch and crossing anomaly.
[0067] For example, the component category of a track component is the ballast bed. The specific defect types of the track component - ballast bed can include 3 defect types such as abnormal superelevation and twist, and ballast bed anomaly.
[0068] In addition, the defect data can also include: rail status data and maintenance decision support data. Among them, the rail status data refers to the real-time information of the track geometric parameters and component defects after digital processing (for example, it can include the current defect information mentioned above), reflecting the current health status and performance indicators of the track. The maintenance decision support data refers to the input data and model analysis results, providing suggestions and optimization schemes for maintenance activities, including the timing, type, and priority of maintenance, etc.
[0069] It can be understood that the digital twin model can truly reflect the implementation status of each track component. The digital twin model can store and manage the input data to provide data support during the training and application process of the reinforcement learning model. In addition, the digital twin model can also store the output results of the reinforcement learning model (such as optimized maintenance activity suggestions) for recording and further model training.
[0070] Optionally, the digital twin model can analyze the on-site data of each track component to obtain the status data of each track component. Specifically, it includes the following modules:
[0071] 1. Data integration and synchronization module: used to integrate data from different sources to ensure data consistency and real-time performance, providing an accurate basis for subsequent analysis and decision-making.
[0072] 2. Data cleaning and preprocessing module: used to clean the original data, remove noise and outliers, fill or interpolate missing data, making the data more complete and reliable.
[0073] 3. Status evaluation and analysis module: used to evaluate the geometric parameters and component defects of the track using the model, analyze the operating status and potential problems of the track, and identify key areas that need attention and maintenance.
[0074] 4. Historical data storage and traceability module: used to store historical data, facilitating the analysis and traceability of the long-term change trends and maintenance effects of the track, and providing references for optimizing maintenance strategies.
[0075] The analysis process involves the judgment of whether there are defects in the track. For track component defects, it can directly judge whether there are defects in the track components. If so, the defects are determined. For track geometric parameter defects, the geometric parameters of the track can be compared with the thresholds. If the parameters exceed the thresholds, they are regarded as defects.
[0076] Except that the correlation between gauge and other track geometric defects is relatively low, other track geometric defects are highly correlated. This indicates that if a track section has some track geometric defects (except gauge), it often also has other defects. Compared with track geometric defects, the correlation of track component defects seems to be lower, which means that track component defects seem to be more independent. However, it is worth noting that the correlation between turnout, crossing and fastener defects and track geometric defects is very high, with a correlation greater than 0.90. This finding will help the maintenance party investigate track defects and formulate maintenance plans.
[0077] The data and information that can be linked and stored in the digital twin model include track position, track cross-section, track geometry, track component defects or maintenance activities. Then, the data stored in the digital twin model is connected to the developed reinforcement learning model.
[0078] The data involved in this application is relatively complex. For data that is not suitable for storage in the digital twin model, such as data with large sizes or long sequences, certain parts of the data can be called through the digital twin model in a hyperlink manner. To store data in the digital twin model, the design of the model must support the data to be stored.
[0079] S203. Input the status data of each track component into the reinforcement learning model respectively. The reinforcement learning model determines the target maintenance strategy for each track component based on the service life and reward function of each track component, and the reward function is determined based on the maintenance cost and the occurrence of defects.
[0080] Optionally, the reinforcement learning model can judge whether maintenance is required according to the status data of each track component. Specifically, the reinforcement learning model will use the data extracted from the digital twin model to decide the appropriate maintenance activities to be performed during the next maintenance period. The maintenance activities recommended by the reinforcement learning model can be used to compile a maintenance plan or a maintenance report.
[0081] It can be understood that the components of reinforcement learning mainly include five parts, namely, the agent, the environment, the state, the action, and the reward. The agent interacts with the environment by taking actions and obtaining rewards. The environment refers to everything except the agent. In addition, the environment also defines the rules and properties of the reinforcement model. The state is the time step of the information of the environment or provided to the agent. The action refers to how the agent interacts with the environment. Depending on the environment, the action can be discrete or continuous. In this application, the action specifically refers to the maintenance strategy executed on the track components. The reward is a scalar value obtained by the agent from the environment and can be calculated using the reward function. It represents the success of the agent.
[0082] After taking the maintenance action, the environment will generate a new set of states based on the on-site data and consider the appropriate values for each state. This process will be repeated until the training ends.
[0083] When the training of the reinforcement learning model ends, the target maintenance strategy for each track component can be determined using the trained reinforcement learning model. Specifically, the agent (also known as the intelligent agent) of the reinforcement learning model can be used to determine the action with the highest maintenance efficiency (i.e., the target maintenance strategy) for each track component.
[0084] Furthermore, S203 may include: inputting the state data of each track component into the reinforcement learning model respectively to obtain the target maintenance strategy determined by the reinforcement learning model for each track component based on the service life and the reward function of each track component. Specifically, the maintenance decision of the reinforcement learning model is executed according to the service life and the reward function of the track section. That is, the service life and the reward function of the track section can be used as the basis for influencing the target maintenance strategy determined by the reinforcement learning model for the track component.
[0085] For example, for any track component, the reinforcement learning model can perform the following steps: execute multiple decisions to obtain multiple candidate maintenance strategies for the track component decision, evaluate each candidate maintenance decision according to the service life and the reward function of the track section of the track component to obtain the evaluation scores corresponding to the multiple candidate maintenance strategies respectively, and determine the candidate maintenance strategy with the highest evaluation score as the target maintenance strategy of the track component.
[0086] Of course, in a possible design, there can be N target maintenance strategies. The evaluation scores of the candidate maintenance strategies can be sorted, and the top N candidate maintenance strategies with the highest evaluation scores can be selected as the N target maintenance strategies, where N is a positive integer greater than or equal to 1.
[0087] The results of the reinforcement learning model can better respond to the current situation of the track section. This can reduce unnecessary maintenance activities, thus saving maintenance costs and time. At the same time, the number of track geometry and track component defects is significantly reduced, and track component defects are more common than track geometry defects.
[0088] In the reinforcement learning model, in terms of sensitivity analysis and hyperparameter tuning, there are three hyperparameters worthy of attention, namely the learning rate, loss ε, and discount factor. After analysis, it is found that the smaller the learning rate, the better the performance of the reinforcement learning model. ε also has this characteristic. When log(ε) > -2 or ε > 0.01, the loss will also increase rapidly. However, this does not happen in the case of the discount factor, and the discount factor has no effect on the loss. In this embodiment, the above hyperparameters can be appropriately adjusted and optimized to improve the prediction effect of the reinforcement learning model.
[0089] Optionally, the state (i.e., maintenance strategy) and maintenance actions of the reinforcement learning model are stored in the digital twin model through the property set definition. In the property set definition, the state and operations can be defined according to their types, such as real numbers, integers, or the true value (true)-false value (false) of binary states (defect found or not found / maintenance executed or not executed).
[0090] S204. Output the target maintenance strategy for each track component.
[0091] Optionally, the target maintenance decisions for each track component can be divided into different types. For example, the target maintenance strategy can be empty, that is, no maintenance is required. The target maintenance strategy can refer to the maintenance of a local area of the track component, and this target maintenance strategy is local maintenance. The target maintenance strategy can refer to the overall maintenance of the track component, and this target maintenance strategy is comprehensive maintenance.
[0092] It can be understood that the target maintenance strategy for the track component can include at least one maintenance action to be performed on the target track component. The target maintenance strategy is composed of this at least one maintenance action. The track component is repaired through at least one maintenance action in the target maintenance strategy.
[0093] In the technical solution of this application, on-site data of at least one track component of the track to be maintained is collected. The on-site data can cover the geometric parameters, component defects, and historical maintenance data of each track component. The data types are relatively diverse, providing comprehensive and detailed original materials for subsequent analysis. By inputting the on-site data into the digital twin model, the on-site data is deeply analyzed through the digital twin model, and the status data of each track component is output. After the defect data is sorted and integrated by the twin model, the overall picture of the problem can be presented more accurately and systematically, enabling the staff to quickly locate the key fault points. The status data is input into the reinforcement learning model to predict the target maintenance strategy, and the best maintenance plan that suits the current condition of the track components is given, rather than simply following fixed rules. The target maintenance strategies of each track component are output, providing clear and direct action guidelines for track maintenance personnel. By detecting the defects of the track in a timely manner and obtaining accurate maintenance strategies for the defects, the track repair efficiency and accuracy are improved, and casualties and property losses can be effectively avoided.
[0094] The embodiments of this application are mainly applied to the track maintenance scenario, and the maintenance of the track relies on rich data. Therefore, the embodiments of this application provide a rich data surface. The following embodiments introduce the specific composition of various types of data.
[0095] 1. The geometric parameters include at least one of the following: measurement position, superelevation, first longitudinal level, second longitudinal level, first plane, second plane, gauge, and twist.
[0096] Among them, the first longitudinal level is, for example, a 10-meter chord, the second longitudinal level is, for example, a 20-meter chord, the first plane is, for example, a 10-meter chord, the second plane is, for example, a 20-meter chord, the gauge and twist are, for example, a 20-meter chord. Of course, the above parameters can also take other values, and this embodiment does not limit them too much.
[0097] Among them, superelevation refers to the height difference between the inner and outer rail surfaces of the track in the curve section. Its purpose is to balance the centrifugal force through the centripetal force, making the train run more stable and safe in the curve section. By maintaining or trimming the superelevation, the lateral force of the train in the curve section is reduced, the running stability and safety of the train are improved, and the wear of the track and the vehicle is reduced.
[0098] The first longitudinal level, that is, the longitudinal level (10-meter chord), refers to the longitudinal unevenness of the track within the range of a 10-meter chord length. By measuring the height change of the track within 10 meters, the smoothness of the track is evaluated. By maintaining or trimming the longitudinal level, the smoothness of the train running in the straight section is ensured, the vertical vibration of the train is reduced, the comfort of passengers is improved, and the service life of the track and the vehicle is extended.
[0099] The second longitudinal level, i.e., the longitudinal level (20-meter chord), has a similar meaning to the 10-meter chord, but the measurement range is 20 meters. That is, within the 20-meter chord length, the longitudinal unevenness of the track. It provides an evaluation of the longitudinal unevenness of the track over a longer range. By maintaining or trimming the longitudinal level (20-meter chord), a more comprehensive evaluation of track smoothness can be provided, especially in terms of long-wavelength unevenness, which helps to identify and address a wider range of track problems.
[0100] The first plane, i.e., the plane (10-meter chord), refers to the lateral position unevenness of the track within the 10-meter chord length. By measuring the lateral offset of the track within 10 meters, the straightness of the track is evaluated. By maintaining or trimming the plane (10-meter chord), the stability of the train when running on a straight section can be ensured, reducing the lateral vibration of the train, improving the comfort of passengers, and reducing the wear of the track and the vehicle.
[0101] The second plane, i.e., the plane (20-meter chord), is similar to the 10-meter chord, but the measurement range is 20 meters. That is, within the 20-meter chord length, the lateral position unevenness of the track. It provides an evaluation of the lateral position unevenness of the track over a longer range. By maintaining or trimming the plane (20-meter chord), a more comprehensive evaluation of track straightness can be provided, especially in terms of long-wavelength unevenness, which helps to identify and address a wider range of track problems.
[0102] The gauge is the distance between the inner sides of the two rails of the track. The standard gauge is usually 1435 millimeters (4 feet 8.5 inches). By maintaining or trimming the gauge, it can be ensured that the train wheels can run correctly on the track, providing stable support, preventing derailment, and ensuring the safe operation of the train.
[0103] Twist (20-meter chord) refers to the combined effect of the lateral and longitudinal unevenness of the track within the 20-meter chord length. It evaluates the degree of torsion of the track within 20 meters. By maintaining or trimming the twist (20-meter chord), it can be ensured that the geometric shape of the track remains consistent over a long range, reducing the complex vibration of the train during operation, and improving the running stability and safety of the train.
[0104] 2. Component defects include at least one of the following: multiple track defect information. Track defect information refers to the defect information of any track component, as well as the date and location where the defect information exists. The defect information of the component includes the component type and / or defect type. The component type includes at least one of the following: roadbed, fastener, rail, sleeper, turnout, and crossing. The defect type includes at least one of the following: track geometric parameter defect, track component defect. The track geometric parameter defect includes at least one of the following: superelevation anomaly, longitudinal level anomaly, plane anomaly, gauge anomaly, twist anomaly. The track component defect includes at least one of the following: roadbed anomaly, fastener anomaly, rail anomaly, sleeper anomaly, turnout and crossing anomaly.
[0105] Among them, the superelevation represents that it may be necessary to replace the sleepers or unload the roadbed to adjust the geometry of the track. The longitudinal level anomaly (such as the first longitudinal level for a 10-meter chord and the second longitudinal level for a 20-meter chord) represents that it may be necessary to tamp or clean the roadbed to improve the level state of the track. The plane anomaly (such as the first plane for a 10-meter chord and the second plane for a 20-meter chord) represents that it may be necessary to tamp or clean the roadbed to improve the plane state of the track. The gauge anomaly represents that it may be necessary to adjust the gauge or replace the sleepers to correct the gauge. The twist anomaly (such as the twist for a 20-meter chord) represents that it may be necessary to replace the rails or sleepers to eliminate the twist.
[0106] Among them, the roadbed anomaly represents that it may be necessary to clean or unload the roadbed to repair the roadbed defect. The fastener anomaly represents that it may be necessary to replace the fasteners to repair the fastener defect. The rail anomaly represents that it may be necessary to grind or replace the rails to repair the rail defect. The sleeper anomaly represents that it may be necessary to replace the sleepers to repair the sleeper defect. The turnout and crossing anomaly represents that specific maintenance activities for the turnout and crossing may be required, such as turnout grinding or turnout replacement.
[0107] 3. The historical maintenance data includes at least one of the following: the maintained track components, the specific maintenance strategy, and the maintenance time. The specific maintenance strategy includes one or more maintenance actions.
[0108] In this embodiment, by defining detailed data for various parameters or data, it is possible to have more detailed and comprehensive data support during the track maintenance process, obtain a more accurate and effective maintenance strategy, thereby improving the scientificity, accuracy, and efficiency of the track maintenance decision-making as a whole, reducing the uncertainty of manual experience judgment, and ensuring the stable and safe operation of the track system.
[0109] Figure 3 It is a flowchart of another embodiment of a track maintenance method provided by an embodiment of the present application. The track maintenance method may include the following steps:
[0110] S301. Obtain on-site data of at least one track component of the track to be maintained. The on-site data includes relevant data of each track component, and the relevant data includes at least one of the following: geometric parameters, component defects, service life, or historical maintenance data.
[0111] It can be understood that in this embodiment, various data, such as on-site data and status data, can be represented by binary data or decimal data. In this embodiment, the data type is not overly limited. Various data can be directly obtained through binary conversion or through encoding. For example, one-hot encoding or label encoding can be used to encode the relevant data.
[0112] Optionally, in terms of track geometric parameters, the status is a real number representing irregular size. In terms of track component defects, the status is a binary number representing whether there are defects. To update the status of the reinforcement learning model, it is necessary to use the on-site data in track geometric measurements, defect inspection reports, and maintenance records to consider the change of each status. The change of the status is based on these data and uses a normal distribution related to the specific maintenance activities performed above.
[0113] For example, geometric parameters, component defects, and historical maintenance data can all be represented in binary. Taking several defects specifically included in track geometric parameter defects: superelevation anomaly, longitudinal level anomaly, plane anomaly, gauge anomaly, and twist anomaly as examples, the above defects can be encoded using methods such as one-hot encoding or label encoding.
[0114] S302. Input the on-site data of each track component into the digital twin model respectively, and extract the status data of each track component through the digital twin model. The status data includes the component category and defect data of the corresponding track component.
[0115] S303. Determine at least one maintenance strategy according to at least one maintenance action. The maintenance strategy includes one or more maintenance actions selected from at least one maintenance action. The maintenance action is any one of the following: tamping, rail grinding, ballast cleaning, sleeper replacement, rail replacement, fastener replacement, or ballast unloading.
[0116] Optionally, when the maintenance action consists of 7 actions such as tamping, rail grinding, ballast cleaning, sleeper replacement, rail replacement, fastener replacement, or ballast unloading, at least one maintenance strategy can be obtained by combining the above 7 actions. The calculation formula for the final quantity is:
[0117]
[0118] Among them, n is the number of options, and r is the size of the combination. It can be seen from the formula that n is equal to 7, and r can vary from 0 to 7. The sum of combinations or possible actions is 128. From the maintenance records, there are more than 1 million sets of data, and it can be considered that these data are sufficient to comprehensively compare maintenance activities, changes in track geometry parameters, and the occurrence of track component defects in the form of normal distribution.
[0119] S304. Input the status data of each track component and at least one maintenance strategy into the trained reinforcement learning model, and select a target maintenance strategy for each track component from at least one maintenance strategy through the reinforcement learning model.
[0120] Specifically, the reinforcement learning model can first select at least one candidate maintenance strategy for the track component from at least one maintenance strategy, and then, according to the service life of the track component and the reward function, conduct decision evaluation for each candidate maintenance decision to obtain evaluation scores corresponding to multiple candidate maintenance strategies, and determine the top N candidate maintenance strategies with the highest evaluation scores as the N target maintenance strategies of the track component, where N is a positive integer greater than or equal to 1.
[0121] S305. Output the target maintenance strategies of each track component.
[0122] In the embodiments of the present application, through a series of specific maintenance actions and combining various maintenance strategies accordingly, this exhaustive strategy generation method ensures coverage of all feasible maintenance means and provides a comprehensive plan reserve for track repair. Thus, with the help of the trained reinforcement learning model and combined with the real-time status data of the track components, the target maintenance strategy is selected from numerous maintenance strategies. The reinforcement learning model "learns" experience based on a large amount of past track maintenance data, weighs the advantages and disadvantages of different strategies in terms of cost, efficiency, and effect, and makes the best recommendation that fits the actual situation, overcoming the subjectivity and limitations of manual decision-making.
[0123] As mentioned above, the purpose of the reinforcement learning model is to output the optimal maintenance decision for each track component. Based on the above objective, the reinforcement learning model can be trained.
[0124] Figure 4 The flowchart of a training method for a reinforcement learning model provided by the embodiments of the present application is shown, and the method may include the following steps:
[0125] S401. Obtain a plurality of training data, where the training data includes the status data of the track components participating in the training.
[0126] Optionally, the status data of the track components available for training, which is the training status data, can be obtained from the data storage device of the digital twin model.
[0127] S402. Take at least one maintenance strategy as the state of the reinforcement learning model respectively, input each piece of training data into the reinforcement learning model, and select the first maintenance strategy for the track components corresponding to each piece of training data through the reinforcement learning model.
[0128] The first maintenance strategy may refer to the maintenance strategy with the highest maintenance benefit decided by the agent of the reinforcement learning model for the track components during the current iterative calculation process.
[0129] S403. Take the highest total maintenance benefit obtained by the first maintenance strategy of the track components corresponding to each piece of training data as the iterative target, and iteratively train the reinforcement learning model until the reinforcement learning model meets the iterative target.
[0130] As Figure 5 shown, it is an example diagram of the working process of the reinforcement learning model.
[0131] For the reinforcement learning model, its input data may include: geometric parameters, component defects, historical maintenance data, and thresholds. The threshold may refer to the threshold for judging whether to perform maintenance on the track components. Then, input the data (such as geometric parameters, component defects, historical maintenance data) into the RL model. Determine the current state data of the track components. And judge whether the track components need to perform maintenance according to the threshold. If so, the maintenance cost needs to be increased. If not, the maintenance cost does not need to be increased. Then, the action response in the next stage can be determined to obtain the optimal maintenance decision for each track component. Thus, calculate the total maintenance benefit. The calculation process of the total maintenance benefit includes the replacement cost or repair cost corresponding to the remaining service life of the track components. By judging whether the total maintenance benefit reaches the benefit target. If it is greater than the threshold, the benefit target is reached and the iteration is terminated. If it is less than or equal to the threshold, the benefit target is not reached, and return to the step of inputting the data (such as geometric parameters, component defects, historical maintenance data) into the RL model to continue iterative training.
[0132] Optionally, it can Figure 5 shown that the working process can be linked with the digital twin model. The digital twin model is integrated with the reinforcement learning model, thereby promoting the overall efficiency of the project life cycle, rather than only considering all risks and vulnerabilities at a specific stage of the project.
[0133] Optionally, Figure 5 in the shown example, judge whether the track components need to perform maintenance according to the threshold, and the threshold may include one or more. When setting one threshold, it can be judged directly. When setting multiple thresholds, associate a maintenance priority with each threshold. The higher the maintenance priority, the higher the need for maintenance, and the more necessary it is to maintain the track components. The lower the maintenance priority, the lower the need for maintenance, and the lower the need to maintain the track components.
[0134] Taking the setting of four priorities as an example, when Priority 1 is greater than Priority 2, Priority 2 is greater than Priority 3, and Priority 3 is greater than Priority 4, Priority 1 means that the track geometric parameters are very poor and the track section needs to be maintained as soon as possible, while Priority 4 means that the track section needs to be incorporated into the regular maintenance plan.
[0135] Suppose the reinforcement learning model can estimate the maintenance values of each track component. The larger the maintenance value, the higher the need for maintenance, and the smaller the maintenance value, the lower the need for maintenance.
[0136] Thus, the corresponding maintenance priorities can be associated according to the threshold size. Suppose the threshold size is proportional to the priority level. When setting 3 thresholds, the first threshold, the second threshold, and the third threshold can form 4 threshold value ranges. Each threshold value range can be associated with a maintenance priority. The larger the threshold, the higher the level of the maintenance priority required.
[0137] For example, suppose there are four thresholds: 10, 20, 30, and 40. Among them, 10 is associated with Priority 4, 20 is associated with Priority 3, 30 is associated with Priority 2, and 40 is associated with Priority 1.
[0138] Optionally, Priority 4 is a trigger level with relatively low attention, and the threshold corresponding to Priority 4 can also be used as the overall threshold for maintenance triggering. That is, as long as the maintenance value of the track component is greater than or equal to the threshold corresponding to Priority 4, it is determined that maintenance is required. Taking the above four thresholds as an example, if the maintenance value of the track component is greater than 10, then maintenance is required.
[0139] The total maintenance benefit can refer to the sum of the maintenance benefits of each track component. The maintenance benefits of each track component can be obtained by calculating the reward function.
[0140] In the embodiments of the present application, the training status data of the track components is obtained. Such data covers component categories and defect data, providing highly targeted materials for subsequent training. Each training data is input into the reinforcement learning model for training the reinforcement learning model. During the training of the reinforcement learning model, with the goal of maximizing the total maintenance benefit, the reinforcement learning model is driven to continuously improve itself. Iterative training based on this goal enables the model to continuously weigh the advantages and disadvantages of different strategies, discard those solutions that are effective in the short term but have high long-term costs or poor effects, until the decision-making mode with the optimal overall benefit is found, so that the strategy suggestions given by the finally trained model in the actual track maintenance scenario are more economically beneficial and practical.
[0141] As an example, the total maintenance benefit of the reinforcement learning model refers to the sum of the maintenance benefits of each track component; the maintenance benefit of a track component refers to the sum of the maintenance cost of the track component and the maintenance benefit that can be obtained by implementing the first maintenance strategy when a defect occurs in the track component; the maintenance cost includes the replacement cost or repair cost corresponding to the remaining service life of the track component.
[0142] It can be understood that the replacement cost or repair cost corresponding to the remaining service life of the track component can be determined according to the service life of the track component. Specifically, the remaining service life of the track component can be determined according to the service life of the track component, and the replacement cost or repair cost of the track component can be determined according to the remaining service life of each track component in combination with the first maintenance strategy of the track component.
[0143] For example, a mapping relationship of the remaining service life, maintenance strategy, replacement cost or repair cost can be set. After obtaining the remaining service life and the first maintenance strategy of the track component, query the mapping relationship to obtain the replacement cost or repair cost of the track component.
[0144] Of course, a calculation formula can also be set. The input parameters of the calculation formula are the remaining service life and the maintenance strategy, and the output parameter is the replacement cost or repair cost of the track component. The calculation formula can be set according to experience or obtained by formula solving, and this embodiment does not limit this too much.
[0145] In the embodiment of the present application, the total maintenance benefit is defined as the sum of the maintenance benefits of each track component. This cumulative calculation method can control the maintenance benefit of the entire track system from a macro level. In addition, by associating with the remaining service life of the track component to consider the replacement and repair costs, the effective acquisition of costs can be realized. Thus, precise decisions can be made according to the component status in different stages. For example, for components approaching the end of their service life, replacement is preferred rather than repair, while for newly laid components, a low-cost repair strategy is preferred. Thus, in the track maintenance scenario, by continuously comparing the benefit situations under different strategies, the most beneficial and highest-yielding maintenance plan for the track components can be screened out, ensuring the safety of track operation while improving economic efficiency.
[0146] Further, the maintenance benefit of the track component is obtained by calculating a reward function. The reward function is as follows: reward = w1 * C + w2 * D, where reward is the reward result of the reward function, C is the maintenance cost, D is the maintenance benefit, w1 is the weight of the maintenance cost, and w2 is the weight of the maintenance benefit.
[0147] The reward function comprehensively considers the maintenance cost and the occurrence of defects, and can have a higher decision-making reference value, thereby improving the decision-making accuracy.
[0148] It can be understood that the maintenance revenue or cost in this application can be represented by a numerical value. Quantifying the cost numerically can more intuitively display the level of maintenance benefits.
[0149] Among them, the maintenance cost refers to the cost required for maintaining the track cross-section. It is related to the type of maintenance decision and the frequency of maintenance activities, and the value of the maintenance cost is estimated based on real data. The defect occurrence situation refers to the probability of defects occurring in the track cross-section within a certain period of time. It is related to the service life of the track cross-section and the type of maintenance decision, and the value of the defect occurrence situation is simulated based on the digital twin model. The goal of this reward function is to reduce the maintenance cost while ensuring railway safety. Therefore, the weight w1 of the maintenance cost is negative, and the weight w2 of the defect occurrence situation is positive. The values of the two weights are determined based on experimental results.
[0150] Compared with the maintenance cost, the penalty when a defect occurs is set relatively high, that is, the absolute value of w2 can be set to be greater than the absolute value of w1, because the purpose of agent training is to minimize costs and reduce the number of defects.
[0151] In addition, loss and policy entropy will also be used to display the performance of the reinforcement learning model. The loss represents the defects that occur in the track section, while the policy entropy shows the degree of response of the agent to the problem.
[0152] In the embodiments of this application, by setting a clear reward function "reward = w1*C + w2*D", the relatively vague and difficult-to-intuitively-compare maintenance cost and maintenance revenue are transformed into a specific quantitative value. This value accurately reflects the comprehensive benefits brought by each maintenance behavior of the track components, making the overall benefits of each iteration have a clear and measurable standard, and making the effectiveness evaluation of track maintenance work more scientific and accurate.
[0153] As an embodiment, after outputting the target maintenance strategies of each track component, it further includes:
[0154] After each track component performs track maintenance according to the corresponding target maintenance strategy, collect the track maintenance results of each track component;
[0155] Update the on-site data of each track component according to the maintenance results of each track component.
[0156] Optionally, updating the on-site data of each track component according to the maintenance results of each track component may include: updating the historical maintenance data in the on-site data of each track component according to the maintenance results of each track component.
[0157] In the embodiments of the present application, after outputting the target maintenance strategy and performing the maintenance operation, the actual track maintenance results are collected. This link transforms the entire track maintenance process from a simple "decision-making - execution" into a closed-loop system. Updating the on-site data of each track component based on the track maintenance results enables the system to understand the true effectiveness of the strategy, and then make targeted adjustments and optimizations. This allows the model to continuously absorb the latest practical feedback, continuously self-optimize, make the track maintenance strategy more in line with the actual working conditions, improve the reliability and effectiveness of the entire track maintenance process, and reduce unnecessary maintenance costs and potential risks.
[0158] Figure 6 FIG. 4 is a schematic structural diagram of a track maintenance device provided by an embodiment of the present application. The track maintenance device 600 may include:
[0159] A data acquisition unit 601, configured to acquire on-site data of at least one track component of a track to be maintained. The on-site data includes relevant data of each track component, and the relevant data includes at least one of the following: geometric parameters, component defects, service life, or historical maintenance data.
[0160] A state extraction unit 602, configured to input the on-site data of each track component into a digital twin model respectively, and extract the state data of each track component through the digital twin model. The state data includes the component category and defect data of the corresponding track component.
[0161] A strategy acquisition unit 603, configured to input the state data of each track component into a reinforcement learning model respectively. The reinforcement learning model makes a decision on the target maintenance strategy for each track component based on the service life and reward function of each track component, and the reward function is determined based on the maintenance cost and the occurrence of defects.
[0162] A strategy prompt unit 604, configured to output the target maintenance strategy of each track component.
[0163] As an embodiment, the geometric parameters include at least one of the following: measurement position, superelevation, first longitudinal level, second longitudinal level, first plane, second plane, gauge, and twist; the superelevation refers to the height difference between the inner and outer rail surfaces of the track in the curve section; the first longitudinal level refers to the longitudinal level of a 10-meter chord, specifically, within the range of a 10-meter chord length, the longitudinal unevenness of the track; the second longitudinal level refers to the longitudinal level of a 20-meter chord, specifically, within the range of a 20-meter chord length, the longitudinal unevenness of the track; the first plane refers to the plane of a 10-meter chord, specifically, within the range of a 10-meter chord length, the lateral position unevenness of the track; the second plane refers to the plane of a 20-meter chord, specifically, within the range of a 20-meter chord length, the lateral position unevenness of the track;
[0164] Component defects include at least one of the following: multiple track defect information, where track defect information refers to the defect information of any track component, as well as the date and location where the defect information exists. The defect information of the component includes the component type and / or defect type. The component type includes at least one of the following: roadbed, fastener, rail, sleeper, turnout, and crossing. The defect type includes at least one of the following: track geometry parameter defect, track component defect. The track geometry parameter defect includes at least one of the following: superelevation anomaly, longitudinal level anomaly, planar anomaly, gauge anomaly, twist anomaly. The track component defect includes at least one of the following: roadbed anomaly, fastener anomaly, rail anomaly, sleeper anomaly, turnout and crossing anomaly;
[0165] Historical maintenance data includes at least one of the following: the maintained track components, the specific maintenance strategy, and the maintenance time. The specific maintenance strategy includes one or more maintenance actions.
[0166] As another embodiment, the policy acquisition unit may include:
[0167] A policy enumeration module, configured to determine at least one maintenance strategy according to at least one maintenance action. The maintenance strategy includes one or more maintenance actions selected from at least one maintenance action. The maintenance action is any one of the following: tamping, rail grinding, ballast cleaning, sleeper replacement, rail replacement, fastener replacement, or ballast unloading;
[0168] A policy analysis module, configured to input the status data of each track component and at least one maintenance strategy into the trained reinforcement learning model, and select a target maintenance strategy for each track component from at least one maintenance strategy through the reinforcement learning model.
[0169] As another embodiment, the reinforcement learning model is trained with the highest total maintenance benefit of each track component as the training objective; the device further includes:
[0170] A sample acquisition unit, configured to acquire multiple training data. The training data includes the status data of the track components participating in the training. The training status data includes the component category and defect data of the track components corresponding to the training.
[0171] A model training unit, configured to use at least one maintenance strategy as the state of the reinforcement learning model respectively, input each training data into the reinforcement learning model, and select a first maintenance strategy for the track component corresponding to each training data through the reinforcement learning model;
[0172] An iteration judgment unit, configured to take the highest total maintenance benefit obtained by the first maintenance strategy of the track component corresponding to each training data as the iteration objective, and iteratively train the reinforcement learning model until the reinforcement learning model meets the iteration objective.
[0173] As another embodiment, the total maintenance benefit of the reinforcement learning model refers to the sum of the maintenance benefits of each track component;
[0174] The maintenance benefit of a track component refers to the sum of the maintenance cost of the track component and the maintenance benefit that can be obtained by executing the first maintenance strategy when a defect occurs in the track component;
[0175] The maintenance cost includes the replacement cost or repair cost corresponding to the remaining service life of the track component.
[0176] As another embodiment, the maintenance benefit of the track component is obtained by calculating with a reward function. The reward function is as follows: reward = w1*C + w2*D, where reward is the reward result of the reward function, C is the maintenance cost, D is the maintenance benefit, w1 is the weight of the maintenance cost, and w2 is the weight of the maintenance benefit.
[0177] As another embodiment, it further includes:
[0178] After each track component performs track maintenance according to the corresponding target maintenance strategy, collect the track maintenance results of each track component;
[0179] According to the maintenance results of each track component, update the on-site data of the track to be maintained.
[0180] As Figure 7 shown, it is a schematic block diagram of a computer device provided by an embodiment of the present application. The computer device may include: a memory 701, a processor 702, and a computer program stored on the memory 701. The processor executes the computer program to implement the steps of any of the above methods.
[0181] In some embodiments of this embodiment, a computer program product is provided, including a computer program / instructions. When the computer program is executed by a processor, it implements the steps of the method in the above embodiment.
[0182] The processor may include, but is not limited to, for example, one or more processors or microprocessors, etc. Each processor may be implemented by an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a controller, a microcontroller, a microprocessor, or other electronic components, and is used to execute the methods in the above embodiments.
[0183] The computer-readable storage medium may be implemented by any type of volatile or non-volatile storage device or a combination thereof. The computer-readable storage medium may include, but is not limited to, for example, random access memory (RAM), read-only memory (ROM), flash memory, EPROM memory, EEPROM memory, registers, computer storage media (such as hard disks, floppy disks, solid state drives, removable disks, CD-ROMs, DVD-ROMs, Blu-ray discs, etc.).
[0184] The computer-readable storage medium may also store at least one computer-executable program / instructions, which are, for example, computer-readable instructions. The computer-readable storage medium includes, but is not limited to, for example, volatile memory and / or non-volatile memory. Volatile memory may, for example, include random access memory (RAM) and / or cache memory, etc. The computer-readable storage medium may, for example, include read-only memory (ROM), hard disks, flash memory, etc. For example, the non-transitory computer-readable storage medium may be connected to a computing device such as a computer. Then, when the computing device runs the computer-readable instructions stored on the computer-readable storage medium, the various methods described above may be performed.
[0185] In addition, the computer device may also include (but is not limited to) a data bus, an input / output (I / O) bus, a display, and input / output devices (such as a keyboard, a mouse, a speaker, etc.).
[0186] The processor may communicate with external devices via the I / O bus through a wired or wireless network.
[0187] In one embodiment, the at least one computer-executable instruction may also be compiled into or form a software product / computer program product, and when one or more computer-executable instructions are run by a processor, each function and / or method step in the embodiments described in this technology is executed.
[0188] In the embodiments provided in this disclosure, it should be understood that the disclosed devices and methods may also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of this disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code includes one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or actions, or may be implemented by a combination of dedicated hardware and computer instructions.
[0189] It should be noted that in this disclosure, the terms "include", "comprise", or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article, or device. Without further limitation, an element limited by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device including the element.
[0190] Although the disclosed embodiments are as above, the above content is only an embodiment adopted for the convenience of understanding this disclosure and is not intended to limit this disclosure. Any person skilled in the art within the technical field to which this disclosure pertains may make any modifications and changes in the form of implementation and details without departing from the spirit and scope disclosed by this disclosure. However, the scope of patent protection of this disclosure shall still be subject to the scope defined by the appended claims.
Claims
1. An orbit maintenance method, characterized in that, Including: Obtaining on-site data of at least one track component of the track to be maintained, where the on-site data includes relevant data of each track component, and the relevant data includes at least one of the following: geometric parameters, component defects, service life, or historical maintenance data; Respectively inputting the on-site data of each track component into the digital twin model, and extracting the state data of each track component through the digital twin model, where the state data includes the component category and defect data of the corresponding track component; Respectively inputting the state data of each track component into the reinforcement learning model, and the reinforcement learning model determines the target maintenance strategy for each track component based on the service life and reward function of each track component, and the reward function is determined based on the maintenance cost and the occurrence of defects; Outputting the target maintenance strategy of each track component.
2. The method according to claim 1, wherein The geometric parameters include at least one of the following: measurement position, superelevation, first longitudinal level, second longitudinal level, first plane, second plane, gauge, and twist; the superelevation refers to the height difference between the inner and outer rail surfaces of the track in the curve section, and the first longitudinal level refers to the longitudinal level of a 10-meter chord, specifically referring to the longitudinal unevenness of the track within the range of a 10-meter chord length; the second longitudinal level refers to the longitudinal level of a 20-meter chord, specifically referring to the longitudinal unevenness of the track within the range of a 20-meter chord length; the first plane refers to the plane of a 10-meter chord, specifically referring to the lateral position unevenness of the track within the range of a 10-meter chord length; the second plane refers to the plane of a 20-meter chord, specifically referring to the lateral position unevenness of the track within the range of a 20-meter chord length; The component defects include at least one of the following: multiple track defect information, where the track defect information refers to the defect information existing in any track component and the date and location where the defect information exists, and the defect information existing in the component includes the component type and / or defect type, and the component type includes at least one of the following: ballast bed, fastener, rail, sleeper, turnout, and crossing; the defect type includes at least one of the following: track geometric parameter defect, track component defect, and the track geometric parameter defect includes at least one of the following: abnormal superelevation, abnormal longitudinal level, abnormal plane, abnormal gauge, abnormal twist; the track component defect includes at least one of the following: abnormal ballast bed, abnormal fastener, abnormal rail, abnormal sleeper, abnormal turnout, and abnormal crossing; The historical maintenance data includes at least one of the following: the track component maintained, the specific maintenance strategy, and the maintenance time, and the specific maintenance strategy includes one or more maintenance actions.
3. The method according to claim 1, wherein The step of respectively inputting the state data of each track component into the reinforcement learning model includes: Determining at least one maintenance strategy according to at least one maintenance action, where the maintenance strategy includes one or more maintenance actions selected from the at least one maintenance action, and the maintenance action is any one of the following: tamping, rail grinding, ballast cleaning, sleeper replacement, rail replacement, fastener replacement, or ballast unloading; Input the status data of each track component and at least one maintenance strategy into the trained reinforcement learning model, and select a target maintenance strategy for each track component from the at least one maintenance strategy through the reinforcement learning model.
4. The method according to any one of claims 1 to 3, characterized in that, The reinforcement learning model is trained with the highest total maintenance benefit of each track component as the training objective; The training steps of the reinforcement learning include: Obtain a plurality of training data, where the training data includes the status data of the track components participating in the training; Use at least one maintenance strategy as the state of the reinforcement learning model respectively, input each training data into the reinforcement learning model, and select a first maintenance strategy for the track component corresponding to each training data through the reinforcement learning model; Take the highest total maintenance benefit obtained by the first maintenance strategy of the track component corresponding to each training data as the iteration objective, and iteratively train the reinforcement learning model until the reinforcement learning model meets the iteration objective.
5. The method according to claim 4, wherein The total maintenance benefit of the reinforcement learning model refers to the sum of the maintenance benefits of each track component; The maintenance benefit of the track component refers to the sum of the maintenance cost of the track component and the maintenance benefit that can be obtained by executing the first maintenance strategy when a defect occurs in the track component; The maintenance cost includes the replacement cost or repair cost corresponding to the remaining service life of the track component.
6. The method according to claim 5, wherein The maintenance benefit of the track component is calculated through a reward function, and the reward function is as follows: reward = w1*C + w2*D, where reward is the reward result of the reward function, C is the maintenance cost, D is the maintenance benefit, w1 is the weight of the maintenance cost, and w2 is the weight of the maintenance benefit.
7. The method according to claim 1, characterized in that, It further includes: After each track component performs track maintenance according to the corresponding target maintenance strategy, collect the track maintenance results of each track component; Update the on-site data of the track to be maintained according to the maintenance results of each track component.
8. An orbital maintenance device, characterized in that, It includes: A data acquisition unit for acquiring on-site data of at least one track component of the track to be maintained, where the on-site data includes relevant data of each track component, and the relevant data includes at least one of the following: geometric parameters, component defects, service life, or historical maintenance data; A status extraction unit for inputting the on-site data of each track component into the digital twin model respectively, and extracting the status data of each track component through the digital twin model, where the status data includes the component category and defect data of the corresponding track component; A strategy acquisition unit for inputting the status data of each track component into the reinforcement learning model respectively, and the reinforcement learning model makes a decision on the target maintenance strategy for each track component based on the service life of each track component and the reward function, and the reward function is determined based on the maintenance cost and the occurrence of defects; A strategy prompt unit for outputting the target maintenance strategy of each track component.
9. A computer device, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 7.