A Method and System for Dynamic Monitoring of Energy Efficiency Balance in Rail Transit

By adopting fuzzy clustering classification model, life cycle cost analysis and Bayesian network technologies in the rail transit system, the misjudgment and misjudgment problems in rail transit energy efficiency monitoring are solved, and the accuracy of the root cause analysis of energy loss and the pertinence of the solution are improved.

CN119648206BActive Publication Date: 2025-05-30NINGBO YIKATONG TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510173823.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-05-30
Estimated Expiration
2045-02-18

AI Technical Summary

Technical Problem

The existing technology has problems of misjudgment and misjudgment in the monitoring of rail transit energy efficiency, and the expert system is highly subjective, making it difficult to accurately locate the source of energy loss, and the solution is not targeted.

Method used

A dynamic monitoring method for energy efficiency balance of rail transit is adopted. By obtaining relevant data from external devices, inputting them into the fuzzy clustering classification model, outputting device classification results, and using life cycle cost analysis methods to determine the personalized energy efficiency standard threshold. Combining Bayesian network and fault tree analysis, the root causes of energy loss are located.

Benefits of technology

It effectively improves the level of rail transit energy efficiency management, reduces misjudgment and misjudgment, improves the accuracy of the root cause analysis of energy loss, and provides highly targeted solutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119648206B_ABST
    Figure CN119648206B_ABST
Patent Text Reader

Abstract

This application relates to a method and system for dynamic monitoring of energy efficiency balance in rail transit, belonging to the field of rail transit energy conservation and monitoring. It solves the problem that in the energy consumption monitoring link, simple statistical analysis and rule-based systems are difficult to adapt to complex and changeable operating environments due to fixed thresholds and rules, and are prone to misjudgment and missed judgment. The method includes: determining whether there is an abnormal energy consumption in the corresponding equipment according to the comparison result between the actual energy consumption of the equipment and the personalized energy efficiency standard threshold; if so, combining the personalized threshold of the equipment and the historical energy consumption data, using the constructed comprehensive energy loss analysis model based on fault tree and Bayesian network, starting from the fault mode of the equipment, combining the operation history, maintenance records, and environmental factors of the equipment, through Bayesian network reasoning and dynamic fault tree analysis, locating the root cause of energy loss, and sending it to the terminal held by the person in charge. This application has the following effects: effectively improving the level of rail transit energy efficiency management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of rail transit energy conservation and monitoring, and in particular to a method and system for dynamically monitoring the energy efficiency balance of rail transit. Background Art

[0002] At present, in the aspect of rail transit energy efficiency monitoring, a variety of technologies have been developed. Some technologies collect data in real time through sensors installed on equipment and use simple statistical analysis methods to determine whether the energy consumption is within the normal range. There are also some that use rule-based systems to identify abnormal energy consumption based on pre-set energy consumption standards and equipment operation parameters. In addition, in the field of fault diagnosis and energy loss analysis, there are technologies that use expert systems, relying on the experience and knowledge of domain experts to build a knowledge base to reason and judge the possible causes of energy loss. In formulating solutions, generally, based on the handling experience of past similar cases, general improvement suggestions are given.

[0003] However, the existing technologies have obvious deficiencies. In the energy consumption monitoring link, simple statistical analysis and rule-based systems are difficult to adapt to complex and changeable operating environments due to fixed thresholds and rules, and are prone to misjudgment and missed judgment. When analyzing the root causes of energy loss, expert systems rely on expert experience, are highly subjective, and lack sufficient multi-source data fusion, making it difficult to accurately locate problems. In terms of solution generation, the suggestions based on past cases lack pertinence and cannot meet the unique needs of projects, resulting in poor implementation effects. Summary of the Invention

[0004] In order to effectively improve the energy efficiency management level of rail transit, the present application provides a method and system for dynamically monitoring the energy efficiency balance of rail transit.

[0005] In a first aspect, the present application provides a method for dynamically monitoring the energy efficiency balance of rail transit, adopting the following technical solutions:

[0006] A method for dynamically monitoring the energy efficiency balance of rail transit includes:

[0007] Obtain relevant data of various external devices;

[0008] Input the relevant data of various external devices into a pre-constructed fuzzy clustering classification model, and output the classification results of various external devices;

[0009] According to the device classification results, use the life cycle cost analysis method to determine the personalized energy efficiency standard thresholds for various devices;

[0010] According to the comparison result between the actual energy consumption of the device and the personalized energy efficiency standard threshold, determine whether there is abnormal energy consumption in the corresponding device;

[0011] If it is yes, then combining the personalized threshold of the device and the historical energy consumption data, using the constructed comprehensive energy loss analysis model based on fault tree and Bayesian network, starting from the fault mode of the device, combining the operation history, maintenance records, and environmental factors of the device, through Bayesian network reasoning and dynamic fault tree analysis, locate the root cause of energy loss and send it to the terminal held by the person in charge;

[0012] If it is no, then no notification is made.

[0013] Optionally, using the life cycle cost analysis method to determine the personalized energy efficiency standard threshold of various devices includes:

[0014] According to the real-time relevant data of various external devices, analyze whether there is a change in the device state;

[0015] If it is no, then maintain the original settings, and the original settings are to use the life cycle cost analysis method to determine the personalized energy efficiency standard threshold of various devices;

[0016] If it is yes, then use the device with the changed state as the target device, and collect various information of the target device. Among them, the various information of the target device includes the real-time operation state, energy consumption data, and usage frequency;

[0017] Use the life cycle cost analysis method to determine the personalized energy efficiency standard threshold of the target device as the initial personalized threshold, and input the initial personalized threshold of the target device and the various information of the target device into the pre-set and trained adjustment value analysis model based on the reinforcement learning algorithm to output the adjustment value;

[0018] Based on the adjustment value and the initial personalized threshold, calculate and determine the personalized energy efficiency standard threshold of the target device.

[0019] Optionally, according to the real-time relevant data of various external devices, analyzing whether there is a change in the device state includes:

[0020] Obtain the scene category where the external device is located. Among them, the scenes include daily operation scenes, peak operation scenes, equipment maintenance scenes, extreme weather scenes, and equipment upgrade and transformation scenes;

[0021] According to the corresponding relationship between the scene category where the external device is located and the device state analysis algorithm, match the device state analysis algorithms used by different external devices;

[0022] Based on the real-time relevant data of the external device, use the matched device state analysis algorithm to analyze whether there is a change in the device state of the external device.

[0023] In the second aspect, the present application provides a rail transit energy efficiency balance dynamic monitoring system, adopting the following technical solutions:

[0024] A dynamic monitoring system for rail transit energy efficiency balance includes a memory, a processor, and a program stored on the memory and executable on the processor. When the program is loaded and executed by the processor, it can implement the dynamic monitoring method for rail transit energy efficiency balance described in the first aspect. Description of the Drawings

[0025] Figure 1 It is a schematic diagram of the overall process of a dynamic monitoring method for rail transit energy efficiency balance in an embodiment of the present application.

[0026] Figure 2 It is a schematic diagram of the process of using the life cycle cost analysis method to determine the personalized energy efficiency standard thresholds of various types of equipment in another embodiment of the present application.

[0027] Figure 3 It is a schematic diagram of the process of analyzing whether there is a change in the equipment status based on the real-time relevant data of various external devices in another embodiment of the present application. Detailed Embodiments

[0028] The present application will be further described in detail below with reference to the accompanying drawings.

[0029] Referring to Figure 1 , a dynamic monitoring method for rail transit energy efficiency balance disclosed in the present application includes:

[0030] Step S100, obtaining relevant data of various external devices.

[0031] External devices: In the rail transit system, various facilities related to the dynamic monitoring of energy efficiency balance. For example, the traction motor of a train, which is the power core of train operation and has a high energy consumption ratio; the escalator in a station, which continuously operates and consumes electricity, and its energy consumption affects the overall energy efficiency of the station; and the signal system equipment, which has a large number and operates throughout the day, and the total energy consumption cannot be ignored.

[0032] Relevant data: It includes information that can reflect the operating status, energy consumption characteristics, and other factors affecting energy efficiency of external devices. It involves both real-time operating parameters of the equipment, such as current and voltage during operation, and historical data accumulated during the long-term operation of the equipment, such as monthly average energy consumption and the number of faults in the past year, and also includes equipment operating environment data, such as humidity and altitude.

[0033] There are several ways to obtain relevant data of various external devices, as follows: 1. Sensor collection: Install different sensors on various types of equipment to collect data. For example, pressure sensors are installed in the train braking system to monitor the braking pressure and determine the system status and energy consumption; flow sensors are used in the station ventilation duct to measure the ventilation volume and help optimize the energy saving of the ventilation system. 2. Equipment with built-in system interface: Many rail transit equipment now have intelligent control systems and data interfaces, which are connected to the monitoring system through communication protocols to obtain data. Small equipment such as intelligent lighting control modules may use Bluetooth interfaces, and monitoring personnel can obtain working parameters at a certain distance; large equipment such as train traction control systems often use industrial Ethernet interfaces to achieve high-speed and stable data transmission. 3. Data storage system reading: The data generated by rail transit operations is stored in a special system, and historical data can be read for analysis. Relational databases store structured data, such as basic equipment information and maintenance records, which can be quickly obtained and analyzed using SQL statements; distributed file systems store massive unstructured data, such as equipment operation logs, from which information such as start and stop times and abnormal alarms can be extracted.

[0034] Step S200: input relevant data of various external devices into a pre-built fuzzy clustering classification model, and output classification results of various external devices.

[0035] Fuzzy clustering classification model: Based on fuzzy mathematics, it uses "membership degree" (between 0 and 1) to measure the degree to which a data point belongs to a certain category, breaking the traditional clear classification boundaries, flexibly classifying rail transit external equipment data, and mining the potential characteristics of the data.

[0036] The information related to model construction and training is as follows: 1. Construction: Based on a large amount of historical data of rail transit external equipment, covering operating conditions, energy consumption, load and other information, the model is built by mining data patterns. 2. Training: Divide historical data into training sets and test sets, use the training set to tune the internal parameters of the model, such as cluster centers and membership functions, and then use the test set to evaluate. If the performance does not meet the requirements, adjust the parameters or increase the data for retraining until it meets the requirements.

[0037] The relevant contents of data input and processing are as follows: 1. Input: The real-time parameters, historical energy consumption, environmental parameters, etc. of the equipment obtained in step S100 are formatted and pre-processed before inputting into the model. 2. Processing: Standardize to eliminate dimensional differences, then calculate the degree of membership between data points and cluster centers using the fuzzy clustering algorithm, and iterate and adjust to make the similarity of the same type of data high and the similarity of different types low, and determine the equipment classification.

[0038] The relevant content of the output and application is as follows: 1. Output: Provide the equipment classification results and membership degrees, such as being classified into categories like "highly energy-efficient", "normal energy consumption", "high energy consumption that needs to be optimized", etc., and the corresponding values. 2. Application: Provide a basis for energy efficiency management and equipment optimization, summarize the advantages of "highly energy-efficient" equipment, analyze the problems of "high energy consumption that needs to be optimized" equipment and transform them to improve the system energy efficiency.

[0039] Step S300: According to the equipment classification results, use the life cycle cost analysis method to determine the personalized energy efficiency standard thresholds for various types of equipment.

[0040] Life cycle cost analysis method: Comprehensively consider the costs of the entire process of the equipment from birth to scrapping, such as the expenses in procurement, use, maintenance, scrapping treatment and other links, and analyze energy consumption and other costs based on this to balance the costs and energy efficiency of the entire life cycle of the equipment.

[0041] Personalized energy efficiency standard threshold: Based on the equipment characteristics, usage frequency, operating environment and life cycle cost analysis results, formulate exclusive energy consumption judgment criteria for different equipment to accurately manage the energy efficiency of the equipment.

[0042] The process of determining the personalized energy efficiency standard threshold is as follows: 1. Use the equipment classification results: Refer to the equipment classification in step S200, and classify the equipment with similar characteristics and usage patterns into one category, such as one category for train traction motors and one category for station ventilation equipment. 2. Conduct life cycle cost analysis: Sort out the costs at each stage of each type of equipment, calculate the purchase and transportation costs during procurement; count the energy consumption and maintenance costs during use; estimate the demolition and recycling costs during scrapping. 3. Determine the threshold: Combine the total cost of the entire life cycle of the equipment, industry energy-saving standards and technological trends, and set thresholds for equipment under different working conditions. For example, for equipment with frequent starts and stops, the starting energy consumption standard is appropriately relaxed, and strict control is implemented during stable operation.

[0043] Step S400: According to the comparison result between the actual energy consumption of the equipment and the personalized energy efficiency standard threshold, determine whether there is abnormal energy consumption in the corresponding equipment. If yes, execute step S500; if no, execute step S600.

[0044] Among them, the actual energy consumption: refers to the amount of energy consumed by the equipment during actual operation, and is obtained by real-time collection and recording through various metering devices installed on the equipment, such as electricity meters, gas meters, etc.

[0045] The comparison process is as follows: During the operation of rail transit equipment, continuously monitor and collect the actual energy consumption data of the equipment. According to a certain time period, such as every hour, every day, compare the collected actual energy consumption with the personalized energy efficiency standard threshold determined in step S300. During the comparison, focus on the magnitude relationship of the energy consumption values and the degree of deviation of the actual energy consumption from the threshold.

[0046] Step S500, combining the personalized threshold of the equipment and the historical energy consumption data, using the built comprehensive energy loss analysis model based on fault tree and Bayesian network, starting from the failure mode of the equipment, combined with the equipment's operation history, maintenance records, and environmental factors, through Bayesian network reasoning and dynamic fault tree analysis, locate the root cause of the energy loss and send it to the terminal held by the person in charge.

[0047] The analysis process is as follows: 1. Use historical energy consumption data to calculate the mean and standard deviation, determine the energy consumption boundary, analyze historical energy consumption according to multiple dimensions, and establish a normal energy consumption mode. 2. Sort out the failure modes of equipment from systems to components, determine the probability of occurrence of each subdivided mode, and build a dynamic fault tree with energy consumption exceeding the preset proportion of the normal boundary as the top event. Use logic gates to describe causal relationships, find all possible fault paths, and select high-frequency fault bottom events as the direction of troubleshooting. 3. Convert fault tree events into Bayesian network nodes and edges, determine the prior probability based on historical equipment failures and industry statistics, input multi-source data to update the posterior probability, and determine the cause of the failure through probability changes.

[0048] The positioning results and applications are as follows: Integrate the fault tree and Bayesian network analysis results, establish an evaluation matrix to score the fault mode, and select factors whose comprehensive scores exceed the preset scores and whose energy consumption increases exceed the preset ratio as the root causes of energy loss. Send the positioning results to the responsible terminal to provide a basis for the subsequent formulation of solutions.

[0049] Step S600: no notification.

[0050] Among them, no notification means that the equipment's operating energy consumption is within a normal and reasonable range, and there is no need to provide special reminders to relevant personnel or initiate additional analysis and processing procedures.

[0051] Reference Figure 2 , using the life cycle cost analysis method to determine the personalized energy efficiency standard thresholds for various types of equipment include:

[0052] Step S310, based on the real-time related data of various external devices, analyze whether there is a change in the device status. If not, execute step S320; if yes, execute step S330.

[0053] Among them, real-time relevant data of external equipment refers to the data generated by the current operation of various equipment in the rail transit system, such as operating voltage, current size, equipment operating speed, etc., which can intuitively reflect the current operating status of the equipment.

[0054] Equipment status changes: covers changes in equipment operating modes, such as switching from normal operation to energy-saving mode; load changes, such as different loads of ventilation equipment during peak and off-peak periods; also includes equipment failures, starts and stops, etc.

[0055] The general process is as follows: 1. The monitoring system continuously receives real-time data transmitted from sensors and the device's own system. 2. Using preset data analysis algorithms, the real-time data is compared with the standard data set for the normal operation of the device. 3. Based on the comparison results, it is judged whether the device status has changed.

[0056] For example: Taking the subway ventilation equipment as an example, the current data of the ventilator is collected in real time through a current sensor. When operating normally, the current is stable at 5 - 8A. If at a certain moment, it is monitored that the current suddenly rises to 12A, far exceeding the normal range, through algorithm analysis, it can be determined that the status of the ventilation equipment has changed, which may be caused by a blocked air duct resulting in an increased load.

[0057] Step S320, maintain the original settings, and the original settings are to determine the personalized energy efficiency standard thresholds for various devices using the life cycle cost analysis method.

[0058] Step S330, take the device whose status has changed as the target device, and collect various information of the target device.

[0059] Among them, the various information of the target device includes the real-time operating status, energy consumption data, and usage frequency.

[0060] The general process is as follows: 1. Once the target device is determined, the system immediately starts the information collection program for this device. 2. Various sensors transmit the real-time monitored data to the data acquisition terminal, and the device management system also uploads relevant information to the monitoring platform synchronously. 3. The monitoring platform preliminarily sorts and classifies the collected data, removes abnormal data, and ensures the accuracy and effectiveness of the data to prepare for subsequent analysis.

[0061] Step S340, use the life cycle cost analysis method to determine the personalized energy efficiency standard threshold of the target device as the initial personalized threshold, and input the initial personalized threshold of the target device and the various information of the target device into the pre-set and trained adjustment value analysis model based on the reinforcement learning algorithm to output an adjustment value.

[0062] Among them, the adjustment value analysis model based on the reinforcement learning algorithm: A trained intelligent model that can continuously optimize and learn according to the input information of the target device using the reinforcement learning algorithm, and output an adjustment value to correct the initial personalized threshold to make it more suitable for the actual operation of the device.

[0063] The general process is as follows: 1. Use the life cycle cost analysis method to process the life cycle cost data of the target device to obtain the initial personalized threshold. 2. Integrate the initial personalized threshold with the various information of the target device collected in step S330 and input it into the trained adjustment value analysis model. 3. The model analyzes and calculates using the reinforcement learning algorithm based on the input data and outputs an adjustment value.

[0064] Taking the escalator in the subway as an example, first calculate the initial personalized threshold by analyzing its procurement cost, electricity cost over the years, regular maintenance cost, etc. Then, input information such as the current running speed (real-time running status), daily power consumption (energy consumption data), and the number of starts and stops within a day (usage frequency) of the escalator, together with the initial personalized threshold, into the model. Finally, the model outputs an adjustment value for subsequent correction of the personalized energy efficiency standard threshold of the escalator.

[0065] Step S350, based on the adjustment value and the initial personalized threshold, calculate and determine the personalized energy efficiency standard threshold of the target device.

[0066] The general process is as follows: 1. Obtain the initial personalized threshold and the adjustment value from the corresponding steps and the model. 2. According to the preset calculation rule, perform an operation on the adjustment value and the initial personalized threshold. Usually, add or subtract the adjustment value based on the initial personalized threshold to obtain the final personalized energy efficiency standard threshold. 3. Store the determined personalized energy efficiency standard threshold in the system database for subsequent comparison when monitoring the device energy consumption.

[0067] Refer to Figure 3 , according to the real-time relevant data of various external devices, analyze whether there are changes in the device status, including:

[0068] Step S311, obtain the scene category where the external device is located.

[0069] Among them, the scene category where the external device is located refers to different environments or working stages when the external device of rail transit is running. The scenes include daily operation scenes, peak operation scenes, equipment maintenance scenes, extreme weather scenes, and equipment upgrade and transformation scenes.

[0070] The daily operation scene is the normal running state of the device; the peak operation scene is the period when the passenger flow is large and the device is running at high load; the equipment maintenance scene is the period when the equipment is under inspection and maintenance; the extreme weather scene is the running environment of the equipment under bad weather such as heavy rain and heavy snow; the equipment upgrade and transformation scene is the stage when the equipment is undergoing function upgrade or technical transformation.

[0071] The acquisition method is disclosed as follows:

[0072] Time and passenger flow data: Obtain daily time information and passenger flow data of each station through the operation management system of rail transit. For example, during the morning and evening rush hours every day, the passenger flow significantly increases, and combined with the time, it can be determined that it is in the peak operation scenario. Equipment maintenance plan and records: Obtain the equipment maintenance plan arrangement and actual maintenance records from the equipment maintenance management platform to clarify when the equipment is in the equipment maintenance scenario. Meteorological information docking: Docking with the data interface of the meteorological department to obtain real-time meteorological data to determine whether it is in an extreme weather scenario, such as heavy rain or heavy snow weather. Engineering construction and renovation documents: Consult the engineering construction and equipment renovation related documents of rail transit to understand the time and progress of equipment upgrade and renovation, and determine the equipment upgrade and renovation scenario. The general process is as follows: 1. The system regularly obtains relevant information from each data source. 2. Integrate and analyze the obtained data such as time, passenger flow, meteorology, and maintenance plan. 3. According to the preset scenario judgment rules, map external equipment to the corresponding scenario categories.

[0073] Step S312: According to the corresponding relationship between the scenario category where the external equipment is located and the equipment status analysis algorithm, match the equipment status analysis algorithms adopted by different external equipment.

[0074] Equipment status analysis algorithm: A method used to process real-time data of external equipment and identify whether the equipment status has changed through specific calculation logics and models. Such as algorithms based on data threshold judgment, machine learning model analysis, etc.

[0075] Corresponding relationship between scenario category and equipment status analysis algorithm: A preset rule that clarifies which algorithm should be used to analyze the status change of different external equipment in various operation scenarios. The characteristics of equipment status changes in different scenarios are different, and different algorithms need to be adapted to accurately judge.

[0076] Specifically, for the daily operation scenario, the equipment status analysis algorithm adopts an algorithm based on threshold judgment. For example, the normal temperature range in the carriage is set to 24 - 26 °C. When the real-time temperature collected by the temperature sensor exceeds this range and the air conditioner operation mode is normal, the algorithm determines that there may be an abnormality in the cooling or heating effect of the air conditioner.

[0077] For the peak operation scenario, the equipment status analysis algorithm uses the decision tree algorithm in machine learning. Taking current, speed, acceleration, passenger flow, etc. as input features, the decision tree builds a model through learning historical data. For example, when the train accelerates, if the deviation of the current and acceleration from the normal situation is too large and the passenger flow is in the peak period, the decision tree algorithm judges that there may be an abnormality in the traction system.

[0078] Regarding the equipment maintenance scenario, the equipment status analysis algorithm adopts a rule-based expert system algorithm. Rules are set according to equipment maintenance standards and expert experience. For example, if the wear depth of the track exceeds a certain value or the contact pressure of the pantograph is lower than the standard range, the equipment is determined to be abnormal.

[0079] Regarding the extreme weather scenario, the equipment status analysis algorithm uses a fuzzy logic algorithm. Rainfall, humidity, fan operation parameters, etc. are used as inputs, and through fuzzy processing, fuzzy reasoning, and defuzzification, a judgment on the equipment status is obtained. For example, when the rainfall and humidity exceed a certain range and the fan operating current is abnormal, the fuzzy logic algorithm determines that there may be a risk of water ingress or blockage in the ventilation system.

[0080] Regarding the equipment upgrade and transformation scenario, the equipment status analysis algorithm adopts a comparative analysis algorithm. The operating parameters of the equipment after upgrade and transformation are compared with the design standard parameters, and at the same time, the key performance indicators of the equipment before and after transformation are compared. For example, the signal strength should reach more than 90% of the design standard. If the actual signal strength is lower than this value, the signal system status is determined to be abnormal.

[0081] The general process is as follows: The system uses the scenario category and equipment type as retrieval conditions for precise matching in the configuration file. For example, if it is determined that a subway ventilation equipment is in an extreme weather scenario, the system will search for the record of "extreme weather scenario - ventilation equipment" in the configuration file to obtain the corresponding equipment status analysis algorithm name.

[0082] According to the algorithm name obtained from the configuration file, the system retrieves in the algorithm library. After finding the corresponding algorithm, the code of the algorithm and related dependencies are loaded into the memory to complete the initialization of the algorithm, preparing for the next step of processing the real-time data of external devices.

[0083] Step S313, based on the real-time relevant data of external devices, use the matching equipment status analysis algorithm to analyze whether there is a change in the equipment status of external devices.

[0084] The general process is as follows: 1. Data collection: Continuously obtain the real-time relevant data of external devices from sensors and the equipment's own system. 2. Algorithm operation: Input the collected data into the equipment status analysis algorithm matched in step S312, and the algorithm processes and analyzes the data according to the preset calculation logic and model. 3. Result judgment: According to the algorithm analysis result, judge whether the equipment status has changed. If the result output by the algorithm exceeds the set range of the normal state or conforms to a specific state change pattern, it is determined that the equipment status has changed; otherwise, the equipment status is considered normal.

[0085] A method for dynamically monitoring the energy efficiency balance of rail transit also includes steps after collecting various information of the target equipment, specifically as follows:

[0086] Step S3A0: Analyze whether there are other devices that have a collaborative relationship with the target device. If the answer is no, then execute Step S3B0; if the answer is yes, then execute Step S3C0.

[0087] Collaborative relationship: In the rail transit system, there is a working connection where devices cooperate with and influence each other. For example, in a train, the traction system and the braking system. The traction system is responsible for providing power to make the train move forward, while the braking system controls the train to decelerate or stop. The two are closely related and have a collaborative relationship.

[0088] Collaborative device: Other devices that have a collaborative relationship with the target device. If the target device is the air - conditioning system of a train, then the ventilation system can be regarded as its collaborative device because they work together to regulate the air environment in the carriage.

[0089] The acquisition methods are disclosed as follows: 1. Equipment layout and connection drawings: By referring to the equipment layout diagram and connection circuit diagram of the rail transit system, clarify the physical connection and layout relationship between devices, and thus judge whether there is a collaborative relationship between devices. 2. System function specifications: The function specifications of the equipment detail the functions of the equipment and the interaction methods with other devices, from which information on the collaborative work of the equipment can be obtained.

[0090] The general process is as follows: 1. After determining the target device, first sort out the devices that are physically located or circuit - connected to the target device from the equipment layout and connection drawings. 2. Then, in combination with the system function specifications, analyze the cooperation relationship between these associated devices and the target device during the function implementation process to judge whether there is a collaborative relationship.

[0091] Suppose the target device is the ticket vending machine at a subway station. It is found from the equipment layout diagram that the ticket vending machine is connected to the station's network server. By checking the system function specifications, it is known that the ticket vending machine needs to obtain real - time ticket information, fare updates and other data from the server. The two have an obvious collaborative relationship. So in Step S3A0, it is judged that there is a device that has a collaborative relationship with the ticket vending machine, that is, the station network server.

[0092] Step S3B0: Maintain the original settings and continue with the subsequent steps.

[0093] Among them, the original settings refer to the various parameters and standards set according to the existing rules, models and methods after determining the target device, including the initial personalized threshold determined by using the life - cycle cost analysis method, and the established equipment operation monitoring and analysis processes in the system.

[0094] The general process is as follows:

[0095] After step S3A0 determines that there are no other devices in a collaborative relationship with the target device, the system first confirms all the original settings from the system configuration file. According to the data collection frequency and method specified in the original settings, it continues to collect relevant data of the target device, such as real-time operating status, energy consumption data, etc.

[0096] The collected data is processed and analyzed subsequently according to the established analysis process and criteria in the original settings to complete the entire dynamic monitoring process of energy efficiency balance.

[0097] Step S3C0: Collect various information of other devices in a collaborative relationship with the target device and the existing collaborative relationship data.

[0098] Collaborative relationship data: Specific data describing the interaction and influence between the target device and the collaborative device. For example, the degree of influence of the startup and stop of the collaborative device on the energy consumption of the target device, the data transmission volume and frequency between the two during collaborative work, etc.

[0099] Various information of collaborative devices: Information such as the real-time operating status, energy consumption data, and usage frequency of other devices in a collaborative relationship with the target device. This information helps to comprehensively understand the overall situation when devices work together.

[0100] The general process is as follows:

[0101] After determining other devices in a collaborative relationship with the target device, first establish a data connection with these collaborative devices and start collecting collaborative relationship data through the device interface.

[0102] Synchronously start collecting data from various monitoring sensors on the collaborative devices to obtain information such as their operating status and energy consumption.

[0103] Regularly extract the historical data of the collaborative devices from the system log records to supplement and verify the collected real-time data to ensure the integrity and accuracy of the information.

[0104] Step S3D0: Input the initial personalized threshold of the target device, various information of the target device, various information of other devices in a collaborative relationship with the target device, and the existing collaborative relationship data into the pre-trained system collaborative reinforcement learning threshold adjustment model to output an adjustment value.

[0105] System collaborative reinforcement learning threshold adjustment model: An intelligent model constructed based on the reinforcement learning algorithm, specifically used to handle the impact of the collaborative relationship between devices on the energy efficiency standard threshold. It optimizes its own parameters and strategies by continuously learning the operating data and collaborative relationship data of the devices to output a more realistic adjustment value.

[0106] Adjustment value: The value output by the system's collaborative reinforcement learning threshold adjustment model, which is used to correct the initial personalized threshold of the target device, so that the finally determined personalized energy efficiency standard threshold can more accurately reflect the energy consumption of the device in the collaborative working environment.

[0107] The general process is as follows:

[0108] The system sorts and preprocesses the relevant data of the target device and collaborative devices in the format required by the model, removes outliers and noise data, and ensures data quality.

[0109] Call the system's collaborative reinforcement learning threshold adjustment model and input the preprocessed data into the model. Inside the model, based on the reinforcement learning algorithm, it deeply analyzes and learns the input data to explore the potential relationship between the collaborative relationship between devices and energy consumption.

[0110] The model calculates and outputs the adjustment value according to the learning and analysis results, and this adjustment value reflects the influence degree of the collaborative work of the devices on the energy consumption standard of the target device.

[0111] Step S3E0, based on the adjustment value and the initial personalized threshold, calculate and determine the personalized energy efficiency standard threshold of the target device.

[0112] The general process is as follows: 1. The system extracts the adjustment value output in step S3D0 from the temporary data storage area, and at the same time retrieves the initial personalized threshold of the target device from the database. 2. Perform operations on these two values according to the pre-set calculation rules. Generally, if the adjustment value is positive, add this adjustment value to the initial personalized threshold; if the adjustment value is negative, subtract this adjustment value to obtain the final personalized energy efficiency standard threshold. 3. Update and store the calculated personalized energy efficiency standard threshold in the system database as the comparison basis for monitoring the energy consumption of the target device later, so as to judge in real time whether the device energy consumption is within the normal range.

[0113] Among them, the construction process of the system's collaborative reinforcement learning threshold adjustment model is as follows:

[0114] Step S3D1, collect the historical operation data of rail transit equipment, covering the operation status, energy consumption data of each device and the corresponding system energy consumption situation in different time periods and working conditions.

[0115] Historical operation data of rail transit equipment: It refers to a series of data generated by various equipment in the rail transit system during past operations, including the operating status of the equipment (such as on / off, operating speed, operating mode, etc.), energy consumption data (power consumption, gas consumption, etc.), and the energy consumption of the entire rail transit system corresponding to the operation of these equipment under different time periods (such as weekdays, holidays, morning and evening rush hours, etc.) and different operating conditions (normal operation, equipment maintenance, fault repair, etc.). These data are the basis for constructing a system collaborative reinforcement learning threshold adjustment model and can reflect the operating characteristics and energy consumption laws of the equipment under various conditions.

[0116] The general process is as follows: 1. Determine the time range and equipment types for data collection. For example, collect the operation data of all subway trains and station equipment in the past year. 2. Extract the corresponding data from the equipment monitoring system, energy management system, and operation and maintenance management records according to the predetermined time range and equipment list. 3. Conduct preliminary sorting of the extracted data and store it classified according to equipment types, time sequence, etc., for convenient subsequent data preprocessing and model training.

[0117] Step S3D2: Perform data preprocessing on the collected historical operation data of rail transit equipment. Among them, data preprocessing includes data cleaning and normalization processing.

[0118] Data cleaning: It refers to the review and correction of the collected historical operation data of rail transit equipment to remove incomplete, incorrect, duplicate, or noisy data. For example, abnormal energy consumption data caused by sensor failures or incorrect markings in the equipment operation status records need to be corrected through data cleaning to ensure the accuracy and reliability of the data and provide a high-quality data basis for subsequent analysis and model training.

[0119] Normalization processing: The process of converting data with different features to the same scale range. In the operation data of rail transit equipment, the numerical range of energy consumption data may vary greatly from that of equipment operation speed data. Through normalization processing, these data with different magnitudes can be made comparable, facilitating the model to better learn and process the relationships between data and improving the training effect and accuracy of the model.

[0120] Step S3D3: Perform an initialization operation on the system collaborative reinforcement learning threshold adjustment model. The initialization operation is to randomly assign values to the weight matrices and bias vectors of the policy network and the value network.

[0121] Policy Network: In reinforcement learning, the policy network is used to generate a series of possible actions and their corresponding probability distributions based on the current environmental state, that is, to determine what actions the agent should take in the current state. It is the core part of the model's decision-making. For example, in the scenario of energy efficiency management of rail transit equipment, the policy network will give the probabilities of actions such as adjusting the operating parameters of the equipment (such as changing the traction power of the train, adjusting the wind speed of the ventilation system) according to the current operating state of the equipment, energy consumption data, etc.

[0122] Value Network: It is used to evaluate the expected cumulative reward that the agent can obtain in the future after taking a certain action in a certain state, that is, to measure the value of the current state and action. In rail transit, the value network can evaluate the expected effect of adjusting the operating parameters of the equipment on reducing the overall energy consumption or improving the energy efficiency of the system, and help the model judge the quality of the actions.

[0123] Weight Matrix and Bias Vector: The weight matrix and bias vector are key parameters in neural networks (including the policy network and the value network). The weight matrix determines the connection strength between the input data and the neurons, and the bias vector adds a fixed value to the output of the neurons. They jointly affect the output result of the network. During the training process, these parameters will be continuously adjusted to optimize the model performance.

[0124] The general process is as follows: 1. Determine the network structures of the policy network and the value network, including the number of neuron layers, the number of neurons in each layer, etc. For example, construct a simple policy network that includes an input layer, two hidden layers, and an output layer. 2. According to the network structure, use a random number generator to generate random values for the weight matrix and bias vector of each layer of the policy network and the value network. For example, for a neural network with 10 neurons in the input layer and 20 neurons in the hidden layer, a 10×20 weight matrix and a bias vector with a length of 20 need to be generated, and their initial values are filled with random numbers. 3. Apply the generated random weight matrix and bias vector to the policy network and the value network to complete the initialization operation of the model, so that the model has preliminary processing capabilities and is ready to accept input data for training.

[0125] Step S3D4, starting from the initial state, the policy network gives the action probability distribution based on the current state, selects an action using the roulette wheel algorithm and applies it to the simulation environment, records the state, action, reward, and next state information, and collects a set of records reaching the preset quantity to form a trajectory set.

[0126] Initial State: At the starting stage of the reinforcement learning model training, it represents the environmental state where the agent is located. For the energy efficiency management of rail transit equipment, it is the initial operating state of the equipment when the model starts to run, including various parameters of the equipment (such as the initial speed of the train, the initial wind speed of the ventilation system, etc.) and the surrounding environmental conditions (such as the initial temperature in the station, the passenger flow, etc.).

[0127] Action probability distribution: A set of occurrence probabilities of various possible actions generated by the policy network based on the current state. For example, in the subway lighting system, possible actions include turning on / off some lights, adjusting the brightness of lights, etc. The action probability distribution indicates the likelihood of each action being selected in the current state.

[0128] Roulette wheel algorithm: A probability-based random selection algorithm used to select a specific action from the action probability distribution. It is like spinning a roulette wheel divided into different regions, each region corresponding to an action, and the size of the region is proportional to the action probability. The action corresponding to the region where the pointer stops when the roulette wheel stops is the selected action.

[0129] Simulation environment: A virtual environment constructed to train the reinforcement learning model, simulating the operation of the real rail transit system. In this environment, various actions of the equipment can be simulated, and the resulting effects can be observed, such as changes in energy consumption, changes in equipment status, etc.

[0130] Trajectory set: Composed of a series of records, each record containing state, action, reward, and next state information. These records reflect the running trajectory of the model in the simulation environment and are an important data basis for model learning and optimization.

[0131] The general process is as follows: 1. The model starts from the initial state and inputs the current state information into the policy network. 2. The policy network calculates and outputs the action probability distribution based on the input state. 3. Use the roulette wheel algorithm to randomly select an action according to the action probability distribution. 4. Apply the selected action to the simulation environment, observe and record the changes in the environment, including the obtained rewards (such as the degree of energy consumption reduction, improvement in equipment operation stability, etc.) and the next state entered. 5. Continuously repeat the above steps, collect records reaching a preset quantity (such as 1000 records), and organize these records into a trajectory set for subsequent model training and optimization.

[0132] In step S3D5, use the generalized advantage estimation method, first calculate the TD error, then calculate the advantage function, and judge the quality of the action.

[0133] TD error: That is, the temporal difference error, used to measure the difference between the predicted value and the actual value in the current state. In reinforcement learning, the TD error is calculated by comparing the value estimation in the current state with the value estimation in the next state plus the reward, which reflects the estimated deviation of the model from the environmental changes.

[0134] Advantage function: Represents the degree of advantage of taking a certain action compared to the average action in the current state. The larger the value of the advantage function, the better the return that the action can bring compared to other actions. The advantage function can be used to judge the quality of the action and provide a direction for model optimization.

[0135] The acquisition methods are disclosed as follows: 1. Value network output: Obtain the value estimates of the current state and the next state from the value network. The value network outputs an evaluation of the value in that state based on the input state information, and these output values are the basic data for calculating the TD error and the advantage function. 2. Reward value: In the simulation environment, the feedback reward obtained after executing an action, such as in the optimization of rail transit equipment, the reward brought by reduced energy consumption, or the reward brought by reduced equipment failures. These reward values directly participate in the calculation of the TD error and the advantage function.

[0136] The general process is as follows:

[0137] First, according to the value estimates of the current state and the next state by the value network, and the reward obtained after executing the action, use the formula to calculate the TD error. The formula is generally: TD error = reward + γ × value estimate of the next state - value estimate of the current state, where γ is the discount factor, which is used to measure the importance of future rewards.

[0138] Next, based on the TD error, use the Generalized Advantage Estimation (GAE) method to calculate the advantage function. This method comprehensively considers the TD errors of multiple time steps to more accurately evaluate the advantage of an action. For example, the advantage function value is obtained by accumulating the discounted TD errors of multiple time steps.

[0139] According to the calculated advantage function values, judge the pros and cons of each action. If the advantage function value is positive, it means that this action is better than the average action; if the advantage function value is negative, it means that this action is relatively poor.

[0140] Still taking the model for optimizing the energy consumption of subway train operation as an example. Suppose the current train is running at a speed of 40 km / h, and after executing an acceleration action, it enters the next state with a speed of 45 km / h. The value estimate of the value network for the current state is 0.5, the value estimate for the next state is 0.6, the reward obtained by executing the action is 0.1 (because the energy consumption has decreased), and the discount factor γ is set to 0.9. First, calculate the TD error: 0.1 + 0.9 × 0.6 - 0.5 = 0.14. Then use the Generalized Advantage Estimation method to calculate the advantage function. Suppose the advantage function value is 0.2 after a series of calculations. This indicates that this acceleration action has an advantage over the average action, and in subsequent model training, the probability of selecting such actions will tend to increase.

[0141] In step S3D6, use the PPO-Clip objective function in the Proximal Policy Optimization (PPO) algorithm to calculate the policy loss, and adjust the parameters through gradient descent to minimize the policy loss.

[0142] Proximal Policy Optimization Algorithm (PPO): An algorithm for optimizing the policy network, aiming to improve the performance of the policy in reinforcement learning. It restricts the magnitude of policy updates to prevent drastic changes in the policy during the update process, thereby enhancing the stability and efficiency of training.

[0143] PPO-Clip Objective Function: The core part of the PPO algorithm, used to calculate the policy loss. It clips the advantage ratio of the new and old policies to control the scope of policy updates, ensuring that the policy can develop in a more optimal direction during optimization without sacrificing performance due to excessive updates.

[0144] Policy Loss: Measures the gap between the current policy and the optimal policy. By minimizing the policy loss, the parameters of the policy network are optimized so that the policy network can output a more optimal action probability distribution, thereby improving the model's performance in practical applications.

[0145] The acquisition methods are disclosed as follows: 1. Advantage Function Value: The advantage function value calculated in step S3D5, which reflects the advantage degree of each action compared to the average action and is the key data for calculating the PPO-Clip objective function. 2. Action Probabilities of the New and Old Policies: The action probabilities output by the policy network for the same state before and after the update. By comparing the action probabilities of the new and old policies, the advantage ratio can be calculated and then used in the calculation of the PPO-Clip objective function.

[0146] The general process is as follows:

[0147] First, obtain the advantage function value calculated in step S3D5 and the action probabilities output by the policy network for the current state before and after the update.

[0148] Calculate the advantage ratio based on the action probabilities of the new and old policies. The formula is: Advantage Ratio = New Policy Action Probability / Old Policy Action Probability.

[0149] Substitute the advantage ratio and the advantage function value into the PPO-Clip objective function for calculation. The general form of the PPO-Clip objective function is: L = min(Advantage Ratio × Advantage Function Value, clip(Advantage Ratio, 1 - ε, 1 + ε) × Advantage Function Value), where the clip function is used to clip the advantage ratio, and ε is a hyperparameter, usually set to a small value such as 0.2, to control the magnitude of policy updates.

[0150] Using the gradient descent algorithm, adjust the parameters of the policy network to minimize the policy loss. During the gradient descent process, based on the gradient information of the objective function, continuously update the weight matrix and bias vector of the policy network to make the policy network gradually develop in a more optimal direction.

[0151] Step S3D7: Calculate the value loss using the mean squared error loss function and update the parameters through gradient descent.

[0152] Mean squared error loss function: Commonly used to measure the degree of difference between two numerical sequences. In model training, it calculates the average of the sum of the squares of the errors between the predicted values and the true values. By minimizing the mean squared error loss function, the predicted values of the model can be made as close as possible to the true values, improving the accuracy of the model.

[0153] Value loss: In reinforcement learning, the value network predicts the future value based on the current state. The value loss is the gap between the predicted value and the actual obtained value. By adjusting the parameters of the value network to reduce the value loss, the value network can more accurately evaluate the state value.

[0154] The acquisition method is disclosed as follows: 1. Predicted value of the value network: Obtain the predicted value of the future value in the current state from the value network. The value network calculates based on the input state information using its internal parameters and algorithms and outputs the predicted value. 2. Actually obtained value: After performing an action in the simulation environment, calculate the actually obtained value according to the actual rewards obtained and the value feedback of the subsequent state. For example, in the rail transit scenario, the actual energy consumption reduction and operation efficiency improvement after adjusting the operation parameters of the equipment are converted into the actually obtained value.

[0155] The general process is as follows:

[0156] First, obtain the predicted value of the value network for the current state and the actually obtained value after performing an action in the simulation environment.

[0157] Substitute the predicted value and the actually obtained value into the mean squared error loss function for calculation.

[0158] Through the gradient descent algorithm, update the parameters of the value network according to the calculated value loss. The gradient descent algorithm adjusts the weight matrix and bias vector of the value network along the opposite direction of the gradient of the loss function, making the value loss gradually decrease and the prediction of the value network more accurate.

[0159] Step S3D8: Run the mixed-integer linear programming model, calculate the theoretical optimal solution of the system energy consumption based on the physical characteristics of the rail transit system and the equipment operation constraint conditions, and calculate the gap between the current system energy consumption and the theoretical optimal solution.

[0160] Mixed-integer linear programming model: A mathematical optimization model used to handle linear programming problems containing integer variables and continuous variables. In the field of rail transit, this model can be used to optimize the calculation of the system energy consumption. The integer variables may represent the number of equipment startups, the number of operation mode switches, etc., and the continuous variables may be the operation parameters of the equipment (such as power, speed, etc.).

[0161] Theoretical optimal solution of system energy consumption: Based on the mixed-integer linear programming model, comprehensively considering the physical characteristics of the rail transit system (such as the energy conversion efficiency of equipment, line losses, etc.) and the operating constraints of equipment (such as the maximum and minimum operating power of equipment, operating time limits, etc.), the theoretical value with the minimum system energy consumption or the highest energy efficiency is calculated, which provides an ideal reference standard for evaluating the current system energy consumption.

[0162] The acquisition method is disclosed as follows: 1. Data collection: Collect various data of the rail transit system, including technical parameters of equipment (such as motor efficiency, transformer loss, etc.), operating constraints (such as the maximum operating speed of trains, the maximum air volume of the ventilation system, etc.), and the topological structure of the system (such as the power supply network layout, equipment connection relationship, etc.). These data are obtained from technical documents provided by equipment manufacturers, design drawings of the rail transit system, and actual operation monitoring data. 2. Model call: Use professional mathematical optimization software (such as Gurobi, CPLEX, etc.) to call the mixed-integer linear programming model, and use the collected data as the input parameters of the model to calculate the theoretical optimal solution of system energy consumption.

[0163] The general process is as follows: 1. Collect and organize various data related to the energy consumption of the rail transit system to ensure the accuracy and integrity of the data. 2. Input the organized data into the mixed-integer linear programming model, and set the objective function of the model (such as minimizing system energy consumption) and constraint conditions (such as equipment operation limits, physical law constraints, etc.). 3. Use mathematical optimization software to solve the model to obtain the theoretical optimal solution of system energy consumption. At the same time, obtain the energy consumption data of the current system's actual operation, which is monitored in real time through sensors or extracted from the energy management system. 4. Calculate the gap between the current system energy consumption and the theoretical optimal solution, which can be measured by the difference or percentage between the two.

[0164] Step S3D9, determine the adjustment strategy according to the corresponding relationship between the gap range in which the gap between the current system energy consumption and the theoretical optimal solution falls and the adjustment strategy.

[0165] Gap range: Different intervals divided according to the magnitude of the difference between the current system energy consumption and the theoretical optimal solution. For example, the difference within 0 - 5% is divided into one range, and 5% - 15% is divided into another range. Different ranges reflect the degree of deviation of the system energy consumption from the optimal state.

[0166] Adjustment strategy: A series of model optimization and strategy adjustment methods formulated for different gap ranges. For example, when the gap is small, only fine-tune the model parameters; when the gap is large, it is necessary to re-examine the model structure or optimization strategy.

[0167] The acquisition method is disclosed as follows:

[0168] Gap calculation result: Obtain the gap value between the current system energy consumption and the theoretical optimal solution from step S3D8. This value is obtained through mathematical calculations and reflects the degree of difference between the actual energy consumption of the system and the ideal energy consumption.

[0169] Strategy correspondence table: A pre-established correspondence table between the gap range and adjustment strategies, stored in the system configuration file or database. This table clearly stipulates the specific adjustment strategies corresponding to each gap range, facilitating quick query and invocation by the system.

[0170] The general process is as follows: First, obtain the gap value between the current system energy consumption and the theoretical optimal solution from step S3D8. According to the gap value, find the gap range to which it belongs in the pre-established strategy correspondence table. After determining the gap range, obtain the adjustment strategy corresponding to this range from the strategy correspondence table.

[0171] Step S3DA, change the action probability distribution and reward function output by the policy network according to the determined adjustment strategy.

[0172] Action probability distribution output by the policy network: A set of occurrence probabilities of various possible actions calculated by the policy network based on the current environmental state. In the rail transit scenario, such as the operation strategy of a subway train, the policy network will output the probabilities of different actions such as accelerating, decelerating, and maintaining a constant speed.

[0173] Reward function: Used to measure the benefits obtained after taking a certain action in a certain state, and it is an important basis for guiding the agent to learn in the reinforcement learning model. In energy consumption optimization, a positive reward is obtained for a decrease in energy consumption, and a negative reward is obtained for an increase in energy consumption.

[0174] The general process is as follows:

[0175] First, receive the determined adjustment strategy from step S3D9. This adjustment strategy is obtained based on the gap analysis between the current system energy consumption and the theoretical optimal solution, and is an important basis for subsequent operations.

[0176] Deeply analyze the impact of the adjustment strategy on the action probability distribution output by the policy network. For example, if the adjustment strategy emphasizes improving the train operation efficiency, then it is necessary to increase the probability of actions that can achieve this goal (such as optimizing the train acceleration and deceleration modes, reasonably distributing power output, etc.) in the action probability distribution. The specific implementation method is to adjust the parameters of the policy network to increase the probability values of these beneficial actions.

[0177] Precisely change the reward function according to the adjustment strategy. For example, to encourage reducing the wheelspin time of the train, when the action successfully reduces wheelspin, increase the reward value given by the reward function; when the action causes an increase in wheelspin time, increase the penalty, that is, reduce the reward value or give a negative reward.

[0178] Fully apply the adjusted action probability distribution and reward function to the policy network and the entire reinforcement learning model. This process involves reloading new parameters and configurations to ensure that the model can learn and optimize according to the new rules during subsequent training and decision-making processes.

[0179] Taking the escalator system of the subway as an example, assume that the previously determined adjustment strategy is to reduce the energy consumption of the escalator. In the action probability distribution output by the current policy network, the probability of the energy-saving action of switching the escalator to the low-speed operation mode during low passenger flow periods is 0.4. According to the adjustment strategy, to more effectively achieve the energy-saving goal, this probability is increased to 0.7. For the reward function, previously when the energy consumption of the escalator decreased by 15%, the reward value was 6, and now it is increased to 10; when the energy consumption remains unchanged, the reward value is reduced from the original 1 to -3. Through these adjustments, the model will more frequently select to switch the escalator to the low-speed operation mode during low passenger flow periods during subsequent training, thereby achieving the purpose of reducing the energy consumption of the escalator.

[0180] Step S3DB: Re-train the model based on the adjusted policy network action probability distribution and reward function.

[0181] Adjusted policy network action probability distribution: The new distribution obtained by adjusting the original action probability distribution of the policy network according to the adjustment strategy in step S3DA. It reflects the change in the model's decision-making tendency and guides the model to more frequently select actions that are beneficial to achieving system goals (such as reducing energy consumption and improving operation efficiency) during subsequent training.

[0182] Adjusted reward function: The new function formed by modifying the original reward function based on the adjustment strategy in step S3DA. The new reward function more effectively encourages the model to learn the optimal strategy by changing the reward values corresponding to different actions and states.

[0183] Model training: Use a large amount of training data and optimization algorithms to enable the model to continuously learn and improve to enhance its performance and accuracy in practical applications. In this scenario, model training is to enable the reinforcement learning model to find a better decision-making strategy according to the adjusted policy network action probability distribution and reward function.

[0184] The general process is as follows:

[0185] Read the adjusted policy network action probability distribution and reward function from step S3DA to ensure the accuracy of the data.

[0186] Prepare the training data and preprocess the data, such as cleaning outliers, normalizing, etc., to improve the training effect.

[0187] Apply the adjusted action probability distribution and reward function to the reinforcement learning model, and set the training parameters, such as the number of training epochs, learning rate, etc.

[0188] Start model training. In each training epoch, the model selects actions according to the current action probability distribution, obtains reward feedback according to the reward function, and then updates the model parameters through an optimization algorithm (such as gradient descent) to gradually improve the model performance.

[0189] Regularly evaluate the model during the training process, such as calculating metrics such as the loss value and accuracy of the model on the validation set, and judge whether the model converges or reaches the expected performance.

[0190] Step S3DC. After each training cycle ends, use the validation set data to evaluate the model performance. When the model reaches the set performance metrics on the validation set, determine that the system collaborative reinforcement learning threshold adjustment model has completed training.

[0191] Training cycle: The time period for the model to perform a complete training process, which includes multiple rounds of parameter updates and learning. During each training cycle, the model will traverse all training data or perform multiple trainings according to specific batches.

[0192] Validation set data: A part of the data divided from all the data, which is specifically used to evaluate the performance of the model during the training process. It does not participate in the training of the model. The purpose is to test the generalization ability of the model for unseen data and avoid overfitting of the model.

[0193] Set performance metrics: The model performance measurement criteria determined in advance according to the actual application requirements. For example, in the rail transit energy efficiency balance dynamic monitoring model, metrics such as the energy consumption prediction error within a certain range and the action decision accuracy reaching a certain value may be set. When the performance of the model on the validation set reaches these metrics, the model training is considered successful.

[0194] Validation set data: Usually in the data preprocessing stage, it is divided from the historical operation data of rail transit equipment collected according to a certain proportion (such as 20%-30%). The division method can use random sampling or stratified sampling to ensure that the validation set data can represent the characteristics of the overall data.

[0195] Set performance indicators: Determined by the model developer according to the project objectives and actual business requirements. For example, if the goal is to reduce the energy consumption of the rail transit system by 10%, one of the performance indicators can be set as the average absolute error between the predicted energy consumption of the model and the actual energy consumption within 5%.

[0196] The general process is as follows:

[0197] After each training cycle ends, apply the trained model to the validation set data. The model outputs actions according to the state information in the validation set data according to the currently learned policy network and calculates the corresponding rewards.

[0198] Based on the calculation results on the validation set, calculate various performance indicators of the model. For example, calculate the error between the predicted energy consumption of the model and the actual energy consumption in the validation set, and evaluate the consistency between the model's decision-making actions and the actual optimal actions, etc.

[0199] Compare the calculated performance indicators with the pre-set performance indicators. If the model reaches the set performance indicators on the validation set, it means that the model has been trained to the expected level, and determine that the system collaborative reinforcement learning threshold adjustment model has completed training; if not, continue the next training cycle, adjust the model parameters, and repeat the training and evaluation process.

[0200] The root causes of energy loss localization include:

[0201] Step S510, obtain multi-source data, and the multi-source data includes device personalized thresholds, historical energy consumption, operation history, maintenance records, and environmental factors.

[0202] Device personalized threshold: A standard value or range in terms of energy consumption, operation parameters, etc. customized for each device according to factors such as the device's own characteristics, operating environment, and historical data. For example, the personalized threshold of the traction energy consumption of each subway train will vary depending on the vehicle type and line characteristics.

[0203] Historical energy consumption: The energy consumption data of the device over a past period, which reflects the energy consumption pattern and trend of the device. By analyzing the historical energy consumption, the energy consumption level and energy consumption fluctuation of the device during normal operation can be understood.

[0204] Operation history: Records information such as when the device starts and stops, operation duration, and operation mode. These data help to judge the operation status and operation stability of the device.

[0205] Maintenance record: Covers information such as the maintenance time, maintenance content, and replaced parts of the device, which can reflect the health status and maintenance situation of the device and has important reference value for analyzing the possible device failure reasons for abnormal energy consumption.

[0206] Environmental factors: Include external environmental conditions such as temperature, humidity, altitude where the equipment operates. These factors may affect the energy consumption of the equipment. For example, in a high-temperature environment, the energy consumption of ventilation equipment may increase.

[0207] The general process is as follows:

[0208] Determine the scope of equipment and time range for which data needs to be obtained. For example, for all trains on a certain subway line, obtain data for the past year.

[0209] Export historical energy consumption and operation history data from the equipment monitoring system, and organize them according to time sequence and equipment number.

[0210] Access the equipment management database, extract personalized thresholds and maintenance records of the corresponding equipment to ensure the integrity and accuracy of the data.

[0211] Connect to the environmental monitoring system to collect environmental factor data in the equipment operation area during the same time period.

[0212] Integrate the obtained multi-source data and store it in a unified data storage platform for use in subsequent steps.

[0213] Step S520: Calculate the mean and standard deviation using historical energy consumption data to determine the energy consumption boundary, analyze historical energy consumption in multiple dimensions, and establish a normal energy consumption mode.

[0214] Mean: In statistics, the mean is the sum of a set of data divided by the number of data, representing the average level of the data. In historical energy consumption data, the mean reflects the average energy consumption status when the equipment operates normally.

[0215] Standard deviation: A statistic that measures the degree of dispersion of data, used to describe the fluctuation size of data relative to the mean. The larger the standard deviation, the greater the fluctuation of the energy consumption data; conversely, the smaller the fluctuation.

[0216] Energy consumption boundary: The energy consumption range determined based on the mean and standard deviation, generally centered on the mean, plus or minus a certain multiple of the standard deviation to define. The energy consumption within this boundary can be considered as the energy consumption range when the equipment is in normal operation.

[0217] Normal energy consumption mode: The pattern and mode presented by the equipment's energy consumption under normal circumstances summarized through analyzing historical energy consumption data in multiple dimensions (such as time, operating conditions, etc.), which is an important basis for judging whether the energy consumption is abnormal.

[0218] The general process is as follows:

[0219] Extract historical energy consumption data from the equipment monitoring system to ensure the integrity and accuracy of the data.

[0220] Use statistical methods to calculate the mean and standard deviation of historical energy consumption data. For example, use the mean() and std() functions in the NumPy library of Python for calculation.

[0221] Determine the energy consumption boundary based on the calculated mean and standard deviation. Generally, it can be set as the mean ± 2 times the standard deviation (the specific multiple can be adjusted according to the actual situation).

[0222] Analyze the historical energy consumption data from the time dimension (such as daily, weekly, monthly, quarterly, etc.) and the operating condition dimension (such as full load, no load, different speed ranges, etc.), and draw energy consumption trend charts, energy consumption distribution histograms, etc.

[0223] Through the summary and induction of the multi-dimensional analysis results, establish a normal energy consumption mode, and clarify the normal energy consumption range and change rules of the equipment under different times and operating conditions.

[0224] Step S530: Sort out the fault modes of the equipment from the system to the components, determine the occurrence probability of each sub-mode, construct a dynamic fault tree with the energy consumption exceeding the preset proportion of the normal boundary as the top event, describe the causal relationship with logic gates, and then find all possible fault paths through qualitative analysis. Take the events whose occurrence times exceed the preset proportion of the total number of faults as the high-frequency fault bottom events, and take the paths of the high-frequency fault bottom events as the troubleshooting directions.

[0225] Fault mode: Various fault manifestation forms that may occur in the equipment from the system level to the component level. For example, the fault modes of the braking system of a subway train may include excessive wear of brake pads, brake fluid leakage, etc.

[0226] Occurrence probability: The likelihood of each fault mode occurring during the operation of the equipment, usually obtained through statistical analysis of the equipment's historical fault data.

[0227] Top event: In fault tree analysis, the final fault event that is concerned and whose causes need to be analyzed. Here, the energy consumption exceeding the preset proportion of the normal boundary is used as the top event. For example, the energy consumption exceeds the normal boundary by 20%.

[0228] Dynamic fault tree: A logical model used to describe the causal relationship of system faults. Different from the traditional fault tree, it can consider the time sequence and dynamic characteristics of event occurrence. The causal relationship between various fault events is represented by logical gates (such as AND gates, OR gates, etc.).

[0229] Fault path: The path from the top event, through logical gates, back to a series of bottom events that cause the top event to occur. Each fault path represents a possible combination of faults that can cause the top event to occur.

[0230] High-frequency failure basic events: Among all failure paths, the basic events whose occurrence times exceed a preset proportion (such as 30%) of the total failure times. These basic events are common causes of equipment failures and are the key directions for investigation.

[0231] The acquisition methods are as follows: 1. Equipment historical failure data: Obtained from equipment maintenance records and fault reporting systems, which record various fault information that has occurred to the equipment in the past, including fault time, fault phenomenon, fault cause, etc. 2. Industry statistical data: Obtained by consulting industry reports, communicating with other operating units of the same type of equipment, etc., and used to supplement and verify the fault probability analysis of the equipment in this unit.

[0232] The general process is as follows:

[0233] Comprehensively sort out all possible failure modes of the equipment from the system to the components, and the technical manuals and maintenance experience of the equipment can be referred to.

[0234] According to the equipment historical failure data and industry statistical data, count the occurrence probability of each failure mode.

[0235] Determine that the energy consumption exceeding the preset proportion of the normal boundary is the top event. Starting from the top event, use logic gates (such as "AND gate" indicating that multiple events occur simultaneously to cause the top event, and "OR gate" indicating that as long as one event occurs, it will cause the top event) to connect each failure mode and construct a dynamic fault tree.

[0236] Conduct qualitative analysis on the dynamic fault tree, and find all possible failure paths through logical reasoning.

[0237] Count the occurrence times of the basic events in each failure path, calculate their proportion of the total failure times, determine the basic events exceeding the preset proportion as high-frequency failure basic events, and take the paths of these high-frequency failure basic events as the main investigation directions.

[0238] Step S540: Convert the fault tree events into Bayesian network nodes and edges, determine the prior probability based on the equipment historical failures and industry statistics, input the multi-source data as evidence, use Bayes' theorem to update the posterior probability, and consider that the possibility of the fault cause increases significantly when the probability change exceeds 0.1, and calculate the fault probability accordingly.

[0239] Bayesian network: A graphical network model based on probabilistic reasoning, composed of nodes and edges. Nodes represent random variables. In energy loss analysis, these variables can be equipment failure events; edges represent the dependence relationships between variables, and this relationship is quantified through a conditional probability table. It can effectively process uncertain information and update the judgment of the occurrence probability of an event through known information.

[0240] Prior probability: The preliminary estimate of the probability of an event based on past experience or statistical data without additional evidence. For example, based on the historical failures of the equipment and industry statistics, the probability of a certain failure is estimated in advance.

[0241] Posterior probability: The probability of an event obtained by revising the prior probability after obtaining new evidence (such as multi-source data such as the current operating data and environmental data of the equipment). It more accurately reflects the likelihood of the event occurring under the current circumstances.

[0242] The general process is as follows:

[0243] One-to-one correspondence between the events in the fault tree is transformed into the nodes of the Bayesian network, and the causal relationship between the events in the fault tree is transformed into the edges between the nodes of the Bayesian network.

[0244] Based on the historical failures of the equipment and industry statistical data, determine the prior probability of each node (fault event), and establish the corresponding conditional probability table to describe the dependence relationship between the nodes.

[0245] Take the multi-source data obtained from step S510 as evidence and input it into the Bayesian network.

[0246] Use Bayes' theorem for probability inference to update the posterior probability of each node.

[0247] Set a probability change threshold, such as 0.1. When the posterior probability of a certain fault event changes by more than this threshold relative to the prior probability, it is considered that the likelihood of the fault cause has increased significantly, and it is taken as the key focus object.

[0248] Step S550, fuse the logical structure of the fault tree and the fault probability of the Bayesian network, establish an evaluation matrix to score the fault modes from two dimensions, and screen out the factors with a comprehensive score exceeding the preset value and an energy consumption increase exceeding the preset ratio as the root causes.

[0249] Evaluation matrix: A tool for comprehensive analysis and evaluation, presented in tabular form. In this step, the rows and columns of the evaluation matrix represent different fault modes and evaluation dimensions respectively. By scoring each fault mode in each dimension, the key fault factors can be intuitively compared and screened out.

[0250] Fault mode scoring: Quantitatively evaluate each fault mode from two dimensions: the logical structure of the fault tree and the fault probability calculated by the Bayesian network. The scoring of the fault tree logical structure dimension can be based on the complexity of the fault path, the number of key equipment involved, etc.; the fault probability dimension directly refers to the posterior probability calculated by the Bayesian network.

[0251] Comprehensive Score: The weighted sum of the scores of each failure mode in different dimensions of the evaluation matrix, where the weights are set according to actual requirements and the importance of each dimension. Through the comprehensive score, the impact degree of each failure mode on energy loss can be comprehensively measured.

[0252] The general process is as follows:

[0253] Construct an evaluation matrix. The rows of the matrix are various failure modes, and the columns are two evaluation dimensions: the fault tree logic structure and the Bayesian network failure probability.

[0254] For each failure mode, in the dimension of the fault tree logic structure, analyze the complexity of its fault path. For example, a complex fault path involving multiple "AND gates" connections is given a lower score; a simple and direct path leading to the top event is given a higher score. In the dimension of failure probability, score according to the size of the posterior probability calculated by the Bayesian network. The higher the probability, the higher the score.

[0255] Set weights for the two evaluation dimensions. For example, the weight of the fault tree logic structure is set to 0.4, and the weight of the Bayesian network failure probability is set to 0.6. Calculate the comprehensive score of each failure mode. The formula is: Comprehensive Score = Fault Tree Logic Structure Score × 0.4 + Bayesian Network Failure Probability Score × 0.6.

[0256] Set a preset score for the comprehensive score, such as 60 points, and at the same time set a preset proportion of energy consumption increase, such as 15%. Screen out the factors corresponding to the failure modes whose comprehensive scores exceed the preset score and cause the energy consumption increase to exceed the preset proportion, and determine these factors as the root causes of energy loss.

[0257] Still taking the subway ventilation system as an example, construct an evaluation matrix. The failure modes include fan blade damage, motor burnout, and air duct blockage. In the dimension of the fault tree logic structure, the fault path of fan blade damage is simple, and it is scored 8 points; the motor burnout involves multiple component associations, and the fault path is complex, and it is scored 4 points; the fault path of air duct blockage is relatively simple, and it is scored 7 points. In the dimension of the Bayesian network failure probability, according to the previous calculation, the posterior probability of fan blade damage is 0.15, corresponding to a score of 3 points; the posterior probability of motor burnout is 0.1, scored 2 points; the posterior probability of air duct blockage is 0.38, scored 7 points. Calculate the comprehensive score according to the above weights. The comprehensive score of fan blade damage = 8 × 0.4 + 3 × 0.6 = 5; the comprehensive score of motor burnout = 4 × 0.4 + 2 × 0.6 = 2.8; the comprehensive score of air duct blockage = 7 × 0.4 + 7 × 0.6 = 7. Assuming the preset score is 5 points and the preset proportion of energy consumption increase is 15%, after detection, the air duct blockage causes an energy consumption increase of 20%, exceeding the preset proportion. Therefore, the factors related to the air duct blockage are determined as the root causes of energy loss.

[0258] A method for dynamically monitoring the energy efficiency balance of rail transit also includes steps parallel to those sent to the terminals held by the person in charge, specifically as follows:

[0259] Step S700: Input the located root cause, relevant data of the equipment, and industry standard information of the equipment into the solution analysis model, and output a solution. Among them, the solution includes problem description, solution overview, implementation steps, expected effects, and risk assessment.

[0260] Solution analysis model: An intelligent model based on data analysis and algorithms that can generate targeted solutions according to various input information. This model has been trained with a large amount of data and has the ability to analyze problems and provide effective solution strategies.

[0261] Problem description: Elaborate in detail on the located root cause of energy loss, including the equipment where the fault occurred, the fault phenomenon, the scope of influence, etc., so that relevant personnel can quickly understand the essence of the problem.

[0262] Solution overview: Briefly summarize the overall idea and main methods for solving the energy loss problem, and provide a macro solution framework.

[0263] Implementation steps: Refine the solution into specific operation steps, clarify the execution order, responsible person, and time node of each step to ensure the operability of the solution.

[0264] Expected effects: Predict the possible effects in terms of reducing energy loss and improving the operation efficiency of the equipment after implementing the solution, and provide a basis for evaluating the effectiveness of the solution.

[0265] Risk assessment: Identify and analyze the possible risks during the implementation of the solution, including technical risks, operation risks, cost risks, etc., and propose corresponding countermeasures.

[0266] The general process is as follows: 1. Extract the located root cause of energy loss from step S550. 2. Collect relevant data of the equipment from the equipment monitoring system and the equipment management database, such as operation duration, number of faults, energy consumption fluctuation conditions, etc. 3. Obtain the industry standard information of the equipment, including energy consumption standards, operation parameter standards, etc., through channels such as Internet search and industry conference materials. 4. Organize the root cause, equipment-related data, and industry standard information into a standardized format and input it into the solution analysis model. 5. The model generates a solution, including problem description, solution overview, implementation steps, expected effects, and risk assessment, etc., according to the input information by using internal algorithms and logics.

[0267] Step S800: Send the solution to the terminal held by the person in charge.

[0268] Terminal held by the person in charge: Generally refers to the electronic devices used by relevant personnel responsible for the management, maintenance or operation of rail transit equipment, such as mobile phones, tablets, laptops, etc., which are used to receive and view important work-related information.

[0269] Further considering that before sending the solution to the terminal held by the person in charge, the solution can also be simulated and verified. If there are problems, timely adjustments can be made to better ensure that the solution received by the person in charge meets the requirements of the scenario. Therefore, after the solution is output, the following steps should also be included:

[0270] Step SA00, build a simulation environment similar to the actual situation on the simulation platform, and use digital twin and 3D modeling technologies to create a virtual scene according to the actual rail transit information. Import the detailed device parameters, historical operation data and environmental data to ensure that the simulation environment truly reflects the actual situation.

[0271] Among them, the simulation platform: Select professional software, such as OMNeT++ and SUMO. OMNeT++ is good at simulating communication networks and can simulate the communication between trains and the dispatching center in rail transit; SUMO focuses on traffic flow simulation and can present the changes in the train operation trajectory.

[0272] Digital twin technology: Create a virtual model corresponding to the real device or system, and reflect its state in real time. For example, build a digital twin model for each subway train, and update the state of the virtual model according to the real-time data of the train sensors (such as speed, position).

[0273] 3D modeling technology: Used to build a realistic virtual scene, covering track lines, stations and equipment. For example, build a 3D model of a subway station to accurately present the layout of the platform, passageways and the location and appearance of equipment such as escalators.

[0274] Detailed device parameters: Include the physical and operation parameters of the device. Obtained from the device technical documents, such as the rated power of the train traction motor; can also be collected by device sensors, such as the wear data of the train brake pads. For example, the rated power of a certain train traction motor is 3000kW, which affects the simulation of train power and energy consumption.

[0275] Historical operation data: From the rail transit operation database, including the operation status and energy consumption data of the equipment under different time periods (weekdays, peak hours, etc.) and working conditions (normal, faulty, etc.). For example, for the lighting system of a certain station, the historical data shows that the opening time is long and the energy consumption is high during peak hours, providing a reference for simulation.

[0276] Environmental data: Include meteorological and geographical information. Obtain temperature and humidity through the meteorological data interface; obtain altitude and track gradient from the geographical information system. For example, a certain section of the track has a high altitude and a large gradient, which affects the train energy consumption and needs to be accurately set during simulation.

[0277] Step SB00: Adopt a simulation test scheme based on orthogonal experimental design. Set multiple levels with operating conditions, environment, and fault scenarios as factors. Use the discrete event simulation algorithm to collect energy consumption and equipment operation parameter data at a frequency of once per second. Repeat each test combination 10 times and take the average value.

[0278] Orthogonal experimental design: This is a scientific experimental design method. By reasonably arranging the test combinations of multiple factors and multiple levels, it can obtain comprehensive and representative results with fewer test times. For example, in rail transit simulation, considering factors such as operating conditions, environment, and fault scenarios, different levels are set for each factor. The operating conditions can be divided into three levels: peak, off-peak, and low-peak; the temperature in the environmental factor can be set to three levels: high temperature, normal temperature, and low temperature; the fault scenarios can include different types such as key component failures of the train and signal system failures as levels. Through combination with the orthogonal table, ensure that the test covers all possible situations, effectively reduce the test workload, and at the same time ensure the reliability and effectiveness of the results.

[0279] Discrete event simulation algorithm: This algorithm regards various activities in the rail transit system as discrete events, such as the arrival and departure of trains, the startup and stop of equipment, the occurrence of faults, etc. It is simulated based on the time sequence of event occurrence. By constructing an event queue and corresponding processing logic, it accurately simulates the dynamic change process of the system over time. For example, when simulating the train operation, after the train departure event is triggered, according to the predetermined operation plan and speed curve, during the advancement of the simulation time, successively process events such as the acceleration, constant speed, and deceleration of the train, as well as possible equipment failures or environmental changes on the way, so as to truly reproduce the operation process of the rail transit system.

[0280] Data collection: During the simulation process, collect energy consumption and equipment operation parameter data at a frequency of once per second. The energy consumption data covers train operation energy consumption, station equipment energy consumption, etc., such as train traction energy consumption, power consumption of station lighting and ventilation equipment, etc. The equipment operation parameters include the speed, acceleration, temperature inside the carriage, working current, voltage, rotation speed of the equipment, etc. The acquisition method of these data is mainly through the data acquisition module built in the simulation platform, and automatically collect according to the set frequency and parameter categories. Repeat each test combination 10 times and take the average value, aiming to reduce the influence of random errors in the simulation process, make the collected data more stable and representative, and thus provide a reliable basis for subsequent analysis and scheme adjustment.

[0281] Step SC00: Use a deep neural network (DNN) to construct a multi-layer perceptron (MLP) structure model. Train it with the mean squared error (MSE) as the loss function using the stochastic gradient descent (SGD) algorithm. Compare the simulation results with the expected results to find the differences. If there are differences, use a genetic algorithm to optimize and adjust the solution. Use the implementation steps and technical parameters as genes, and perform iterative optimization through selection, crossover, and mutation operations. Adjust the solution according to the optimal solution.

[0282] Deep Neural Network (DNN) and Multi-Layer Perceptron (MLP): DNN is a type of neural network that contains multiple hidden layers and can automatically learn complex patterns and relationships in data. MLP is a common structure of DNN, consisting of an input layer, multiple hidden layers, and an output layer. In this step, the input layer receives the energy consumption and equipment operation parameter data collected by the simulation, such as the speed and energy consumption values of the train at different times. The hidden layer processes and extracts features from the data through a non-linear activation function, and the output layer outputs the relevant results for comparison with the expected results. For example, when inputting the speed, acceleration, and energy consumption data of a train during a certain operation period, the MLP learns and calculates to output a judgment result on whether the energy consumption of the train meets the expectation.

[0283] Mean Squared Error (MSE) and Stochastic Gradient Descent (SGD): MSE is an index that measures the error between the predicted value and the true value of a model, and evaluates the accuracy of the model by calculating the average of the squares of the differences between the predicted value and the true value. In this scenario, it is used to measure the deviation degree between the simulation results and the expected results. SGD is an optimization algorithm. When training the MLP model, a small part of the data (a batch) is randomly selected each time to calculate the gradient, and the model parameters are updated accordingly, so that the model is continuously optimized to reduce the MSE value. For example, the MSE value between the initial model's prediction of the train's energy consumption and the actual energy consumption is relatively high. After multiple SGD iterations to update the model parameters, the MSE value gradually decreases, and the prediction accuracy of the model improves.

[0284] Genetic Algorithm: This is an optimization algorithm based on natural selection and genetic mechanisms. Acquisition method and quantification: Implement it using a genetic algorithm library in Python (such as DEAP). First, determine the encoding method of the genes. For example, for the set value of the train running speed, real number encoding can be used, and the value range is set according to the actual situation (such as 0 - 120 km / h). Define the fitness function, and calculate the fitness of an individual according to the difference degree between the simulation results and the expected results. For example, if the expected energy consumption is reduced by 10%, and the simulation result corresponding to a certain individual is only reduced by 5%, then its fitness is relatively low. Set the parameters of the genetic algorithm, such as the population size (such as 50 individuals), crossover probability (such as 0.8), mutation probability (such as 0.01), etc. During the iterative process, in each generation, individuals are selected, crossed, and mutated according to their fitness. After evolving through multiple generations (such as 100 generations), a better solution is found.

[0285] For example, in the verification of the energy efficiency solution for a certain subway line, data such as the energy consumption of trains and the operating parameters of equipment obtained through simulation at different time periods are input into the constructed MLP model, and the model is trained using the SGD algorithm to calculate the MSE. If it is found that in a certain simulation scenario, the deviation between the predicted value and the expected value of the train's energy consumption during peak hours is large, and the MSE reaches 0.3. At this time, the genetic algorithm is used to optimize the train operation plan and equipment control parameters in the solution. Suppose the original planned stop time of the train at a certain station is 30 seconds. After optimization by the genetic algorithm, the stop time is adjusted to 25 seconds, and the operating power of the equipment is also adjusted accordingly. Then, the simulation verification is carried out again to observe whether the MSE decreases and whether the simulation results are closer to the expected effect. After multiple iterative adjustments, until the MSE is reduced to an acceptable range (such as below 0.05), and the simulation results meet the expected goals of energy consumption reduction and equipment operation efficiency improvement, the optimized solution is finally determined.

[0286] Based on the same inventive concept, an embodiment of the present invention provides a rail transit energy efficiency balance dynamic monitoring system, including a memory and a processor. A program that can be run on the processor to implement any of the methods as Figures 1 to 3 any one of the methods is stored on the memory.

[0287] The embodiments of the specific implementation manners are all preferred embodiments of the present application, and do not limit the protection scope of the present application accordingly. Therefore, all equivalent changes made according to the structure, shape, and principle of the present application shall be covered by the protection scope of the present application.

Claims

1. A method for dynamic monitoring of rail transit energy efficiency balance, characterized in that: include: Get relevant data of various external devices; Input relevant data of various external devices into the pre-built fuzzy clustering classification model, and output the classification results of various external devices; Based on the equipment classification results and the real-time related data of various external equipment, analyze whether there is any change in equipment status; If not, use the life cycle cost analysis method to determine the personalized energy efficiency standard thresholds for each type of equipment; If yes, the device whose status has changed is taken as the target device, and the personalized energy efficiency standard threshold of the target device is determined by using the life cycle cost analysis method as the initial personalized threshold, and various information of the target device is collected, and various information of other devices that have a collaborative relationship with the target device and the existing collaborative relationship data are collected; Inputting the initial personalized threshold of the target device, various information of the target device, various information of other devices that have a collaborative relationship with the target device, and existing collaborative relationship data into a pre-trained system collaborative reinforcement learning threshold adjustment model, and outputting an adjustment value; Based on the adjustment value and the initial personalized threshold, calculate and determine the personalized energy efficiency standard threshold of the target device; According to the comparison result between the actual energy consumption of the equipment and the personalized energy efficiency standard threshold, determine whether the corresponding equipment has energy consumption abnormality; If yes, then the system combines the personalized threshold of the equipment and the historical energy consumption data, uses the built comprehensive energy loss analysis model based on fault tree and Bayesian network, starts from the equipment failure mode, combines the equipment operation history, maintenance records, and environmental factors, and locates the root cause of energy loss through Bayesian network reasoning and dynamic fault tree analysis, and sends it to the terminal held by the person in charge; If no, no notification will be made; The construction process of the system collaborative reinforcement learning threshold adjustment model is as follows: Collect historical operation data of rail transit equipment, covering the operation status, energy consumption data and corresponding system energy consumption of each equipment under different time periods and working conditions; Perform data preprocessing on the collected historical operation data of rail transit equipment, wherein the data preprocessing includes data cleaning and normalization processing; Initialize the threshold adjustment model of system collaborative reinforcement learning. The initial operation is to randomly assign values ​​to the weight matrix and bias vector of the policy network and value network. Starting from the initial state, the policy network gives the action probability distribution according to the current state, uses the roulette algorithm to select actions to be applied to the simulation environment, records the state, action, reward and next state information, and collects a preset number of records to form a trajectory set; Using the generalized advantage estimation method, first calculate the TD error, then calculate the advantage function to judge the quality of the action; The PPO-Clip objective function in the proximal policy optimization algorithm is used to calculate the policy loss, and the parameters are adjusted by gradient descent to minimize the policy loss. Use the mean square error loss function to calculate the value loss and update the parameters by gradient descent; Run a mixed integer linear programming model to calculate the theoretical optimal solution for system energy consumption based on the physical characteristics of the rail transit system and the equipment operation constraints, and calculate the gap between the current system energy consumption and the theoretical optimal solution; Determine the adjustment strategy according to the corresponding relationship between the gap range between the current system energy consumption and the theoretical optimal solution and the adjustment strategy; Changing the action probability distribution and reward function output by the policy network according to the determined adjustment strategy; Retrain the model based on the adjusted policy network action probability distribution and reward function; After each training cycle, the model performance is evaluated using the validation set data. When the model reaches the set performance indicator on the validation set, the system collaborative reinforcement learning threshold adjustment model is determined to complete the training; Locating the root causes of energy loss includes: Obtain multi-source data, including device personalized thresholds, historical energy consumption, operation history, maintenance records, and environmental factors; Use historical energy consumption data to calculate the mean and standard deviation to determine energy consumption boundaries, analyze historical energy consumption in multiple dimensions, and establish a normal energy consumption model; Sort out the failure modes of the equipment from the system to the components, determine the probability of occurrence of each subdivided mode, build a dynamic fault tree with the energy consumption exceeding the preset proportion of the normal boundary as the top event, use logic gates to describe the cause and effect relationship, and then find all possible fault paths through qualitative analysis. Events with a frequency exceeding the preset proportion of the total number of faults are regarded as high-frequency fault bottom events, and the paths of high-frequency fault bottom events are used as the troubleshooting direction; Convert fault tree events into Bayesian network nodes and edges, determine prior probabilities based on historical equipment failures and industry statistics, use multi-source data as evidence input, and use Bayesian theorem to update posterior probabilities. When the probability change exceeds 0.1, it is considered that the possibility of the fault cause has increased significantly, and the fault probability is calculated accordingly. By integrating the logical structure of the fault tree and the failure probability of the Bayesian network, an evaluation matrix is ​​established to score the failure modes from two dimensions, and factors with comprehensive scores exceeding the preset scores and energy consumption increases exceeding the preset proportions are screened as root causes.

2. A rail transit energy efficiency balance dynamic monitoring method according to claim 1, characterized in that: Based on the real-time data of various external devices, analyze whether there are any changes in the device status, including: Obtain the scenario category of the external device, where the scenarios include daily operation scenarios, peak operation scenarios, equipment maintenance scenarios, extreme weather scenarios, and equipment upgrade and transformation scenarios; According to the correspondence between the scene category of the external device and the device state analysis algorithm, the device state analysis algorithm adopted by different external devices is matched; Based on the real-time relevant data of external devices, a matching device status analysis algorithm is used to analyze whether there is any device status change in the external device.

3. A method for dynamic monitoring of rail transit energy efficiency balance according to any one of claims 1 to 2, characterized in that: It also includes steps parallel to sending to the terminal held by the responsible person, as follows: Input the root cause of the location, relevant data of the equipment, and industry standard information of the equipment into the solution analysis model, and output the solution, where the solution includes problem description, solution overview, implementation steps, expected results, and risk assessment; The solution is sent to the terminal held by the person in charge.

4. A rail transit energy efficiency balance dynamic monitoring system, characterized in that: It includes a memory, a processor, and a program stored in the memory and executable on the processor, wherein the program can be loaded and executed by the processor to implement a method for dynamic monitoring of energy efficiency balance of rail transit as described in any one of claims 1 to 3.

Citation Information

Patent Citations

  • AI agent generation method, platform system, electronic equipment and storage equipment

    CN119537959A

  • Fine energy consumption analysis and optimization management method and system based on rail transit

    CN119578834A