An interlocking host computer monitoring system and method

CN122653931APending Publication Date: 2026-08-28CASCO SIGNAL LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610742011.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-27
Publication Date
2026-08-28

AI Technical Summary

Technical Problem

[0007]本发明的目的是提供一种联锁上位机监控系统及方法,有效解决了现有联锁上位机故障多发、故障诊断低效、运维处置不规范的技术问题

Benefits of technology

1)本发明能够有效缩短联锁上位机的故障检测与识别耗时,便于运维人员及时开展维护作业。尤其针对联锁上位机长期运行老化引起的各类故障,能够实现提前预警与及时告警,大大减少故障突发工况。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122653931A_ABST
    Figure CN122653931A_ABST
Patent Text Reader

Abstract

The application discloses an interlocking host computer monitoring system and method, a computer interlocking system comprising a plurality of interlocking host computers, the plurality of interlocking host computers comprising a maintenance machine and an operating machine, the maintenance machine being in communication connection with the operating machine, and the system comprising: a plurality of host computer state monitoring modules, respectively arranged on the plurality of interlocking host computers, the host computer state monitoring module being automatically started along with the starting of the interlocking host computer; the host computer state monitoring module comprising: a state monitoring submodule for periodically collecting the running state information of the interlocking host computer; an alarm and early warning submodule configured to generate corresponding alarm and early warning information according to the running state information; the maintenance machine comprising a first processing module configured to diagnose the fault causes of the operating machine and the maintenance machine and recommend corresponding fault handling strategies according to the running state information and the alarm and early warning information of the operating machine and the running state information and the alarm and early warning information of the maintenance machine.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of rail transit technology, and in particular to an interlocking upper computer monitoring method and system. Background Technology

[0002] The computer interlocking system comprises multiple industrial control computers (also known as interlocking host computers), including operator computers and maintenance computers, each running its own software. With prolonged system operation, core components of the industrial control computers, such as the motherboard, memory, and hard drive, will gradually age and experience performance degradation. Simultaneously, issues such as cache redundancy and compatibility degradation generated during software operation will become increasingly apparent, leading to intermittent operational anomalies in the industrial control computers and their supporting software. These anomalies manifest as software response delays, system thread and process lag, and abnormal internal data transmission. If these problems are not addressed promptly and effectively, they can lead to serious consequences such as abnormal display at interlocking stations and failure to execute train control commands, thereby hindering train dispatching, significantly reducing operational efficiency, and even posing a potential threat to train safety.

[0003] To alleviate the aforementioned problems, existing technologies commonly employ a software watchdog mechanism to achieve routine fault detection and operational control of the operator and maintenance software. This mechanism continuously monitors the software's operational status by embedding a timed detection module. When a software crash, deadlock, or severe lag is detected (i.e., parameter values ​​representing the operational status exceed preset normal thresholds), a software restart process is automatically triggered, attempting to restore normal operation by restarting the software process and reducing the adverse impact of software failures on the overall operation of the interlocking system. However, long-term practical verification has revealed several shortcomings in this technical solution, making it difficult to fundamentally solve operational problems caused by system aging. These shortcomings are specifically manifested in the following aspects: 1) Fault response is delayed, and the software restart process affects system continuity. The software watchdog mechanism relies on periodically running code and comparing it with preset time parameters to complete the operation status detection. However, in the early stages of system aging, the code cycle runtime is not continuously affected, resulting in a significant time lag between the occurrence of a fault and its detection and identification by the watchdog mechanism. During this period, the interlocking function may already be in an abnormal state, which will still interfere with the normal operation of rail transit.

[0004] 2) Fault handling methods may lead to an expansion of the fault scope, resulting in secondary impacts from fault cascading. If the operator malfunctions and the software watchdog triggers a restart, it will not only directly interfere with interlocking control and station display functions, but also spread the fault's impact to higher-level external systems such as CTC (Centralized Traffic Control), TDCS (Train Dispatching and Commanding System), STP (Shunting Train Protection), and marshalling yard integrated automation. If the maintenance machine malfunctions and triggers a software restart, it will directly affect the daily operation and maintenance of interlocking equipment and spread the fault's impact to the signal monitoring system.

[0005] 3) Inability to accurately identify the root cause of the fault, and blindly restarting may mask the core problem. Software watchdogs can only monitor operational anomalies at the software level and cannot determine whether the anomaly is caused by software vulnerabilities or resource leaks, or by underlying causes such as aging industrial control computer hardware (e.g., memory failure, hard drive bad sectors) or external environmental interference. For anomalies caused by hardware aging or issues not stemming from the software itself, simply restarting the software can only achieve short-term recovery, and similar faults will recur. Furthermore, frequent software restarts accelerate hardware wear and tear, while masking the crucial underlying issue of hardware repair or replacement, increasing the difficulty and time cost of troubleshooting.

[0006] The statements herein provide only background information in relation to this invention and do not necessarily constitute prior art. Summary of the Invention

[0007] The purpose of this invention is to provide an interlocking host computer monitoring system and method that effectively solves the technical problems of frequent failures, inefficient fault diagnosis, and non-standard operation and maintenance in existing interlocking host computers. This invention can effectively reduce the probability of interlocking host computer failures and can reliably diagnose and recommend corresponding fault handling strategies after a failure occurs, significantly improving the efficiency of fault repair. This invention avoids frequently restarting the interlocking host computer to mask the root cause of the fault and prevents the fault from spreading within the computer interlocking system, thereby extending the service life of the interlocking host computer and ensuring the overall stable and reliable operation of the computer interlocking system.

[0008] To achieve the above objectives, the present invention provides an interlocking host computer monitoring system. The computer interlocking system includes multiple interlocking host computers, each comprising a maintenance unit and an operator unit. The maintenance unit and the operator unit are communicatively connected. The system includes: Multiple host computer status monitoring modules are deployed on multiple interlocking host computers; each host computer status monitoring module starts automatically when the host computer it belongs to starts. The host computer status monitoring module includes: The status monitoring submodule is used to periodically collect the operating status information of the interlocking host computer to which it belongs; An alarm and early warning submodule is configured to generate corresponding alarm and early warning information based on the operating status information; The repair machine includes a first processing module, which is configured to diagnose the cause of the fault and recommend corresponding fault handling strategies for the operating machine and the repair machine respectively, based on the operating status information and alarm warning information of the operating machine and the operating status information and alarm warning information of the repair machine.

[0009] Optionally, there are multiple operating machines, and each operating machine is connected to the maintenance machine via a dual network; The operator includes a second processing module, which is configured to calculate a health value for the operator based on the communication quality between the operator and the external system and the maintenance machine. The second processing module also corrects the health value based on the alarm and warning information from the operator.

[0010] Optionally, the host computer status monitoring module may further include a communication interface submodule; The first processing module and the second processing module each include a first monitoring and communication submodule and a second monitoring and communication submodule, respectively. Through the communication interface submodule and the first monitoring communication submodule of the repair machine, data interaction between the first processing module and the corresponding host computer status monitoring module is realized inside the repair machine; Through the communication interface submodule and the second monitoring communication submodule of the operator machine, data interaction between the second processing module and the corresponding host computer status monitoring module is realized inside the operator machine; Data interaction between the operator and the maintenance machine is achieved through the communication interface submodule and the first monitoring communication submodule of the operator.

[0011] Optionally, the operating status information includes hardware operating status information; the interlocking host computer is configured with a monitoring base library provided by the interlocking host computer manufacturer, and the status monitoring submodule obtains the hardware operating status information through the monitoring base library; The hardware operating status information includes: static information and dynamic information; The static information includes: CPU model, CPU minimum frequency, CPU maximum frequency, number of physical CPU cores, number of logical CPU cores, number of memory modules, memory module model name, total physical memory capacity, total virtual memory capacity, number of hard drives, logical hard drive name, hard drive model name, total hard drive capacity, number of system fans, number of power supplies; The dynamic information includes: current CPU frequency, CPU utilization, CPU temperature, CPU fan speed, available physical memory capacity, used physical memory capacity, physical memory utilization, available virtual memory capacity, used virtual memory capacity, virtual memory utilization, used hard disk capacity, remaining hard disk capacity, hard disk utilization, system temperature, system fan speed, and power status.

[0012] Optionally, the running status information also includes software running status information; the status monitoring submodule obtains the software running status information through the operating system API of the interlocking host computer; The software running status information includes: process information and service information; The process information includes: process name, process identifier, process status, username, CPU utilization, number of threads, startup command line, and CPU time statistics; The service information includes: service name, associated process ID, and service status.

[0013] Optionally, the alarm warning information includes: CPU usage alarms, memory usage alarms, hard disk space usage alarms, device temperature alarms, and power supply alarms.

[0014] Optionally, the host computer status monitoring module further includes an interface display submodule, used to visually display the operating status information and the alarm warning information, as well as the model of the interlocking host computer, the version number of the host computer status monitoring module, the version number of the monitoring base library, the operating system type value, the communication status between different interlocking host computers, the communication status between the first monitoring communication submodule and the corresponding communication interface submodule, and the communication status between the second monitoring communication submodule and the corresponding communication interface submodule.

[0015] Optionally, the host computer status monitoring module further includes an information storage submodule for storing the operating status information, the alarm and early warning information, the generation time of the alarm and early warning information, and the alarm and early warning status; the alarm and early warning status includes: alarm and early warning occurrence and alarm and early warning recovery.

[0016] Optionally, the interlocking host computer monitoring system further includes an offline analysis device; when maintaining the interlocking host computer, the staff copies the operating status information and the alarm warning information in the information storage submodule to the offline analysis device; The offline analysis device includes a trained neural network model, which is used to generate at least one of the following based on the model of the interlocking host computer, the continuous running time, the running status information, and the alarm warning information: the recommended service life and the optimal restart interval of the corresponding interlocking host computer. The neural network model also updates its own model parameters based on the operating status information and alarm warning information within the most recent preset time period.

[0017] Optionally, the first processing module and the second processing module each include a first interface display unit and a second interface display unit, used to visually display the operating status information and the alarm warning information.

[0018] The present invention also provides an interlocking host computer monitoring method for use in the interlocking host computer monitoring system as described in the present invention, comprising the following steps: S1. Deploy a host computer status monitoring module on each interlock host computer. The host computer status monitoring module starts automatically when the host computer of the interlock is started. S2. The host computer status monitoring module collects the operating status information of the host computer in the interlock and generates corresponding alarm and early warning information; S3. The first processing module diagnoses the cause of the fault and recommends the corresponding fault handling strategy for the operator and the maintenance machine based on the operating status information and alarm warning information of the operator and the maintenance machine, respectively.

[0019] Optionally, the number of operating units is multiple, and each operating unit is communicatively connected to the maintenance unit via a dual network. Each operating unit further includes a second processing module, and the method further includes the following steps: S4. The second processing module calculates a health value for the operator based on the communication quality between the operator and the external system and the maintenance machine; and corrects the health value based on the alarm warning information of the operator; and selects the operator with the highest health value as the master operator.

[0020] Optionally, the interlocking host computer monitoring system further includes an offline analysis device, which includes a trained neural network model, and the method further includes the following steps: S5. When maintaining the interlocking host computer, copy the operating status information and the alarm warning information from the information storage submodule of the host computer status monitoring module to the offline analysis device; The neural network model generates at least one of the following for the interlocking host computer: suggested service life and optimal restart interval, based on the model of the interlocking host computer, continuous running time, operating status information, and alarm warning information. S6. The neural network model updates its own model parameters based on the operating status information and alarm warning information within the most recent preset time period.

[0021] Compared with the prior art, the beneficial effects of the present invention include: 1) This invention can effectively shorten the time required for fault detection and identification of the interlocking host computer, facilitating timely maintenance by operation and maintenance personnel. In particular, it can provide early warning and timely alarm for various faults caused by long-term aging of the interlocking host computer, greatly reducing the occurrence of sudden faults.

[0022] 2) When the interlocking host computer malfunctions, the monitoring data collected by this invention (including operating status information and early warning alarm information) can be used to accurately locate the cause of the malfunction and recommend targeted fault handling strategies, significantly improving the efficiency of fault repair and ensuring the overall stable and reliable operation of the computer interlocking system. It also avoids various operational risks caused by blindly restarting the interlocking host computer. By greatly reducing the number of restarts of the interlocking host computer, the probability of memory failure and hard drive bad sectors is effectively reduced, significantly extending the service life of the interlocking host computer. This invention also prevents frequent restarts of the interlocking host computer from masking the root cause of the fault and prevents the fault from spreading within the computer interlocking system.

[0023] 3) This invention improves the core functions of the existing computer interlocking system. The second processing module calculates accurate health values ​​based on alarm and early warning information and optimizes the judgment logic of the main operator, thereby effectively improving the overall operation and control performance of the computer interlocking system.

[0024] 4) Based on long-term accumulated alarm and early warning information and operational status data, this invention uses an offline processing device to accurately assess the overall quality and actual operating conditions of each interlocking host computer, thereby outputting recommended service life and optimal restart intervals. This data support can effectively guide equipment operation and maintenance management strategies, thus significantly extending the service life of the interlocking host computers. Attached Figure Description

[0025] Figure 1 This is a schematic diagram of the interlocking host computer monitoring system in an embodiment of the present invention.

[0026] Figure 2 This is a flowchart of the interlocking host computer monitoring method in an embodiment of the present invention. Detailed Implementation

[0027] The interlocking host computer monitoring system and method proposed in this invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. The advantages and features of this invention will become clearer from the following description. It should be noted that the drawings are in a very simplified form and use non-precise proportions, only for the purpose of conveniently and clearly illustrating the embodiments of this invention. Please refer to the drawings to make the objectives, features, and advantages of this invention more apparent and understandable. It should be understood that the structures, proportions, sizes, etc., depicted in the accompanying drawings are only for the purpose of assisting those skilled in the art in understanding and reading the content disclosed in the specification, and are not intended to limit the implementation conditions of this invention. Therefore, they have no substantial technical significance. Any modifications to the structure, changes in the proportional relationships, or adjustments to the size, without affecting the effects and objectives achieved by this invention, should still fall within the scope of the technical content disclosed in this invention.

[0028] Figure 1 This is a schematic diagram of the host computer monitoring system for the first interlocking processing module 110 in an embodiment of the present invention. Figure 1 As shown, the interlocking host computer monitoring system includes multiple host computer status monitoring modules 300, which are deployed on multiple interlocking host computers 100 (also called industrial control computers) of the computer interlocking system. The host computer status monitoring module 300 automatically starts when the interlocking host computer 100 it is on starts.

[0029] Multiple interlocking host computers 100 include maintenance computers and operator computers. In this embodiment, the operator computer and the maintenance computer communicate bidirectionally via a dual-network redundancy mechanism based on pre-divided operator computer network segments and maintenance computer network segments.

[0030] There are usually multiple manipulators, one of which is the master manipulator and the rest are backup manipulators. Figure 1 The diagram shows two operator units, one as the primary operator and the other as the backup operator, which can automatically switch between each other. For example, a switching device (not shown in the diagram) can decide and control the switching between the primary and backup operator units based on the operator unit's health value (described in detail later). How the switching between the primary and backup operator units is performed is prior art and is not the focus of this invention, so it will not be elaborated upon here.

[0031] In this embodiment, as Figure 1 As shown, the host computer status monitoring module 300 includes: a status monitoring submodule 301, an alarm and early warning submodule 302, and a communication interface submodule 303.

[0032] The status monitoring submodule 301 is used to periodically collect the operating status information of its host interlocking computer 100. In this embodiment, the host interlocking computer 100 is equipped with a monitoring basic library provided by the host computer manufacturer (such as Advantech, Vtronix, KENGCHENG, etc.), and the status monitoring submodule 301 obtains the operating status information through the monitoring basic library. Through compatibility design, the monitoring basic library can be adapted to all models of industrial control computer hardware. The monitoring basic library is provided in the form of a dynamic library and is also differentiated according to the operating system (such as Windows and Linux systems) to support cross-operating system platform use.

[0033] In this embodiment, the operating status information includes hardware operating status information and software operating status information. The hardware operating status information includes static information and dynamic information.

[0034] In this embodiment, the status monitoring submodule reads the static information every preset period (e.g., 10 seconds, this is only an example and not a limitation of the present invention). The static information rarely changes and includes: CPU information (CPU model, CPU minimum frequency, CPU maximum frequency, number of CPU physical cores, number of CPU logical cores), memory information (number of memory modules, memory module model name, total physical memory capacity, total virtual memory capacity), hard disk information (number of hard disks, logical hard disk name, hard disk model name, total hard disk capacity), and fan power information (number of system fans, number of power supplies).

[0035] In this embodiment, the status monitoring submodule reads the dynamic information every preset period (e.g., 200 milliseconds, this is only an example and not a limitation of the present invention). The dynamic information is prone to change and includes: CPU information (current CPU frequency, CPU utilization, CPU temperature, CPU fan speed), memory information (available physical memory capacity, used physical memory capacity, physical memory utilization, available virtual memory capacity, used virtual memory capacity, virtual memory utilization), hard disk information (used hard disk capacity, remaining hard disk capacity, hard disk utilization), and fan power information (system temperature, system fan speed, power status).

[0036] Table 1 shows a list of monitoring basic library interfaces from Advantech in one embodiment of the present invention. If it is from another manufacturer, the prefix needs to be replaced. For example, in the list of basic library interfaces from VTech, the prefix "Adv" should be replaced with "Iei".

[0037] Table 1

[0038] In this embodiment, the status monitoring submodule 301 obtains the software running status information every preset period (e.g., 500 milliseconds, this is only an example and not a limitation of the present invention) through the operating system API (Application Programming Interface) of the interlocking host computer 100. In this embodiment, the software running status information includes: process information and service information.

[0039] The process information includes: process name, process identifier, process status, username, CPU utilization, number of threads, startup command line, and CPU time statistics. For Windows systems, the process information also includes memory commit size, number of handles, number of user objects, and number of GDI objects. For Linux systems (such as Kylin OS), the process information also includes physical memory, shared memory, and virtual memory values.

[0040] The service information includes: service name, associated process ID, and service status.

[0041] The aforementioned dynamic information, static information, process information, and service information are all transmitted from the status monitoring submodule 301 to other modules of the host computer status monitoring module 300 via an in-process message queue.

[0042] In this embodiment, when the host computer status monitoring module 300 starts up, it reads basic information through the monitoring base library, including the interlock host computer model, the monitoring base library version number, the host computer status monitoring module 300 version number (read through the host computer status monitoring module 300 binary information) and the operating system type value (identified through the operating system API).

[0043] The alarm and early warning submodule 302 is configured to generate corresponding alarm and early warning information based on the operating status information.

[0044] In this embodiment, the alarm and warning information includes: CPU usage alarm and warning information, memory usage alarm and warning information, hard disk space usage alarm and warning information, device temperature alarm and warning information, and power supply alarm and warning information. The following are alarm and warning items in one embodiment of the present invention; specific additions and deletions can be made according to actual needs.

[0045] High CPU usage alarm: A single instance of high CPU usage (e.g., 80%, configurable based on actual needs, this is just an example) is considered excessive. Within a set continuous monitoring window (e.g., 30 seconds), the number of times CPU usage is excessive is counted, and the percentage of such instances relative to the total number of monitoring instances is calculated. If this percentage exceeds a preset alarm ratio threshold (e.g., 50%), a CPU usage alarm is triggered. If, after an alarm occurs, CPU usage is detected to be below the first preset threshold, the alarm is considered to have been resolved.

[0046] High CPU Usage Warning: A single instance of high CPU usage (below the first preset threshold, e.g., 60%, configurable based on actual needs) is considered a high usage event. Within a set continuous monitoring window (e.g., 30 seconds), the number of times high CPU usage occurs is counted, and the percentage of such occurrences relative to the total monitoring events is calculated. If this percentage exceeds a preset warning ratio threshold (e.g., 40%), a CPU usage warning is triggered. If, after a warning is triggered, CPU usage falls below the second preset threshold, the warning is considered to have been lifted.

[0047] High memory usage alarm: If the total memory usage in a single monitoring session reaches the third preset threshold (e.g., 60%, configurable according to actual needs), it is considered excessive and a high memory usage alarm is triggered. If the total memory usage falls below the third preset threshold after the alarm occurs, the alarm is considered to have been resolved.

[0048] High memory usage alert: A high memory usage alert is triggered when the total memory usage reaches the fourth preset threshold (below the third preset threshold, such as 40%, which can be configured according to actual needs). If the total memory usage falls below the fourth preset threshold after the alert is triggered, the alert is considered to have been lifted.

[0049] Low disk space alarm: If the total disk usage reaches the fifth preset threshold (e.g., 70%, configurable according to actual needs), a low disk space alarm is triggered. If the total disk usage falls below the fifth preset threshold after the alarm occurs, the alarm is considered to have been resolved.

[0050] Hard Disk Space Shortage Warning: If the total hard disk usage reaches the sixth preset threshold (below the fifth preset threshold, such as 50%, which can be configured according to actual needs), a hard disk space shortage warning is triggered. If the total hard disk usage falls below the sixth preset threshold after the warning is triggered, the warning is considered to have been lifted.

[0051] Equipment over-temperature alarm: If the monitored system temperature reaches the seventh preset threshold (e.g., 80 degrees Celsius, configurable according to actual needs), it is considered over-temperature and an equipment over-temperature alarm is triggered. If the monitored system temperature falls below the seventh preset threshold after the alarm occurs, the alarm is considered to have been resolved.

[0052] High Equipment Temperature Warning: A high temperature warning is triggered when the monitored system temperature reaches the eighth preset threshold (below the seventh preset threshold, e.g., 60 degrees Celsius, configurable according to actual needs). Alternatively, a high temperature warning is triggered if the average speed of all system fans (fans installed in the chassis airflow to cool the entire chassis) exceeds a specified value (e.g., 4000 RPM, configurable according to actual needs) within a continuous time period (e.g., 10 seconds). If the average speed of the system fans falls below the specified value after the warning is triggered, the warning is considered to have been lifted.

[0053] Power Failure Alarm: If the monitored power status value is a preset alarm value (configurable according to actual needs), a power failure is considered to be triggered, and a power alarm is activated. If the monitored power status value is normal after the alarm occurs, the alarm is considered to have been resolved.

[0054] Power Abnormality Warning: If the monitored power status value is neither normal nor at the preset alarm value, it is considered abnormal and a power warning is triggered. If the power status value returns to normal after the warning is triggered, the warning is considered to have been lifted.

[0055] In this embodiment, as Figure 1 As shown, the maintenance machine includes a first processing module 110, which includes a first monitoring and communication submodule 111. Through the communication interface submodule 303 of the maintenance machine and the first monitoring and communication submodule 111, data interaction between the first processing module 110 and the corresponding host computer status monitoring module 300 is realized within the maintenance machine (e.g., using named pipes).

[0056] In this embodiment, as Figure 1 As shown, the operator unit includes a second processing module 120, which includes a second monitoring and communication submodule 121. Data interaction (e.g., using named pipes) between the second processing module 120 and the corresponding host computer status monitoring module 300 is achieved within the operator unit via the operator unit's communication interface submodule 303 and the second monitoring and communication submodule 121. Data interaction (e.g., data transmission based on UDP protocol) between the operator unit and the maintenance unit is achieved via the operator unit's communication interface submodule 303 and the maintenance unit's first monitoring and communication submodule 111.

[0057] In other words, the communication interface submodule 303 of the repair machine is only used for communication within the repair machine, while the communication interface submodule 303 of the operator machine is used not only for communication within the operator machine, but also for communication between the operator machine and the repair machine.

[0058] In this embodiment, the first processing module 110 is configured to diagnose the cause of the fault and recommend corresponding fault handling strategies for the operator and the repair machine respectively, based on the operating status information and alarm warning information of the operator and the operating status information and alarm warning information of the repair machine.

[0059] In existing technologies, data packets are captured through a dual-network communication system established between the maintenance unit and the operator unit to determine whether communication faults have occurred between the two devices. External maintenance methods such as unplugging and plugging in network ports and checking switching equipment are used for network fault handling. However, these methods cannot effectively identify and locate hardware and operational faults within the maintenance unit or operator unit itself. Monitoring the operation of the maintenance unit and operator unit using a watchdog mechanism only allows for the continuous restarting of the maintenance unit and operator unit, which can lead to a wider range of faults and interference with interlocking control and station displays.

[0060] In this application, in addition to judging dual-network communication failures, the first processing module 110 can also judge whether the operating unit and the maintenance unit have experienced hardware failures, process abnormalities, or other operating conditions based on the operating status information and alarm warning information of the operating unit and the maintenance unit, thereby helping maintenance personnel to take corresponding fault handling strategies, including: "Please check the fan", "Please check the hard disk space", "Please check the CPU usage of the process", and "Please check the system temperature".

[0061] In this embodiment, the first processing module 110 is internally configured with a trained neural network model (e.g., a Transformer network model). The input data for this model includes existing software log information and network packet capture data, as well as the operating status information and alarm warning information of the interlocking host computer 100. The output data of this neural network model is the fault diagnosis result and the fault handling strategy. How to train this neural network model is existing technology and will not be elaborated upon in this invention.

[0062] This invention can effectively shorten the time required for fault detection and identification of the interlocking host computer 100, facilitating timely maintenance by operation and maintenance personnel. In particular, it can provide early warning and timely alarm for various faults caused by long-term aging of the interlocking host computer 100, greatly reducing the occurrence of sudden faults.

[0063] When the interlocking host computer 100 malfunctions, the operating status information and early warning alarm information collected by this invention can be used to accurately locate the cause of the malfunction and recommend targeted fault handling strategies, significantly improving the efficiency of fault repair and ensuring the overall stable and reliable operation of the computer interlocking system. It also avoids various operational risks caused by blindly restarting the interlocking host computer 100. By greatly reducing the number of restarts of the interlocking host computer 100, the probability of memory failure and hard drive bad sectors is effectively reduced, significantly extending the service life of the interlocking host computer 100. This invention also prevents frequent restarts of the interlocking host computer 100 from masking the root cause of the malfunction and prevents the malfunction from spreading within the computer interlocking system.

[0064] like Figure 1 As shown, in this embodiment, the first processing module 110 further includes a first interface display unit 112, used to render and visualize the first processing interface. This first processing interface includes: a device sub-monitoring page, an alarm diagnosis sub-page, and an information query sub-page. The device monitoring sub-page displays real-time monitoring information (operating status information), and the alarm diagnosis sub-page displays alarm warning information, fault diagnosis results, and corresponding fault handling strategies. The information query sub-page allows users to query historical monitoring information previously reported by the host computer status monitoring module 300.

[0065] In this embodiment, the second processing module 120 of the operator is configured to calculate a health value for the operator based on the communication quality between the operator and external systems and the maintenance unit. In this embodiment, different initial health values ​​X and Y are first configured for the primary operator and the backup operator, respectively, where X is greater than Y. If the communication status between the operator and the maintenance unit is single-network normal, the health value of the operator is subtracted by 'a', where 'a' is a first preset value. If the communication status between the operator and the maintenance unit changes from single-network normal to dual-network disconnected, the health value of the operator is further subtracted by 'b'. If the communication status between the operator and the maintenance unit changes from dual-network disconnected to single-network normal, the health value of the operator is added by 'b'. If the communication status between the operator and the maintenance unit changes from single-network normal to dual-network normal, the health value of the operator is further added by 'a'. When the communication between the operator and an external system is disconnected, the health value of the operator is subtracted by 'c', where 'c' is a third preset value. If the communication with the external system returns to normal, the health value of the operator is added by 'c'. Health values ​​can also be calculated in other ways, and this invention does not impose limitations.

[0066] In the present invention, the second processing module 120 further corrects the health value based on the alarm and early warning information of the operating machine. For example, when the upper computer status monitoring module 300 generates an alarm information of excessively high CPU usage, the health value of the operating machine is decreased by d, where d is a fourth preset value; if the alarm of excessively high CPU usage is recovered, the health value of the operating machine is increased by d. When the upper computer status monitoring module 300 generates early warning information of excessively high CPU usage, the health value of the operating machine is decreased by e, where e is a fifth preset value, e<d; if the early warning of excessively high CPU usage is recovered, the health value of the operating machine is increased by e. An operating machine with a high health value is used as the main operating machine.

[0067] The present invention improves the core function of the existing computer interlocking system, the second processing module 120 calculates an accurate health value according to the alarm and early warning information, and optimizes the determination logic of the main operating machine, thereby effectively improving the overall operation control performance of the system.

[0068] As Figure 1 shown, in this embodiment, the second processing module 120 further includes a second interface display unit 122, configured to render and visually display a second processing interface. The second processing interface includes: an equipment monitoring submenu, an alarm window and an early warning window. The equipment monitoring submenu is configured to display all real-time and historical monitoring information. Alarm information generated by the upper computer status monitoring module 300 of the operating machine is displayed through the alarm window, and early warning information generated by the upper computer status monitoring module 300 of the operating machine is displayed through the early warning window.

[0069] In this embodiment, data interaction inside the interlocking upper computer 100 and between interlocking upper computers 100 all adopt the same application-layer data type and message format. Each message packet is divided into two parts: a message header and a message body. All fields use big-endian byte order.

[0070] In this embodiment, the format of the message header is shown in Table 2 below: Table 2

[0071] The receiver needs to save the "message sequence number" field, and after receiving the next packet of data, compares its sequence number with the most recently saved sequence number. If it is less than the saved sequence number, it is regarded as an expired packet and discarded. If it is equal to the saved sequence number, it is regarded as a redundant packet and discarded. If it is greater than the saved sequence number, it is regarded as a valid packet.

[0072] The message body shall contain one or more sub-packets. Each sub-packet is one of the following message types: 1) Communication connection request message (its format is shown in Table 3). Communication direction: unidirectionally sent from the first processing module 110 / second processing module 120 to the host computer status monitoring module 300. After the first processing module 110 / second processing module 120 starts, it sends this message to the host computer status monitoring module 300 that requests communication. If the communication fails to establish a connection, it will be resent after a certain interval (e.g., 10 seconds).

[0073] Table 3

[0074] 2) Communication connection reply message (its format is shown in Table 4). Communication direction: The message is sent unidirectionally from the host computer status monitoring module 300 to the first processing module 110 / second processing module 120. After receiving the "communication connection request message," the host computer status monitoring module 300 replies to the first processing module 110 / second processing module 120 with this message, indicating that communication has been successfully established.

[0075] Table 4

[0076] 3) Communication connection heartbeat message (its format is shown in Table 5). Communication direction: Bidirectional transmission between the host computer status monitoring module 300 and the first processing module 110 / second processing module 120. After the communication between the host computer status monitoring module 300 and the first processing module 110 / second processing module 120 is successfully established, if the sender does not send any other messages within a certain time T1 (e.g., 3 seconds), this message should be sent to maintain the communication connection. If no valid message is received within a certain time T2 (which must be greater than T1, e.g., 10 seconds), the communication connection should be considered to have been broken.

[0077] Table 5

[0078] 4) Industrial control computer monitors static messages (the format of which is shown in Table 6). Communication direction: unidirectional transmission from the host computer status monitoring module 300 to the first processing module 110 / second processing module 120. After monitoring the status of the interlocking host computer 100, the host computer status monitoring module 300 packages all current static information (including basic information, CPU information, memory information, hard disk information, and fan power information) into a JSON string and sends it. If the information content changes, it is sent immediately; if there is no change, it is resent after a certain interval (e.g., 120 seconds).

[0079] Table 6

[0080] The following illustrates, in one embodiment, the specific objects and properties of the JSON string of static messages monitored by the industrial control computer. Unless otherwise specified, all property values ​​are of string type. For properties with numerical values, negative numbers represent unknown values.

[0081] ① The basic information object is named Basic and contains the following attributes: Model: Industrial PC model LibVersion: Base library version number SoftVersion: Software version number OSType: Operating system type value.

[0082] ②The CPU information object is named CPU and contains the following attributes: Model: Model MinFreq: Minimum frequency (rounded to two decimal places, e.g., "3.00", unit: GHz) MaxFreq: Maximum frequency (rounded to two decimal places, e.g., "3.00", unit: GHz) PhyCores: Number of physical cores LogicCores: Number of logic cores.

[0083] ③ The memory information object is named Memory and contains the following attributes: Count: Number of entries Model: Model number (values ​​are string arrays) PhyTotal: Total physical memory (rounded to two decimal places, e.g., "3.00", unit: MB) SwapTotal: Total virtual memory (retain two decimal places, such as "3.00", unit: MB).

[0084] ④ The hard disk information object is named Disk and contains the following attributes: Count: quantity Model: Model number (values ​​are string arrays) Name: Name (value is an array of strings) Total: Total capacity (rounded to two decimal places, e.g., "3.00", unit GB).

[0085] ⑤ The fan power information object is named System and contains the following properties: SysFanCount: Number of system fans PowerCount: Number of power supplies.

[0086] 4) Industrial PC monitoring dynamic messages (the format is shown in Table 7). Communication direction: unidirectional transmission from the host computer status monitoring module 300 to the first processing module 110 / second processing module 120. After detecting the status of the interlocked host computer 100, the host computer status monitoring module 300 packages the current dynamic information into a JSON string and sends it. The information content is divided into 5 subcategories: CPU information, memory information, hard disk information, fan power information, and software running status. When the information content of each subcategory changes, the corresponding industrial PC monitoring dynamic message is sent immediately. If there is no change, the industrial PC monitoring dynamic message is resent after a certain interval (e.g., 10 seconds).

[0087] Table 7

[0088] The JSON string contains part or all of the current dynamic information (CPU information, memory information, hard disk information, fan power information, and software running status that are prone to change).

[0089] The following example illustrates the specific objects and attributes of the JSON string used by the industrial control computer to monitor dynamic messages. Unless otherwise specified, all attribute values ​​are of string type. For attributes with numerical values, negative numbers represent unknown values.

[0090] ①The CPU information object is named CPU and contains the following attributes: Usage: Average usage rate (rounded to two decimal places, e.g., "3.00", unit: %) Temp: Temperature (rounded to two decimal places, e.g., "3.00", unit: degrees Celsius) CurFreq: Current frequency (retain two decimal places, such as "3.00", unit: GHz).

[0091] ② The memory information object is named Memory and contains the following attributes: PhyUsed: The amount of physical memory used (rounded to two decimal places, e.g., "3.00", in MB). PhyAvailable: Amount of available physical memory (rounded to two decimal places, e.g., "3.00", in MB). SwapUsed: The amount of virtual memory used (rounded to two decimal places, e.g., "3.00", in MB). SwapAvailable: The amount of virtual memory available (rounded to two decimal places, such as "3.00", in MB).

[0092] ③ The hard disk information object is named Disk and contains the following attributes: Used: Number used (rounded to two decimal places, e.g., "3.00", unit GB) Available: Number of available items (rounded to two decimal places, e.g., "3.00", unit GB).

[0093] ④ The fan power information object is named System and contains the following properties: CPUFan: CPU fan speed (retain two decimal places, such as "3.00", unit RPM) SysFan: Fan (value is an array of strings, each string is rounded to two decimal places, such as "3.00", unit is RPM) Temp: Temperature (rounded to two decimal places, e.g., "3.00", unit: degrees Celsius) PowerStatus: Power status (values ​​are a string array, each string is represented in hexadecimal, such as "0x12AB". 0 indicates normal, 0xFFFFFFFF indicates unknown, and all others indicate a fault).

[0094] ⑤ The software runtime state object is named Software and contains the following attributes: Process: Process information block. The value is an array of sub-objects, each of which should contain the following attributes: ID: Process ID Status: Process status User: Username CPUUsedPercent: CPU utilization rate ThreadCount: Number of threads StartCmd: Launch command line CPUTime: CPU time statistics.

[0095] If it is a Windows system, the following attributes should also be included: MemUsed: The size of the memory committed (rounded to two decimal places, e.g., "3.00", in MB). HandleCount: Number of handles UserObjCount: Number of user objects GDIObjCount: Number of GDI objects If it is a Linux system, the following attributes should also be included: MemPhy: Physical memory values MemShare: Shared memory value MemSwap: Virtual memory value.

[0096] Service: Service information block. The value is an array of sub-objects, each of which should contain the following properties: Name: Service Name ProcessID: The associated process ID Status: Service status.

[0097] like Figure 1 As shown, in this embodiment, the host computer status monitoring module 300 further includes an interface display submodule 304. The interface display submodule 304 is used to visually display the operating status information and the alarm warning information, as well as the model of the interlocking host computer 100, the version number of the host computer status monitoring module 300, the version number of the monitoring base library, the operating system type value, the communication status between different interlocking host computers 100, and the communication status between the first processing module 110 / second processing module 120 and the corresponding host computer status monitoring module 300.

[0098] In this embodiment, a cross-platform graphical framework is used to render all monitoring information in real time onto the interface of the host computer status monitoring module 300. The main interface is displayed in the form of a paginated dialog box, with each page implemented as a dialog box. Based on information relevance, the pages are categorized into seven types: CPU information page, memory information page, hard disk information page, fan power information page, software running status page, alarm and warning page, and other information page.

[0099] In this embodiment, the CPU information page uses a circular progress bar component to display the current real-time CPU utilization, a line graph component to display the CPU utilization in the most recent minute, and other information is displayed using a single-line text box component.

[0100] In this embodiment, the memory information page uses a circular progress bar component to display the real-time physical memory usage, virtual memory usage, and total memory usage. A line graph component displays the real-time physical memory usage, virtual memory usage, and total memory usage over the past 5 minutes. Other information is displayed using single-line text boxes.

[0101] In this embodiment, the hard drive information page uses a circular progress bar component to display the total real-time utilization rate of all hard drives, and a line graph component to display the total utilization rate of all hard drives in the last 10 minutes. Other information is displayed using a single-line text box component.

[0102] In this embodiment, the fan power information page uses a multi-line text box component to display the meaning of the power status, while other information is displayed using a single-line text box component.

[0103] In this embodiment, the software running status page uses a table component to display the current real-time process information block and service information block.

[0104] In this embodiment, the alarm and warning page uses a table component to display alarm and warning information for a recent period. By default, it displays information from the past week; users can adjust the display to show alarms and warnings for other time periods using the time display component.

[0105] In this embodiment, the other information page displays the model of the interlocking host computer 100, the version number of the monitoring base library, the version information of the host computer status monitoring module 300, the operating system type value, the communication status within the interlocking host computer 100, and the communication status between the interlocking host computers 100.

[0106] like Figure 1 As shown in this embodiment, the host computer status monitoring module 300 further includes an information storage submodule 305. The information storage submodule 305 is used to store operating status information, alarm and warning information, the generation time of the alarm and warning information, and the alarm and warning status. The alarm and warning status includes: alarm and warning occurrence and alarm and warning recovery.

[0107] In this embodiment, the information storage submodule 305 uses a database to store all monitoring information (i.e., operating status information) and alarm / early warning information. All monitoring information is categorized and stored in one static information table and multiple dynamic information tables. Alarm / early warning information is stored separately in an alarm / early warning information table.

[0108] The static information table records static information of the interlocking host computer 100 that rarely changes, including time, CPU information (model, minimum frequency, maximum frequency, number of physical cores, number of logical cores), memory information (quantity, model name, total physical capacity, total virtual capacity), hard disk information (quantity, logical name, model name, total capacity), and fan power information (number of fans, number of power supplies). When the static information changes, a record is generated in the static information table. If there is no change, a record is generated only every hour (this is only an example and not a limitation of the invention).

[0109] The dynamic information table records information that is prone to change in the industrial control computer. Because different information is sampled at different frequencies, it can be divided into: a CPU dynamic table (including time, current frequency, instantaneous utilization rate, temperature, and fan speed); a memory dynamic table (including time, physical available capacity, physical used capacity, virtual available capacity, and virtual used capacity); a hard disk dynamic table (including time, used capacity, and remaining capacity); a fan power dynamic table (including time, system temperature, system fan speed, and power status); and a software running status table (including time, process information blocks, and service information blocks). When information changes, a record is generated in the dynamic information table. If there is no change, a record is generated every 60 seconds (this is only an example and not a limitation of the invention).

[0110] The alarm and early warning information table records the alarm and early warning information generated by the alarm and early warning submodule 302, including time, level (alarm or early warning), content of the alarm and early warning information, and status (occurred or recovered).

[0111] In this embodiment, the information storage submodule 305 monitors the database file size every minute. When it exceeds a predetermined value (e.g., 500M), it first adds time information to the filename, then backs up and compresses the file, and then creates a new database file for storage. The information storage submodule 305 also monitors the number of all backup files every minute. If it exceeds a predetermined value (e.g., 20), it deletes the backup files with earlier dates to ensure that the interlocking host computer 100 will not malfunction due to a full disk.

[0112] In this embodiment, as Figure 1 As shown, the interlocking host computer monitoring system also includes an offline analysis device 500. When maintaining the interlocking host computer 100, staff copy the operating status information and alarm warning information from the information storage submodule 305 to the offline analysis device 500.

[0113] The offline analysis device 500 includes a trained neural network model (e.g., a Transformer model; how to train this model is existing technology and will not be elaborated here). This neural network model is configured to use the model of the interlocking host computer 100, its continuous operating time, operating status information, and alarm warning information as multi-dimensional input features to predict at least one of the recommended service life and optimal restart interval of the corresponding interlocking host computer 100. Based on the prediction results, the device can effectively guide the operation and maintenance management strategy, formulate targeted equipment maintenance cycles, reduce the failure risk caused by equipment aging, and thus significantly extend the service life of the interlocking host computer.

[0114] The neural network model in the offline analysis device 500 also updates its own model parameters based on the operating status information and alarm warning information within the most recent preset time period.

[0115] This invention also provides an interlocking host computer monitoring method for use in the interlocking host computer monitoring system as described in this invention. Figure 2 As shown, the interlocking host computer monitoring method includes the following steps: S1. Deploy a host computer status monitoring module on each interlock host computer. The host computer status monitoring module will start automatically when the host computer of the interlock is started.

[0116] S2. The host computer status monitoring module collects the operating status information of the host computer of the interlock and generates corresponding alarm and early warning information.

[0117] S3. The first processing module diagnoses the cause of the fault and recommends the corresponding fault handling strategy for the operator and the maintenance machine based on the operating status information and alarm warning information of the operator and the maintenance machine, respectively.

[0118] In this embodiment, the operator unit communicates with the maintenance unit via dual networks, and there are multiple operator units. For example... Figure 2 As shown, the method further includes the following steps: S4. The second processing module calculates the health value of the operator based on the communication quality between the operator and the external system and the maintenance unit; and corrects the health value based on the alarm and early warning information of the operator; and selects the operator with the highest health value as the master operator.

[0119] In this embodiment, the interlocking host computer system further includes an offline analysis device, which includes a trained neural network model, such as... Figure 2 As shown, the method further includes the following steps: S5. When maintaining the interlocking host computer, the operating status information and the alarm warning information are copied from the information storage submodule of the host computer status monitoring module to the offline analysis device; the neural network model in the offline analysis device generates at least one of the following for the interlocking host computer: recommended service life and optimal restart interval, based on the model, continuous running time, operating status information and alarm warning information.

[0120] S6. The neural network model in the offline analysis device updates its own model parameters based on the operating status information and alarm warning information within the most recent preset time period.

[0121] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0122] In the description of this invention, it should be understood that the terms "center," "height," "thickness," "upper," "lower," "vertical," "horizontal," "top," "bottom," "inner," "outer," "axial," "radial," and "circumferential," etc., indicating orientation or positional relationships, are based on the orientation or positional relationships shown in the accompanying drawings and are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. In the description of this invention, unless otherwise stated, "a plurality of" means two or more.

[0123] In the description of this invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," "linking," and "fixing" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.

[0124] In this invention, unless otherwise explicitly specified and limited, "above" or "below" the second feature can include direct contact between the first and second features, or contact between the first and second features through another feature between them. Furthermore, "above," "over," and "on top" of the second feature includes the first feature directly above or diagonally above the second feature, or simply indicates that the first feature is at a higher horizontal level than the second feature. "Below," "below," and "under" the second feature includes the first feature directly below or diagonally below the second feature, or simply indicates that the first feature is at a lower horizontal level than the second feature.

[0125] Although the present invention has been described in detail through the preferred embodiments above, it should be understood that the above description should not be considered as a limitation of the present invention. Various modifications and substitutions to the present invention will be apparent to those skilled in the art after reading the above description. Therefore, the scope of protection of the present invention should be defined by the appended claims.

Claims

1. A computer-based interlocking monitoring system, comprising multiple interlocking host computers, wherein each interlocking host computer includes a maintenance computer and an operator computer, the maintenance computer and the operator computer being communicatively connected, characterized in that... Include: Multiple host computer status monitoring modules are deployed on multiple interlocking host computers; each host computer status monitoring module starts automatically when the host computer it belongs to starts. The host computer status monitoring module includes: The status monitoring submodule is used to periodically collect the operating status information of the interlocking host computer to which it belongs; An alarm and early warning submodule is configured to generate corresponding alarm and early warning information based on the operating status information; The repair machine includes a first processing module, which is configured to diagnose the cause of the fault and recommend corresponding fault handling strategies for the operating machine and the repair machine respectively, based on the operating status information and alarm warning information of the operating machine and the operating status information and alarm warning information of the repair machine.

2. The interlocking host computer monitoring system as described in claim 1, characterized in that, There are multiple operating machines, and each operating machine is connected to the maintenance machine via a dual network. The operator includes a second processing module, which is configured to calculate a health value for the operator based on the communication quality between the operator and the external system and the maintenance machine. The second processing module also corrects the health value based on the alarm and warning information from the operator.

3. The interlocking host computer monitoring system as described in claim 2, characterized in that, The host computer status monitoring module also includes a communication interface sub-module; The first processing module and the second processing module each include a first monitoring and communication submodule and a second monitoring and communication submodule, respectively. Through the communication interface submodule and the first monitoring communication submodule of the repair machine, data interaction between the first processing module and the corresponding host computer status monitoring module is realized inside the repair machine; Through the communication interface submodule and the second monitoring communication submodule of the operator machine, data interaction between the second processing module and the corresponding host computer status monitoring module is realized inside the operator machine; Data interaction between the operator and the maintenance machine is achieved through the communication interface submodule and the first monitoring communication submodule of the operator.

4. The interlocking host computer monitoring system as described in claim 1, characterized in that, The operating status information includes hardware operating status information; the interlocking host computer is configured with a monitoring base library provided by the interlocking host computer manufacturer, and the status monitoring submodule obtains the hardware operating status information through the monitoring base library; The hardware operating status information includes: static information and dynamic information; The static information includes: CPU model, CPU minimum frequency, CPU maximum frequency, number of physical CPU cores, number of logical CPU cores, number of memory modules, memory module model name, total physical memory capacity, total virtual memory capacity, number of hard drives, logical hard drive name, hard drive model name, total hard drive capacity, number of system fans, number of power supplies; The dynamic information includes: current CPU frequency, CPU utilization, CPU temperature, CPU fan speed, available physical memory capacity, used physical memory capacity, physical memory utilization, available virtual memory capacity, used virtual memory capacity, virtual memory utilization, used hard disk capacity, remaining hard disk capacity, hard disk utilization, system temperature, system fan speed, and power status.

5. The interlocking host computer monitoring system as described in claim 1, characterized in that, The operational status information also includes software operational status information; The status monitoring submodule obtains the software running status information through the operating system API of the interlocking host computer; The software running status information includes: process information and service information; The process information includes: process name, process identifier, process status, username, CPU utilization, number of threads, startup command line, and CPU time statistics; The service information includes: service name, associated process ID, and service status.

6. The interlocking host computer monitoring system as described in claim 1, characterized in that, The alarm and warning information includes: CPU usage alarms, memory usage alarms, hard disk space usage alarms, device temperature alarms, and power supply alarms.

7. The interlocking host computer monitoring system as described in claim 3, characterized in that, The host computer status monitoring module also includes an interface display submodule, which is used to visually display the operating status information and the alarm warning information, as well as the model of the interlocking host computer, the version number of the host computer status monitoring module, the version number of the monitoring base library, the operating system type value, the communication status between different interlocking host computers, the communication status between the first monitoring communication submodule and the corresponding communication interface submodule, and the communication status between the second monitoring communication submodule and the corresponding communication interface submodule.

8. The interlocking host computer monitoring system as described in claim 1, characterized in that, The host computer status monitoring module also includes an information storage submodule, used to store the operating status information, the alarm and warning information, the generation time of the alarm and warning information, and the alarm and warning status; The alarm and warning status includes: alarm and warning occurrence and alarm and warning recovery.

9. The interlocking host computer monitoring system as described in claim 8, characterized in that, It also includes an offline analysis device; when maintaining the interlocking host computer, the staff copies the operating status information and the alarm warning information in the information storage submodule to the offline analysis device. The offline analysis device includes a trained neural network model, which is used to generate at least one of the following based on the model of the interlocking host computer, the continuous running time, the running status information, and the alarm warning information: the recommended service life and the optimal restart interval of the corresponding interlocking host computer. The neural network model also updates its own model parameters based on the operating status information and alarm warning information within the most recent preset time period.

10. The interlocking host computer monitoring system as described in claim 2, characterized in that, The first processing module and the second processing module each include a first interface display unit and a second interface display unit, which are used to visually display the operating status information and the alarm warning information.

11. A method for monitoring interlocking systems on a host computer, used in the interlocking host computer monitoring system as described in any one of claims 1 to 10, characterized in that, Including the following steps: S1. Deploy a host computer status monitoring module on each interlock host computer. The host computer status monitoring module starts automatically when the host computer of the interlock is started. S2. The host computer status monitoring module collects the operating status information of the host computer in the interlock and generates corresponding alarm and early warning information; S3. The first processing module diagnoses the cause of the fault and recommends the corresponding fault handling strategy for the operator and the maintenance machine based on the operating status information and alarm warning information of the operator and the maintenance machine, respectively.

12. The interlocking host computer monitoring method as described in claim 11, characterized in that, The number of operating units is multiple, and each operating unit is communicatively connected to the maintenance unit via a dual network. Each operating unit also includes a second processing module. The method further includes the following steps: S4. The second processing module calculates a health value for the operator based on the communication quality between the operator and the external system and the maintenance machine; and corrects the health value based on the alarm warning information of the operator; and selects the operator with the highest health value as the master operator.

13. The interlocking host computer monitoring method as described in claim 11, characterized in that, The interlocking host computer monitoring system also includes an offline analysis device, which includes a trained neural network model. The method further includes the following steps: S5. When maintaining the interlocking host computer, copy the operating status information and the alarm warning information from the information storage submodule of the host computer status monitoring module to the offline analysis device; The neural network model generates at least one of the following for the interlocking host computer: suggested service life and optimal restart interval, based on the model of the interlocking host computer, continuous running time, operating status information, and alarm warning information. S6. The neural network model updates its own model parameters based on the operating status information and alarm warning information within the most recent preset time period.