A method and system for monitoring OT domain equipment in industrial scenarios
Through physical-communication dual-dimensional modeling and dynamic evaluation models, key equipment is identified. Combined with the impact propagation model and reverse tracing path, the problems of difficult fault location and slow impact assessment in OT domain equipment monitoring are solved, and panoramic monitoring and rapid emergency response are achieved.
Patent Information
- Application Number
- CN202510639136.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-05-19
AI Technical Summary
Existing technologies lack a global perspective on communication links between devices in the industrial OT domain, resulting in difficulty in fault location, disconnection between topology and communication, slow impact assessment, inability to effectively monitor the relationships between devices, and difficulty in predicting the scope of fault impact.
A physical-communication dual-dimensional modeling method is adopted to identify key equipment through a dynamic evaluation model. The fault impact range is predicted by combining the impact propagation model, and a reverse tracing path from the physical layer to the business layer is established. A full-domain communication monitoring map is constructed, and an emergency response is provided using the plan recommendation module.
It realizes panoramic monitoring of the communication status of OT domain equipment, improves fault location accuracy and response speed, accurately identifies key equipment, reduces fault location time, and improves emergency response efficiency.
Smart Images

Figure CN120186074B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of industrial automation and information communication technology, and more particularly relates to a method and system for monitoring OT domain equipment in industrial scenarios. Background Art
[0002] In actual industrial scenarios, the business scope is generally divided into two parts: the OT domain and the IT domain. The OT domain (Operational Technology Domain) refers to the technical field directly related to actual production processes and control, primarily including automated control systems, sensors, actuators, and other equipment. These devices are typically connected via dedicated networks to achieve efficient and secure production operations and are primarily responsible for the real-time control and physical operation of production equipment and processes. A medium-sized factory's OT domain typically includes 150 to 250 production devices, all interconnected and working together to complete a specific task. As domain managers, you must not only ensure the normal and stable operation of individual devices, but also ensure the normal completion of the work.
[0003] With the widespread adoption of industrial intelligence, real-time data collection from various devices enables real-time monitoring and intelligent analysis of individual device status. However, existing technologies for OT (online processing equipment) devices, which emphasize interdependencies, often focus solely on the health of individual devices, failing to delve into the business layer or conduct in-depth analysis of business lines. This monitoring effort falls far short of meeting actual production needs. For example, when a device experiences an anomaly during production, it's often necessary to identify which devices or businesses are affected so contingency plans can be prepared. Alternatively, if a device experiences an anomaly, it's necessary to identify which devices or businesses are affected, avoiding a period of navigating the problem.
[0004] However, the current status of existing monitoring equipment in the industrial OT domain is as follows: isolated equipment monitoring: traditional SCADA systems only monitor the status of a single device and lack a global perspective on the communication links between devices; fault location is difficult: when communication is interrupted, it is impossible to quickly assess the scope of impact and locate the physical layer fault point; topology and communication are disconnected: there is a lack of correlation analysis between the physical connection topology and the logical communication link, and the traceability efficiency is low.
[0005] Therefore, there is an urgent need for a comprehensive industrial scenario OT domain equipment monitoring method that can quickly locate OT domain equipment failures and predict their impact. Summary of the Invention
[0006] The embodiment of the present application provides an industrial scenario OT domain equipment monitoring method and system. The embodiment of the present application uses a physical-communication two-dimensional modeling method to perform panoramic monitoring of the communication status of OT domain equipment; uses a dynamic evaluation model to realize the identification of key equipment; uses an impact propagation model to predict the scope of fault impact; establishes a reverse tracing path from the physical layer to the business layer to achieve multi-dimensional positioning of the root cause of the fault; adopts a progressive construction method from the production line level to the global level to achieve real-time monitoring of a wide range of equipment in the entire OT domain; uses a plan recommendation module to recommend emergency plans to improve the speed and effectiveness of emergency response. It can be seen that this application effectively solves the industry pain points such as "difficult fault location", "slow impact assessment", and "weak panoramic perception" in the communication monitoring of industrial OT domain equipment, and has significant technological progress and industrial application value.
[0007] In a first aspect, an embodiment of the present application provides a method for monitoring OT domain equipment in an industrial scenario, the monitoring method comprising the following steps:
[0008] S1. Identify each production line / business line in the OT domain of the industrial scenario;
[0009] S2. Determine the equipment set included in each production line / business line;
[0010] S3. Build a production line-level physical topology map based on the physical connection relationship between devices;
[0011] S4. Build a communication chain between devices based on the communication connection relationship between devices and generate a communication link diagram;
[0012] S5. Using a dynamic assessment model based on business weight coefficients, identify key equipment in the production line / business line and mark them on the physical topology diagram and / or communication link diagram.
[0013] S6. Integrate the communication link map with the physical topology map to create a communication monitoring map;
[0014] S7. Repeat S2-S6 to process all production lines / business lines in the OT domain.
[0015] S8. Integrate the communication link diagram and physical topology diagram of each production line in the OT domain to form a global communication monitoring map;
[0016] S9. When a communication interruption is detected between the key devices, an impact range analysis is performed based on the communication connection relationship between the devices shown in the communication monitoring map;
[0017] S10. Perform communication fault location and source tracing analysis based on the physical topology relationship between the devices;
[0018] S11. Implement panoramic monitoring of the communication status of OT domain devices based on the global communication monitoring map.
[0019] Furthermore, the physical topology diagram in step S3 includes device type, physical interface parameters, transmission medium type and topology structure characteristic parameters.
[0020] Furthermore, the critical equipment identification in step S5 adopts a dynamic evaluation model based on a business weight coefficient, and the weight coefficient includes three dimensions: equipment downtime impact index, business relevance, and data throughput.
[0021] Furthermore, the dynamic evaluation model uses the following formula to calculate the key value K of the equipment:
[0022] K=α×P+β×R+γ×T,
[0023] Where P is the downtime impact index, R is the service relevance, T is the data throughput, and α, β, and γ are dynamic adjustment coefficients.
[0024] Furthermore, the communication monitoring map includes a monitoring and analysis module and a visual interactive interface, and supports protocols including OPC UA, Modbus TCP, and Profinet industrial protocols.
[0025] Furthermore, the impact range analysis in step S9 adopts an impact propagation model based on a directed graph to establish a device impact association matrix:
[0026] M ij =1 means that the failure of device i will affect device j,
[0027] M ij =0 means that the failure of device i does not affect device j.
[0028] Furthermore, the traceability analysis in step S10 adopts a reverse tracing algorithm to reversely search for potential fault points along the physical topology path and generate a traceability path report including the fault probability.
[0029] Furthermore, the method also includes an emergency plan recommendation module. When a communication interruption is detected, it matches similar cases based on the historical fault library and recommends an emergency plan. The plan includes spare parts information, disposal procedures, and a list of affected equipment.
[0030] In a second aspect, an embodiment of the present application provides an industrial scenario OT domain equipment communication monitoring system, including:
[0031] Physical topology construction module: used to build a production line-level physical topology based on the physical connection relationship between devices;
[0032] Communication link diagram construction module: constructs the communication chain between key devices based on the communication connection relationship between key devices and generates a communication link diagram;
[0033] Monitoring and Analysis Module: This module uses a dynamic assessment model based on business weight coefficients to identify key equipment in the production line / business line and annotate them on the physical topology diagram and / or communication link diagram. When a communication interruption is detected between key equipment, the module performs an impact analysis based on the inter-equipment communication connection relationships shown in the communication monitoring diagram. It also performs communication fault location and source tracing analysis based on the physical topology relationships between the equipment.
[0034] Global map database: used to store and integrate the communication link diagrams and physical topology diagrams of each production line in the OT domain, forming a global communication monitoring map;
[0035] Visual interactive interface: provides a visual interactive interface for panoramic monitoring of the communication status of OT domain devices based on the global communication monitoring map.
[0036] In a third aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the method as described in any one of the above-mentioned first aspects.
[0037] It can be understood that the beneficial effects of the second to third aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here.
[0038] Compared with the prior art, the present invention's embodiments offer the following advantages: The industrial OT domain device monitoring method and system employing the present invention innovatively integrates physical and logical communication topologies to construct a dual-dimensional modeling global graph with a physical and communication dual view, enabling comprehensive monitoring of OT domain device communication status and significantly improving fault location accuracy. The created dynamic assessment model for key devices, based on a dynamic assessment mechanism with multidimensional weight coefficients, addresses the issue of traditional static threshold determination methods lacking a dynamic assessment model that incorporates multiple dimensions, such as business impact and data throughput, making it difficult to accurately identify key devices. This improves key device identification accuracy and overall monitoring effectiveness. Using an impact propagation model, an intelligent impact analysis algorithm combining a directed graph propagation model with an association matrix is employed to predict the impact range of a fault, achieving the technical requirements for rapid and accurate impact range prediction responses. By establishing a reverse tracing path from the physical layer to the business layer and employing a cross-layer tracing mechanism, this solves the current problem of inefficient fault location caused by collecting network traffic solely through the SNMP protocol without integrating physical device connections during OT domain device monitoring. This enables multi-dimensional localization of the fault root cause, significantly reducing average fault location time. Using a progressive construction approach from the production line level to the global level, this system enables real-time monitoring of a wider range of devices across the entire OT domain, with wide coverage and low system latency. The plan recommendation module recommends emergency plans, improving the speed and effectiveness of emergency response. This effectively addresses industry pain points such as difficulty locating faults, slow impact assessment, and weak overall awareness in industrial OT device communication monitoring, demonstrating significant technological advancement and industrial application value. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0040] Figure 1 This is a flowchart of an industrial scenario OT domain equipment monitoring method provided by an embodiment of the present application;
[0041] Figure 2 This is a physical topology diagram of production line 1 provided in an embodiment of the present application;
[0042] Figure 3 This is a communication monitoring graph of production line 1 provided in an embodiment of the present application;
[0043] Figure 4 This is a communication monitoring diagram of a production line device when an abnormality occurs, provided by an embodiment of the present application;
[0044] Figure 5 This is a schematic diagram of production line division within the OT domain provided in one embodiment of the present application;
[0045] Figure 6 This is a schematic diagram of the OT global communication monitoring spectrum provided by an embodiment of the present application;
[0046] Figure 7 This is a schematic diagram of an industrial scenario OT domain equipment monitoring system provided by an embodiment of the present application;
[0047] Figure 8 This is a schematic diagram of an industrial control host that implements an industrial scenario OT domain equipment monitoring method provided by an embodiment of the present application;
[0048] Figure 9 is the device impact correlation matrix M constructed in one embodiment of the present application;
[0049] Figure 10 This is an embodiment of the present application to calculate M 2 (Two-step propagation) calculation matrix. DETAILED DESCRIPTION
[0050] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.
[0051] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.
[0052] It will also be understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0053] As used in this specification and the appended claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.
[0054] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0055] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.
[0056] The embodiment of the present application provides an industrial scenario OT domain equipment monitoring method and system. In one possible implementation, the method can be applied to an industrial control host. The industrial control host has: a physical topology map construction module for constructing a production line-level physical topology map based on the physical connection relationship between devices; a communication link map construction module for constructing a communication chain between key devices based on the communication connection relationship between the key devices and generating a communication link map; a monitoring and analysis module for performing an impact range analysis based on the communication connection relationship between devices shown in the communication monitoring map when a communication interruption is detected between key devices. The monitoring and analysis module can also use a dynamic evaluation model based on a business weight coefficient to identify the key equipment of the production line / business line, and mark them on the above-mentioned physical topology map and / or communication link map, and perform communication fault location and traceability analysis in combination with the physical topology relationship between the devices. Global map database: used to store the global communication monitoring map formed by integrating the communication link map and physical topology map of each production line in the OT domain; and realize visual interaction of panoramic monitoring of the communication status of OT domain devices through a visual interactive interface.
[0057] In one embodiment of the present invention, industrial OT domain equipment monitoring uses a physical-communication dual-dimensional modeling approach to perform panoramic monitoring of OT domain equipment communication status. Combined with the use of innovative technologies, this method effectively addresses industry pain points such as "difficult fault location," "slow impact assessment," and "weak panoramic perception" in industrial OT domain equipment communication monitoring, demonstrating significant technological advancement and industrial application value. Specifically, as shown in Figure 1, the main steps of this monitoring method include:
[0058] S1. Identify each production line / business line in the OT domain of the industrial scenario;
[0059] S2. Determine the equipment set included in each production line / business line;
[0060] S3. Build a production line-level physical topology map based on the physical connection relationship between devices;
[0061] S4. Build a communication chain between devices based on the communication connection relationship between devices and generate a communication link diagram;
[0062] S5. Using a dynamic assessment model based on business weight coefficients, identify key equipment in the production line / business line and mark them on the physical topology diagram and / or communication link diagram.
[0063] S6. Integrate the communication link map with the physical topology map to create a communication monitoring map;
[0064] S7. Repeat S2-S6 to process all production lines / business lines in the OT domain.
[0065] S8. Integrate the communication link diagram and physical topology diagram of each production line in the OT domain to form a global communication monitoring map;
[0066] S9. When a communication interruption is detected between the key devices, an impact range analysis is performed based on the communication connection relationship between the devices shown in the communication monitoring map;
[0067] S10. Perform communication fault location and source tracing analysis based on the physical topology relationship between the devices;
[0068] S11. Implement panoramic monitoring of the communication status of OT domain devices based on the global communication monitoring map.
[0069] The following is a detailed description of each step:
[0070] S1. Identify the production lines / business lines in the OT domain of industrial scenarios:
[0071] Before conducting comprehensive monitoring of the communication status of OT domain devices, it's necessary to first identify the business lines and production lines within the OT domain. This means identifying the number of production lines or business lines within the domain. This comprehensive analysis of the business lines and production lines within the OT domain is the foundation for comprehensive monitoring of the communication status of OT domain devices.
[0072] In one embodiment, the system obtains production line layout data through the equipment asset management system (EAM) and establishes a mapping table between the production line and the PLC controller, sensors, and actuators; or parses the namespace path of the OPC UA server (such as ns=3;s=Line1.Robot1) to extract the production line identifier to identify each production line / business line in the OT domain.
[0073] In another embodiment, the system divides the production lines by device IP address grouping logic: for example, by scanning the IP address segments of devices in the OT domain and combining them with subnet masks to divide the production lines (for example: 192.168.1.0 / 24 for stamping production lines, 192.168.2.0 / 24 for welding production lines).
[0074] In another embodiment, the business line division in the OT domain is achieved by importing production work order data from the MES (Manufacturing Execution System) and binding equipment groups to business processes (for example, work order number WO2023-001 corresponds to the "car body welding" business line).
[0075] like Figure 5 As shown, after sorting out, the OT domain equipment in this embodiment mainly includes production line 1 and production line 2.
[0076] S2. Determine the equipment set included in each production line / business line:
[0077] In one embodiment, when the production line identification of S1 is obtained through the equipment asset management system (EAM), the mapping relationship and set of equipment included in each production line are also obtained simultaneously.
[0078] In another embodiment, the system uses the Modbus TCP protocol to poll device register addresses (e.g., holding registers 40001-40010) to obtain the device type (PLC, sensor, actuator), firmware version, and communication protocol support list. For devices that do not support standard protocols (e.g., proprietary protocol devices), protocol conversion is performed through a gateway.
[0079] S3. Build a production line-level physical topology diagram based on the physical connection relationship between devices:
[0080] The production line physical topology is a topology constructed based on the physical connection relationship of each device in the production line, such as Figure 2 This is a schematic diagram of the physical topology of production line 1 in the above embodiment.
[0081] In one embodiment, the physical topology diagram in step S3 includes annotations of topology parameters such as device type, physical interface parameters, transmission medium type, and topology structure characteristic parameters. These topology parameters help clarify the physical connection characteristics of production line equipment and facilitate the implementation of various additional functions in subsequent global monitoring.
[0082] In one embodiment, the various topology parameters described above are automatically acquired by the system, for example, using a dedicated industrial gateway with information collection capabilities. This method of automatically acquiring topology parameters is applicable to production lines involving numerous devices and complex physical connections.
[0083] Device type: PLC (Siemens S7-1500), industrial switch (Hirschmann RS30), temperature sensor (PT100), motor drive (ABB ACS880);
[0084] Physical interface parameters: Profinet port (RJ45-M12), RS485 interface impedance (120Ω±5%), fiber optic connector (IP67 protection grade);
[0085] Transmission media type: Industrial-grade twisted pair (CAT6e-SF / UTP), armored fiber optic cable (G.657A2), CAN bus cable (ISO 11898 standard);
[0086] Topology features: Ring redundant network (backbone ring network + star-shaped device access).
[0087] In one embodiment, in addition to the system administrator manually establishing a physical topology map based on the physical connection relationships of the production line equipment, the system can also traverse the device connection relationships based on a depth-first search (DFS) to generate a physical topology map represented by an adjacency list, further reducing the workload of the system administrator.
[0088] S4. Build a communication chain between devices based on the communication connection relationship between devices and generate a communication link diagram;
[0089] The communication link diagram primarily documents the topological relationships between devices on a production line. In one embodiment, the communication link diagram uses a computer to automatically and rapidly create a communication link diagram by analyzing the five-tuple structure (source IP, destination IP, port, protocol, and timestamp) of TCP / UDP conversations between devices and constructing a communication adjacency matrix.
[0090] In another embodiment, a system administrator manually creates a communication link diagram based on the communication connection relationships of the production line equipment.
[0091] S5. Use a dynamic assessment model based on business weight coefficients to identify key equipment in the production line / business line and mark them on the physical topology diagram and / or communication link diagram:
[0092] In this application's OT domain device monitoring, accurate identification of critical equipment is crucial. Identifying critical equipment allows for robust design of inter-device links during the communication design phase, and the implementation of various methods such as redundancy and backup to enhance link reliability. Furthermore, identifying critical equipment allows for higher-level monitoring during routine equipment and network monitoring, ensuring efficient and effective production operations without increasing overall monitoring costs.
[0093] In one embodiment of the present application, key equipment identification uses a dynamic evaluation model based on business weight coefficients, which avoids the bias of manual experience and makes the identification of key equipment more accurate. The dynamic evaluation model is as follows:
[0094] Key equipment identification uses a dynamic evaluation model based on business weight coefficients. The weight coefficients include three dimensions: equipment downtime impact index, business relevance, and data throughput. The dynamic evaluation model uses the following formula to calculate the equipment key value K:
[0095] K = α × P + β × R + γ × T;
[0096] Where P is the downtime impact index, R is the service relevance, T is the data throughput, and α, β, and γ are dynamic adjustment coefficients.
[0097] The definition and acquisition method of each dimension are as follows:
[0098] 1. Downtime impact index P:
[0099] P = (number of equipment failures × average repair time) × 100% / total production line operating time;
[0100] For example, if we import historical alarm logs from the SCADA system and find that a PLC has three monthly failures, the average repair time is 30 minutes, and the monthly production line operation time is 720 hours, then P = 3 × 0.5 × 100% / 720 = 0.21%.
[0101] It can be seen that for production line equipment, the larger the value of P, the greater the impact of the equipment shutdown on the operation of the entire production line.
[0102] 2. Business relevance R:
[0103] Business relevance is an important indicator for clarifying the relationship between devices. It can be calculated by counting the number of times the devices are associated in the BOM (Bill of Materials). The formula is as follows:
[0104] R = log (1 + ∑BOM association times);
[0105] For example, if a motor is referenced 15 times in the BOM, then:
[0106] R=log(1+15)=2.77;
[0107] It can be seen that for production line equipment, the larger the value of R, the closer the equipment is to other equipment in the production line, and the greater the scope of impact caused by a failure.
[0108] Data throughput T:
[0109] Key equipment in a production line is the core of the production line and typically requires extensive data exchange to coordinate with other equipment or report production data in a timely manner. Therefore, the device's data throughput per unit time—the number of bytes transmitted per second—is used as one of the key equipment evaluation metrics.
[0110] In one example, to remove deviations caused by differences in production line equipment type, transmitted data protocols, or types, different protocols are weighted to unify statistical caliber (e.g., OPC UA weight = 1.2, Modbus = 1.0).
[0111] Dynamic adjustment coefficients (α, β, γ):
[0112] In order to improve the adaptability of the dynamic model, the dynamic adjustment coefficient is dynamically adjusted according to the production stage (such as normal / peak period). By adjusting the weights of the three dimensions of equipment downtime impact index, business relevance, and data throughput, the identification of key equipment can be made more in line with actual business requirements.
[0113] For example, during peak hours, the α weight is increased (α=0.5→0.7) to emphasize the impact of downtime.
[0114] The dynamic evaluation model of this application can conveniently and accurately identify key equipment. The dynamic evaluation mechanism based on multi-dimensional weight coefficients overcomes the shortcomings of the traditional static threshold judgment method, which is unable to adapt to the dynamic adjustment of production plans, and the problem of inaccurate identification of key equipment caused by relying on manual experience. It improves the accuracy of key equipment identification and enhances the overall monitoring effect.
[0115] In one embodiment, the key equipment on the production line / business line identified by the dynamic assessment model based on the business weight coefficient is annotated on the physical topology diagram completed in step S3 and / or the communication link diagram completed in step S4, and these annotations are retained for subsequent steps. This operation ensures that when the communication monitoring map is merged for all production lines / business lines in the entire OT domain, the problem of key equipment identification distortion caused by the different requirements and business priorities of each production line is avoided.
[0116] S6. Integrate the communication link diagram with the physical topology diagram to create a communication monitoring map:
[0117] The physical topology diagram completed in step S3 and the communication link diagram completed in step S4 are integrated to form the following Figure 3 The communication monitoring map of production line 1 shown in the figure distinguishes the physical topology connection relationship and the communication link connection relationship through different line types. It is a two-dimensional modeling communication monitoring map with a physical-communication dual view.
[0118] S7. Repeat S2-S6 to process all production lines / business lines in the OT domain:
[0119] Repeat S2-S6 for all production lines / business lines in the OT domain to generate a communication monitoring map with a dual physical and communication view for all production lines / business lines.
[0120] S8. Integrate the communication link diagram and physical topology diagram of each production line in the OT domain to form a global communication monitoring map:
[0121] In one embodiment, Figure 6 As shown, the communication link diagrams and physical topology diagrams of production lines 1 and 2 within the OT domain are integrated, and a global communication monitoring map is formed using a progressive construction method from the line level to the global level. The constructed global communication monitoring map shows that the physical topology relationships between devices cannot accurately reflect all communication relationships between devices. When the number of devices within the business scope to be monitored is extremely large (for example, the OT domain of a medium-sized factory typically has 150 to 250 production equipment), simply monitoring the physical topology diagram or performing single-device status monitoring and analysis will not be able to monitor the business chain and core business scope. This can easily lead to omissions or inaccurate event prioritization, thereby increasing the overall losses caused by single-device problems. The method of this application focuses on the global situation, constructing and analyzing the communication relationships between a group of devices. This allows for global management and control from a complete "link" or even "network" perspective. Building on the single-device monitoring and analysis, it expands the perspective, increases the dimension, and differentiates priorities, thereby taking responsibility for the overall interests and providing a new approach to global analysis in the OT domain.
[0122] S9. When a communication interruption is detected between the key devices, an impact range analysis is performed based on the communication connection relationship between the devices shown in the communication monitoring map:
[0123] In one embodiment of the present application, the OT domain equipment monitoring system for industrial scenarios supports industrial protocols including OPC UA, Modbus TCP, Profinet, etc.
[0124] In one embodiment of the present application, the communication monitoring map includes a monitoring and analysis module and a visual interactive interface. The monitoring and analysis module is used to perform data analysis and status monitoring of the OT domain equipment monitoring system. The main functions performed include but are not limited to: using the above-mentioned dynamic evaluation model based on the business weight coefficient to identify the key equipment of the production line / business line and complete the annotation of the key equipment on the above-mentioned physical topology map and / or communication link map; performing impact range analysis based on the communication connection relationship between devices shown in the communication monitoring map; and performing communication fault location and traceability analysis based on the physical topology relationship between the devices.
[0125] In one embodiment, when the monitoring and analysis module detects a communication interruption between the key devices, it performs an impact range analysis based on the inter-device communication connection relationship shown in the communication monitoring graph. Specifically, the impact range analysis is performed by combining the inter-device communication connection relationship shown in the communication monitoring graph with an impact propagation model based on a directed graph, as follows:
[0126] In this embodiment, the impact range analysis adopts an impact propagation model based on a directed graph to establish a device impact association matrix:
[0127] M ij =1 means that the failure of device i will affect device j,
[0128] M ij =0 means that the failure of device i does not affect device j.
[0129] Assume that there are 4 devices (A, B, C, D) in a power system. The fault propagation relationship between the devices is as follows (directed graph):
[0130] A failure in device A will directly affect devices B and C.
[0131] Failure of device B will directly affect C and D.
[0132] Failure of device C will directly affect device D.
[0133] A failure of device D does not affect other devices.
[0134] Step S901: Constructing a device impact correlation matrix M;
[0135] According to the definition:
[0136] M ij =1 means that the failure of device i will affect device j,
[0137] M ij =0 means that the failure of device i does not affect device j.
[0138] Construct a matrix such as Figure 9 The device impact correlation matrix M shown below (rows represent fault source device i, columns represent affected device j):
[0139] explain:
[0140] First row (device A): A failure directly affects B (M AB =1) and C(M AC =1),
[0141] Second row (device B): B failure directly affects C (M BC=1) and D(M BD =1),
[0142] The third row (device C): C failure directly affects D (M CD =1),
[0143] Fourth row (device D): The D fault does not affect any device and is all 0.
[0144] Step S902: Analyze the fault propagation range:
[0145] Assume that device A fails, according to matrix M:
[0146] 1. Direct impact: A → B and A → C (by M AB =1 and M AC =1).
[0147] 2. Indirect impact (cascading):
[0148] After B is affected, B → C and B → D (by M BC =1 and M BD =1).
[0149] After C is affected, C → D (by M CD =1).
[0150] 3. Final impact range: A → B → C → D and A → C → D, all devices are affected.
[0151] Step S903: mathematically represent the propagation path:
[0152] Through the matrix power operation M K The k-step indirect impact can be analyzed:
[0153] Direct influence: represented by M itself.
[0154] Indirect influence: by M 2 .M 3 ......express.
[0155] For example, to calculate M 2 (Two-step propagation) Figure 10 As shown:
[0156] M 2 AD =2 means starting from A, it affects D in two steps through two paths (A→B→D and A→C→D).
[0157] The following problems can be solved by the correlation matrix M:
[0158] 1) Impact range and key device identification: If the row where a device is located has many non-zero elements, then the impact range of its failure is large (such as devices A and B).
[0159] 2) Maintenance Priority: Prioritize protecting high-impact equipment (e.g., repairing A can prevent cascading failures of B, C, and D).
[0160] 3) Fault Isolation Strategy: If D fails, tracing back to M reveals that the source may be B or C.
[0161] Therefore, the correlation matrix M in this embodiment intuitively depicts the direct relationship between fault propagation between devices. Matrix operations (such as exponentiation) can quantify the cascading impact range, making it suitable for automated and accurate impact range estimation, thereby improving the response speed and accuracy of impact range warnings. By analyzing the impact range of communication disruptions between key devices based on the inter-device communication connection relationships depicted in the communication monitoring graph, the impact range analysis of communication disruptions between key devices can be performed. This fully utilizes the inherent advantage of the correlation matrix, as network nodes / hops in the communication monitoring graph can be easily retrieved from network connection logs. This allows for rapid construction of the aforementioned correlation matrix. Furthermore, accurate and comprehensive impact range analysis of key devices enables more targeted and efficient response to communication disruptions.
[0162] S10. Perform communication fault location and source tracing analysis based on the physical topology relationship between the devices:
[0163] In one embodiment, a reverse tracing algorithm is used to reversely search for potential fault points along the physical topology path, and a traceability path report including the fault probability is generated.
[0164] The specific steps are as follows:
[0165] Step S1001: Fault triggering and feature extraction:
[0166] For example, when the 3# motor driver of a steel rolling production line frequently reports communication timeout faults:
[0167] The PLC detects that the Profinet bus cycle is extended to 15ms (standard value 4ms±1ms).
[0168] Based on the physical topology between devices, the affected paths are:
[0169] Motor drive ACS880-03: M12 interface → Fieldbus Segment-C (length 82m, including three T-connectors): Port P5 → Industrial switch SW-07: Fiber optic ring network → Master PLC S7-1500-01.
[0170] Step S1002: Reverse physical path investigation:
[0171] Constructing a reverse detection path:
[0172] Fault phenomenon layer: ACS880-03 drive (communication timeout rate 38%) ← Field layer: Bus Segment-C (measured impedance fluctuation ±8Ω) ← Network layer: SW-07 switch (port P5 packet error rate 0.7%) ← Control layer: S7-1500-01 PLC (CPU load rate 92%).
[0173] Step S1003: Multi-dimensional failure probability assessment:
[0174] The failure probability is calculated based on the characteristics of industrial equipment. The results are shown in Table 1:
[0175] Table 1 Industrial equipment characteristic calculation failure probability analysis table
[0176] Troubleshooting nodes Evaluation Metrics Weight Probability bus Segment-C connector oxidation + length exceeds the standard 0.55 78% SW-07 Switch port buffer overflow + heat dissipation abnormality 0.30 65% PLC controller CPU overload + memory usage exceeds the limit CPU overload 0.15 42%
[0177] Step S1004: Generate industrial-grade traceability report:
[0178] Output a special report containing the following elements:
[0179] 1) Topology heat map: The high-risk path from bus Segment-C to switch P5 is marked in red.
[0180] 2) Repair Priority: Urgently replace the T-connector on bus segment C (78% probability) and upgrade the cooling system of the SW-07 switch (65% probability).
[0181] 3) Calculate the impact range based on step S9: associate 12 devices and 3 production line control signals.
[0182] This example establishes a reverse tracing path from the physical layer to the service layer, leveraging a cross-layer traceability mechanism. This solves the current problem of inefficient fault location caused by monitoring OT devices, which relies solely on collecting network traffic through the SNMP protocol without integrating physical device connections. This enables multi-dimensional root cause location, significantly reducing average fault location time.
[0183] S11. Based on the global communication monitoring map, a panoramic monitoring of the communication status of OT domain devices is realized:
[0184] In one embodiment of the present application, the global communication monitoring map uses a visual interactive interface to perform a panoramic monitoring of the communication status of all OT devices. For example, the device status is pushed in real time via WebSocket, and in order to highlight the status of key devices, the system also supports filtering alarms according to the key value K of the device (devices with K>0.7 are highlighted). Or Figure 4 As shown in the figure, a fault icon is used in the communication monitoring diagram to indicate that a device is abnormal.
[0185] In another embodiment, the system further includes an emergency plan recommendation module for matching similar cases based on a historical fault database and recommending an emergency plan in a timely manner when a communication interruption is detected.
[0186] For example, the historical fault database stores structured case data, which includes field categories such as "fault characteristics", "handling plan", "spare parts information", "affected equipment list", "historical handling records" and other information that represents fault information and corresponding response strategies.
[0187] The emergency plan recommendation module also includes an emergency plan recommendation engine, which analyzes the encountered fault information and automatically matches the "fault characteristics" in the historical fault library in multiple dimensions to accurately match the fault with the historical fault library, thereby providing an accurate emergency plan. In one embodiment, the dimensions include:
[0188] Device type matching: historical fault resolution plans for PLCs of the same model (weight 60%);
[0189] Environmental parameter correlation: Cases of joint oxidation caused by high humidity (>80%RH) (weight 25%);
[0190] Communication Protocol Characteristics: Comparison of Profinet and Modbus TCP anomaly characteristics (weighting 15%).
[0191] Based on the same technical concept, such as Figure 7 As shown, the embodiment of the present application also provides an industrial scenario OT domain equipment communication monitoring system. For the sake of convenience, only the parts related to the embodiment of the present application are shown. The device includes:
[0192] Physical topology construction module: used to build a production line-level physical topology based on the physical connection relationship between devices;
[0193] Communication link diagram construction module: constructs a communication chain between key devices based on the communication connection relationship between the key devices and generates a communication link diagram;
[0194] Monitoring and Analysis Module: This module uses a dynamic assessment model based on business weight coefficients to identify key equipment in the production line / business line and annotate them on the physical topology diagram and / or communication link diagram. When a communication interruption is detected between key equipment, the module performs an impact analysis based on the inter-equipment communication connection relationships shown in the communication monitoring diagram. It also performs communication fault location and source tracing analysis based on the physical topology relationships between the equipment.
[0195] Global map database: used to store and integrate the communication link diagrams and physical topology diagrams of each production line in the OT domain, forming a global communication monitoring map;
[0196] Visual interactive interface: provides a visual interactive interface for panoramic monitoring of the communication status of OT domain devices based on the global communication monitoring map.
[0197] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0198] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of this application. Their specific functions and technical effects can be found in the method embodiment section and will not be repeated here.
[0199] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0200] like Figure 8 As shown, an embodiment of the present application also provides an industrial control host, wherein the industrial control host device includes a memory, a processor, and a computer program stored in the memory and runnable on the processor, and when the processor executes the industrial control host program, the steps in any of the above method embodiments are implemented.
[0201] An embodiment of the present application provides a computer program product. When the computer program product is run on a terminal, the terminal can implement the steps in the above-mentioned various method embodiments when executing the computer program product.
[0202] If the integrated unit is implemented as a software functional unit and sold or used as a standalone product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the process steps in the above-mentioned method embodiments by using a computer program to instruct the relevant hardware. The computer program can be stored in a computer-readable storage medium. When executed by a processor, the computer program can implement the steps of each of the above-mentioned method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium can include at least: any entity or device capable of carrying the computer program code to the camera / target device, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signals, telecommunication signals, and software distribution media. Examples include USB flash drives, removable hard drives, magnetic disks, or optical disks. In some jurisdictions, based on legislation and patent practice, computer-readable media cannot be electric carrier signals or telecommunication signals.
[0203] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0204] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0205] In the embodiments provided in this application, it should be understood that the disclosed devices / network equipment and methods can be implemented in other ways. For example, the device / network equipment embodiments described above are merely illustrative. For example, the division of the modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0206] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected as needed to achieve the purpose of this embodiment.
[0207] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.
Claims
1. A method for monitoring OT domain equipment in industrial scenarios, characterized in that: The monitoring method comprises the following steps: S1. Obtain production line layout data through the equipment asset management system (EAM), parse the namespace path of the OPC UA server to extract the production line identifier, divide each production line by device IP address grouping logic, or implement business line division by importing production work order data from the MES (Manufacturing Execution System) to identify each production line / business line in the industrial scenario's OT domain. S2. Determine the equipment set included in each production line / business line; S3. Build a production line-level physical topology map based on the physical connection relationship between devices; S4. Build a communication chain between devices based on the communication connection relationship between devices and generate a communication link diagram; S5. Use a dynamic assessment model based on business weight coefficients to identify key equipment in the production line / business line and mark them on the physical topology diagram and / or communication link diagram. The weight coefficients include three dimensions: equipment downtime impact index, business relevance, and data throughput. The dynamic assessment model uses the following formula to calculate the equipment key value K: K=α×P+β×R+γ×T, Where P is the downtime impact index, R is the service relevance, T is the data throughput, and α, β, and γ are dynamic adjustment coefficients; S6. Integrate the communication link map with the physical topology map to create a communication monitoring map; S7. Repeat S2-S6 to process all production lines / business lines in the OT domain. S8. Integrate the communication link diagram and physical topology diagram of each production line in the OT domain to form a global communication monitoring map; S9. When a communication interruption is detected between the key devices, an impact range analysis is performed based on the communication connection relationship between the devices shown in the communication monitoring map; S10. Perform communication fault location and source tracing analysis based on the physical topology relationship between the devices; S11. Implement panoramic monitoring of the communication status of OT domain devices based on the global communication monitoring map.
2. The industrial scene OT domain equipment monitoring method according to claim 1 is characterized in that: The physical topology diagram in step S3 includes device type, physical interface parameters, transmission medium type and topology structure characteristic parameters.
3. The industrial scene OT domain equipment monitoring method according to claim 1 is characterized in that: The communication monitoring map includes a monitoring and analysis module and a visual interactive interface, and supports protocols including OPC UA, Modbus TCP, and Profinet industrial protocols.
4. The industrial scene OT domain equipment monitoring method according to claim 1 is characterized in that: The impact range analysis in step S9 adopts an impact propagation model based on a directed graph, and establishes an equipment impact association matrix to analyze the fault propagation range.
5. The industrial scene OT domain equipment monitoring method according to claim 1 is characterized in that: The traceability analysis in step S10 adopts a reverse tracing algorithm to reversely search for potential fault points along the physical topology path and generate a traceability path report including the fault probability.
6. The industrial scene OT domain equipment monitoring method according to claim 1, characterized in that: The method further includes a plan recommendation module, which, when a communication interruption is detected, matches similar cases based on a historical fault database and recommends an emergency plan, the plan including spare parts information, disposal procedures, and a list of affected equipment.
7. An industrial scene OT domain equipment communication monitoring system, characterized in that: include: Based on the global communication monitoring map, a panoramic monitoring of the communication status of OT domain devices is achieved, including: Production Line / Business Line Identification Module: This module obtains production line layout data from the Equipment Asset Management (EAM), parses the namespace path of the OPC UA server to extract the production line identifier, divides production lines by device IP address grouping logic, or imports production work order data from the MES (Manufacturing Execution System) to implement business line division, thereby identifying each production line / business line in the OT domain of industrial scenarios. Equipment determination module: determines the equipment set included in each production line / business line; Physical topology construction module: Builds a production line-level physical topology based on the physical connection relationships between devices; Communication link diagram construction module: constructs the communication chain between devices based on the communication connection relationship between devices and generates a communication link diagram; Monitoring and Analysis Module: A dynamic evaluation model based on business weight coefficients is used to identify key equipment in the production line / business line and annotate them on the physical topology diagram and / or communication link diagram. The weight coefficients include three dimensions: equipment downtime impact index, business relevance, and data throughput. The dynamic evaluation model uses the following formula to calculate the equipment key value K: K=α×P+β×R+γ×T, Where P is the downtime impact index, R is the service relevance, T is the data throughput, and α, β, and γ are dynamic adjustment coefficients; Communication monitoring map integration module: integrates the communication link map with the physical topology map to establish a communication monitoring map; Global communication monitoring map construction module: This module processes all production lines / business lines within the OT domain and integrates the communication link diagram and physical topology diagram of each production line within the OT domain to form a global communication monitoring map. Impact range analysis module: when a communication interruption is detected between the key devices, an impact range analysis is performed based on the communication connection relationship between the devices shown in the communication monitoring map; Positioning and tracing analysis module: performs communication fault location and tracing analysis based on the physical topology relationship between the devices.
8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Sliding revolver pistol
WO2020023001A1
Transmission network fault positioning method, system and device and storage medium
CN119520248A
Network topology and equipment resource information display method, device, equipment, medium and product
CN119676093A