Industrial scene OT domain equipment monitoring method and system

By using physical-communication dual-dimensional modeling and dynamic evaluation model in OT domain equipment monitoring, we identify key equipment and predict the scope of failure impact, solving the problems of difficulty and slow impact assessment in OT domain equipment monitoring, and achieving more efficient fault location and impact scope prediction.

CN120186074AActive Publication Date: 2025-06-20FENGTAI SCI & TECH (BEIJING) CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510639136.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-06-20
Estimated Expiration
2045-05-19

AI Technical Summary

Technical Problem

The existing technology lacks a global perspective in the monitoring of OT domain equipment, has difficulty in fault positioning, and is disconnected from topology and communication, resulting in inefficient assessment of the scope of the fault impact and positioning.

Method used

Through physical-communication dual-dimensional modeling, a panoramic monitoring map of communication status of OT domain devices is constructed, a dynamic evaluation model is used to identify key devices, a impact propagation model is used to predict the range of fault impact, and multi-dimensional fault location is achieved through reverse tracking paths.

Benefits of technology

It improves the accuracy of fault positioning, shortens the fault positioning time, enhances the prediction ability of the fault impact range, and improves the overall monitoring effect and emergency response speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120186074A_ABST
    Figure CN120186074A_ABST
Patent Text Reader

Abstract

The invention discloses an industrial scene OT domain equipment monitoring method and system. OT domain equipment communication state panoramic monitoring is carried out in a physical-communication two-dimensional modeling mode; the key equipment is identified by using the dynamic evaluation model; performing fault influence range prediction by using an influence propagation model; establishing a reverse tracking path from a physical layer to a service layer, and supporting multi-dimensional positioning of a fault root; an initiative production line level to global level progressive construction method is adopted, and real-time monitoring of all-OT-domain wide equipment can be achieved; using a plan recommendation module to recommend an emergency plan; the embodiments effectively solve the industrial pain points of difficult fault positioning, slow influence evaluation, weak panoramic perception and the like in industrial OT domain equipment communication monitoring, and have remarkable technical progress and industrial application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of industrial automation and information communication technology. In particular, it relates to a method and system for monitoring OT domain devices in an industrial scenario. Background Art

[0002] In an actual industrial scenario, from the dimension of business scope, it is generally divided into two parts: the OT domain and the IT domain. Among them, the OT domain (Operational Technology Domain) refers to the technical field directly related to the actual production process and control, mainly including devices such as automated control systems, sensors, and actuators. These devices are usually connected through a private network to achieve efficient and safe production operations, mainly responsible for the real-time control and physical operations of production equipment and processes. In the OT domain of a medium-sized factory, there are usually 150 to 250 production devices. These devices are interconnected and work together to complete a certain task. As a domain manager, not only does one need to ensure the normal and stable operation of individual devices, but also ensure the normal completion of work.

[0003] With the popularization of industrial intelligence, through real-time collection of various device data, it is possible to achieve real-time monitoring and intelligent analysis of the status of individual devices. However, in the prior art, for OT domain devices that emphasize the correlation relationship, most only focus on whether an individual device is normal, and do not go deep into the business level, nor conduct in-depth analysis of business lines. Such monitoring effects far from meet the actual needs of production. For example, in production activities, when a certain device malfunctions, it is often necessary to clarify which devices or services it will affect in order to prepare a plan in advance; or when an abnormality is found in a certain device, which devices or services have affected it, etc., to avoid sticking to the original point when troubleshooting problems.

[0004] However, the current situation of existing industrial OT domain monitoring devices is as follows: isolated device monitoring: traditional SCADA systems only monitor the status of individual devices and lack a global perspective on the communication links between devices; difficult fault location: when communication is interrupted, it is impossible to quickly evaluate the affected range and locate the physical layer fault point; disconnection between topology and communication: there is a lack of correlation analysis between the physical connection topology and the logical communication link, and the tracing efficiency is low.

[0005] Therefore, there is an urgent need for a comprehensive method for monitoring OT domain devices in an industrial scenario, which can achieve rapid fault location and impact prediction of OT domain devices. Summary of the Invention

[0006] The embodiments of the present application provide an industrial scenario OT domain device monitoring method and system. The embodiments of the present application perform panoramic monitoring of the communication status of OT domain devices through physical-communication two-dimensional modeling; use a dynamic evaluation model to identify key devices; use an impact propagation model to predict the scope of fault impact; achieve multi-dimensional positioning of the root cause of faults by establishing a reverse tracing path from the physical layer to the service layer; adopt a progressive construction method from the production line level to the global level to achieve real-time monitoring of a wide range of devices in the entire OT domain; use a pre-plan recommendation module to recommend emergency plans to improve the speed and effectiveness of emergency response. It can be seen that the present application effectively solves the industry pain points such as "difficulty in fault location", "slow impact assessment", and "weak panoramic perception" in industrial OT domain device communication monitoring, and has significant technological progress and industrial application value.

[0007] In the first aspect, the embodiments of the present application provide an industrial scenario OT domain device monitoring method, and the monitoring method includes the following steps: S1. Identify each production line / business line in the industrial scenario OT domain; S2. Determine the set of devices included in each production line / business line; S3. Construct a production line-level physical topology diagram based on the physical connection relationship between devices; S4. Construct a communication link between devices based on the communication connection relationship between devices, and generate a communication link diagram; S5. Use a dynamic evaluation model based on service weight coefficients to identify the key devices of the production line / business line, and mark them on the above physical topology diagram and / or communication link diagram; S6. Integrate the communication link diagram and the physical topology diagram to establish a communication monitoring map; S7. Repeat S2 - S6 to process all production lines / business lines in the OT domain; S8. Integrate the communication link diagrams and physical topology diagrams of each production line in the OT domain to form a global communication monitoring map; S9. When a communication interruption occurs between the key devices is detected, perform an impact range analysis based on the communication connection relationship between devices shown in the communication monitoring map; S10. Combine the physical topology relationship between devices to perform communication fault location and traceability analysis; S11. Achieve panoramic monitoring of the communication status of OT domain devices based on the global communication monitoring map.

[0008] Further, the physical topology diagram in step S3 includes device type, physical interface parameters, transmission medium type, and topology structure characteristic parameters.

[0009] Further, in step S5, the identification of key devices adopts a dynamic evaluation model based on business weight coefficients, and the weight coefficients include three dimensions: equipment downtime impact index, business relevance, and data throughput.

[0010] Further, the dynamic evaluation model calculates the key value K of the device using the following formula: K = α × P + β × R + γ × T, where P is the downtime impact index, R is the business relevance, T is the data throughput, and α, β, and γ are dynamic adjustment coefficients.

[0011] Further, the communication monitoring map includes a monitoring and analysis module and a visualization interaction interface, and the supported protocols include OPC UA, Modbus TCP, and Profinet industrial protocols.

[0012] Further, the impact range analysis in step S9 adopts an impact propagation model based on a directed graph to establish a device impact association matrix: M ij = 1 indicates that the failure of device i will affect device j, M ij = 0 indicates that the failure of device i does not affect device j.

[0013] Further, the traceability analysis in step S10 adopts a reverse tracing algorithm to reverse-search for potential fault points along the physical topology path and generate a traceability path report containing the fault probability.

[0014] Further, the method further includes a pre-plan recommendation module. When a communication interruption is detected, similar cases are matched based on the historical fault library, and an emergency plan is recommended. The emergency plan includes spare part information, disposal procedures, and a list of affected devices.

[0015] In a second aspect, an industrial scenario OT domain device communication monitoring system provided by an embodiment of the present application includes: A physical topology map construction module: used to construct a production line-level physical topology map based on the physical connection relationship between devices; A communication link map construction module: constructs a communication link between key devices based on the communication connection relationship between key devices and generates a communication link map; A monitoring and analysis module: adopts a dynamic evaluation model based on business weight coefficients to identify the key devices of the production line / business line and mark them on the above physical topology map and / or communication link map. When a communication interruption occurs between the key devices is detected, an impact range analysis is performed based on the communication connection relationship between devices shown in the communication monitoring map; communication fault location and traceability analysis are combined with the physical topology relationship between devices; Global Atlas Database: It is used to store the global communication monitoring atlas formed by integrating the communication link diagrams and physical topology diagrams of each production line in the OT domain. Visualization Interaction Interface: It provides a visualization interaction interface for realizing panoramic monitoring of the communication status of OT domain devices based on the global communication monitoring atlas.

[0016] Thirdly, the embodiment of the present application provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, it implements the method described in any one of the above first aspects.

[0017] It can be understood that the beneficial effects of the above second aspect to the third aspect can refer to the relevant descriptions in the above first aspect, and will not be elaborated here.

[0018] The beneficial effects of the embodiment of the present application compared with the prior art are as follows: The industrial scenario OT domain device monitoring method and system adopting the method of the present invention construct a two-dimensional modeling global atlas with physical-communication dual views through innovatively integrating the physical topology and the communication logic topology to conduct panoramic monitoring of the communication status of OT domain devices, significantly improving the fault location accuracy. The created key device dynamic evaluation model, based on the dynamic evaluation mechanism of multi-dimensional weight coefficients, solves the problem that the traditional static threshold determination method lacks a dynamic evaluation model combining multiple dimensions such as business impact and data throughput, making it difficult to accurately identify key devices, improving the recognition accuracy of key devices and enhancing the overall monitoring effect. Using the impact propagation model, the intelligent impact analysis algorithm combining the directed graph propagation model and the correlation matrix is used to predict the fault impact range, achieving the technical requirements of fast and accurate impact range prediction response. By establishing a reverse tracing path from the physical layer to the service layer and through the cross-layer traceability positioning mechanism, the problem that only network traffic is collected through the SNMP protocol and the physical connection relationship of devices is not combined during the current OT domain device monitoring, resulting in low fault location efficiency, is solved, realizing multi-dimensional positioning of the root cause of the fault and greatly reducing the average fault location time. Adopting a progressive construction method from the production line level to the global level, real-time monitoring of a wider range of devices in the entire OT domain is realized, with a wide coverage range and low system latency. Using the pre-plan recommendation module to recommend emergency plans to improve the emergency response speed and effect. It effectively solves the industry pain points such as "difficult fault location", "slow impact assessment", and "weak panoramic perception" in industrial OT domain device communication monitoring, and has significant technological progress and industrial application value. Description of the Drawings

[0019] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0020] Figure 1 It is a schematic flowchart of the industrial scenario OT domain device monitoring method provided by an embodiment of the present application; Figure 2 It is a physical topology relationship diagram of production line 1 provided by an embodiment of the present application; Figure 3 It is a communication monitoring map of production line 1 provided by an embodiment of the present application; Figure 4 It is a communication monitoring map when the device of production line 1 is abnormal provided by an embodiment of the present application; Figure 5 It is a schematic diagram of production line division within the OT domain provided by an embodiment of the present application; Figure 6 It is a schematic diagram of the OT global communication monitoring map provided by an embodiment of the present application; Figure 7 It is a schematic diagram of the industrial scenario OT domain device monitoring system provided by an embodiment of the present application; Figure 8 It is a schematic diagram of an industrial control host for implementing the industrial scenario OT domain device monitoring method provided by an embodiment of the present application; Figure 9 It is the device impact association matrix M constructed by an embodiment of the present application; Figure 10 It is to calculate M in an embodiment of the present application 2 (two-step propagation) calculation matrix. Detailed implementation manners

[0021] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system structures and technologies are presented to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, the detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0022] It should be understood that when used in the specification of the present application and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0023] It should also be understood that the term "and / or" as used in the specification of this application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations.

[0024] As used in the specification of this application and the appended claims, the term "if" can be interpreted, depending on the context, as "when", "once", "in response to determining", or "in response to detecting". Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be interpreted, depending on the context, as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]".

[0025] In addition, in the description of the specification of this application and the appended claims, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and should not be construed as indicating or implying relative importance.

[0026] Reference to "one embodiment" or "some embodiments" or the like described in the specification of this application means that a specific feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in another way. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in another way.

[0027] The embodiments of the present application provide an industrial scenario OT domain device monitoring method and system. In a possible implementation, the method can be applied to an industrial control host. The industrial control host has: a physical topology map construction module for constructing a production line-level physical topology map based on the physical connection relationships between devices; a communication link diagram construction module for constructing a communication link between key devices based on the communication connection relationships between the key devices and generating a communication link diagram; a monitoring and analysis module for performing an impact range analysis based on the communication connection relationships between devices shown in the communication monitoring map when a communication interruption occurs between the key devices. The monitoring and analysis module can also identify the key devices of the production line / business line by using a dynamic evaluation model based on service weight coefficients, mark them on the above physical topology map and / or communication link diagram, and perform communication fault location and traceability analysis in combination with the physical topology relationship between the devices. A global map database: for storing the global communication monitoring map formed by integrating the communication link diagrams and physical topology maps of each production line in the OT domain; and realizing panoramic monitoring of the communication status of OT domain devices through a visual interaction interface.

[0028] In an embodiment of the present invention, the device monitoring in the industrial scenario OT domain is carried out by means of physical-communication two-dimensional modeling to perform panoramic monitoring of the communication status of OT domain devices. Combining the use of innovative technologies, it effectively solves the industry pain points such as "difficult fault location", "slow impact assessment", and "weak panoramic perception" in the communication monitoring of industrial OT domain devices, and has significant technological progress and industrial application value. Specifically, as shown in Figure 1, the main working steps of the monitoring method include: S1. Identify each production line / business line in the industrial scenario OT domain; S2. Determine the set of devices included in each production line / business line; S3. Construct a production line-level physical topology map based on the physical connection relationships between devices; S4. Construct a communication link between devices based on the communication connection relationships between devices and generate a communication link diagram; S5. Use a dynamic evaluation model based on service weight coefficients to identify the key devices of the production line / business line and mark them on the above physical topology map and / or communication link diagram; S6. Integrate the communication link diagram and the physical topology map to establish a communication monitoring map; S7. Repeat S2 - S6 to process all production lines / business lines in the OT domain; S8. Integrate the communication link diagrams and physical topology maps of each production line in the OT domain to form a global communication monitoring map; S9. When a communication interruption occurs between the key devices, perform an impact range analysis based on the communication connection relationships between devices shown in the communication monitoring map; S10. Perform communication fault location and traceability analysis in combination with the physical topology relationship between the devices. S11. Achieve panoramic monitoring of the communication status of OT domain devices based on the global communication monitoring map.

[0029] The following is a detailed description of each step: S1. Identify each production line / business line in the OT domain of the industrial scenario: Before performing panoramic monitoring of the communication status of OT domain devices, it is first necessary to sort out the business lines / production lines in the OT domain, that is, to figure out how many production lines or business lines there are in the domain. A comprehensive sorting out of the business lines / production lines included in the OT domain is the basis for achieving complete panoramic monitoring of the communication status of OT domain devices.

[0030] In one embodiment, the system obtains production line layout data through the Equipment Asset Management System (EAM) and establishes a mapping relationship table between the production line and PLC controllers, sensors, and actuators; or parses the namespace path of the OPC UA server (such as ns=3;s=Line1.Robot1) to extract the production line identifier, so as to identify each production line / business line in the OT domain.

[0031] In another embodiment, the system divides each production line through the device IP address grouping logic: for example, by scanning the IP address segments of the devices in the OT domain and combining with the subnet mask to divide the production lines (for example: 192.168.1.0 / 24 is the stamping production line, and 192.168.2.0 / 24 is the welding production line).

[0032] In another embodiment, the division of business lines in the OT domain is achieved by importing production work order data from the MES (Manufacturing Execution System) and binding the device group to the business process (for example: work order number WO2023-001 corresponds to the "body welding" business line).

[0033] As Figure 5 shown, after sorting out, the OT domain devices in this embodiment mainly include production line one and production line two.

[0034] S2. Determine the set of devices included in each production line / business line: In one embodiment, when obtaining the production line identification of S1 through the Equipment Asset Management System (EAM), the mapping relationship and set of devices included in each production line are synchronously obtained.

[0035] In another embodiment, the system polls the device register addresses (such as holding registers 40001 - 40010) through the Modbus TCP protocol to obtain the device type (PLC, sensor, actuator), firmware version, and communication protocol support list. For devices that do not support standard protocols (such as proprietary protocol devices), protocol conversion is performed through a gateway.

[0036] S3. Construct a production line - level physical topology diagram based on the physical connection relationships between devices: The production line - level physical topology diagram is a topology diagram constructed according to the physical connection relationships of each device in the production line. For example, Figure 2 is the schematic diagram of the physical topology relationship of production line one in the above - mentioned embodiment.

[0037] In one embodiment, the physical topology diagram in step S3 includes annotations of topology parameters such as device type, physical interface parameters, transmission medium type, and topology structure characteristic parameters. These topology parameters help to clarify the physical connection characteristics of production line devices and facilitate the implementation of various additional functions in subsequent global monitoring.

[0038] In one embodiment, the above - mentioned various topology parameters are automatically obtained by the system. For example, a dedicated industrial gateway with information collection functions is used. This method of automatically obtaining topology parameters by the system is applicable to scenarios where there are many devices and complex physical connections in the production line: Device type: PLC (Siemens S7 - 1500), industrial switch (Hirschmann RS30), temperature sensor (PT100), motor drive (ABB ACS880); Physical interface parameters: Profinet port (RJ45 - M12), RS485 interface impedance (120Ω ± 5%), fiber optic connector (IP67 protection level); Transmission medium type: industrial - grade twisted pair (CAT6e - SF / UTP), armored optical cable (G.657A2), CAN bus cable (ISO 11898 standard); Topology structure characteristics: ring - type redundant network (main backbone ring network + star - type access of devices).

[0039] In one embodiment, in addition to manually establishing a physical topology diagram by the system administrator according to the physical connection relationships of production line devices, the system can also generate a physical topology diagram represented by an adjacency list based on a depth - first search (DFS) traversal of device connection relationships, further reducing the workload of the system administrator.

[0040] S4. Construct a communication link between devices based on the communication connection relationships between devices and generate a communication link diagram; The communication link diagram mainly records the communication connection topology among the production line devices. In one embodiment, the communication link diagram performs five-tuple analysis (source IP, destination IP, port, protocol, timestamp) on the TCP / UDP sessions among the devices, and uses a computer by constructing a communication adjacency matrix method to automatically and quickly create a communication link diagram.

[0041] In another embodiment, the system administrator manually creates a communication link diagram according to the communication connection relationship of the production line devices.

[0042] S5. Adopt a dynamic evaluation model based on service weight coefficients to identify the key devices of the production line / service line, and mark them on the above physical topology diagram and / or communication link diagram: In the OT domain device monitoring of this application, the accurate identification of key devices is very crucial. After clarifying the key devices, during the communication construction design stage, the links between key devices can be designed with stronger robustness, and various methods such as redundancy and backup can be adopted in advance to enhance the link reliability; on the other hand, after clarifying the key devices, higher-level monitoring can be given during daily device and network monitoring, and the normal operation of production can be more efficiently guaranteed without increasing the overall monitoring cost.

[0043] In one embodiment of this application, the identification of key devices adopts a dynamic evaluation model based on service weight coefficients, which avoids the deviation of manual experience and makes the identification of key devices more accurate. The specific dynamic evaluation model is as follows: The dynamic evaluation model for identifying key devices based on service weight coefficients, and the weight coefficients include three dimensions: equipment downtime impact index, service relevance, and data throughput. The dynamic evaluation model uses the following formula to calculate the device key value K: K = α×P + β×R + γ×T; Where, P is the downtime impact index, R is the service relevance, T is the data throughput, and α, β, γ are dynamic adjustment coefficients.

[0044] The definitions and acquisition methods of each dimension are as follows: 1. Downtime impact index P: P = (number of equipment failures × average repair time) × 100% / total production line operation time; For example: Import the historical alarm log from the SCADA system, and it is counted that a certain PLC has 3 monthly failures, the average repair time is 30 minutes, and the monthly operation time of the production line is 720 hours, then P = 3×0.5×100% / 720 = 0.21%, It can be seen that for the production line equipment, the larger the value of P, the greater the impact of the downtime of this equipment on the operation of the entire production line.

[0045] 2. Service relevance R: The business association degree is an important indicator for clarifying the association relationship between devices. It can be calculated by calculating the number of associations of devices in the BOM (Bill of Materials). The formula is as follows: R = log(1 + ∑ BOM association times); For example: If a certain motor is cited 15 times in the BOM, then: R = log(1 + 15) = 2.77; It can be seen that for production line devices, the larger the value of R, the closer the association between the device and other devices on the production line, and the greater the impact range caused when a failure occurs.

[0046] Data throughput T: The key devices in the production line are the core devices in the production line. Usually, a large amount of data interaction is required to coordinate with other devices or report production data in a timely manner. Therefore, the data throughput per unit time of the device, that is, the number of bytes transmitted per second by the device, is used as one of the evaluation indicators for key devices.

[0047] In an example, in order to remove the deviation caused by different production line device types, data transmission protocols or types, different protocols are weighted to unify the statistical caliber (for example: OPC UA weight = 1.2, Modbus = 1.0).

[0048] Dynamic adjustment coefficients (α, β, γ): In order to improve the adaptability of the dynamic model, the dynamic adjustment coefficients are used to dynamically adjust according to the production stage (such as normal / peak period). By adjusting the weights of the above three dimensions of device downtime impact index, business association degree, and data throughput, the identification of key devices can be made more in line with the actual business requirements.

[0049] For example, during the peak period, the α weight is increased (α = 0.5 → 0.7) to emphasize the downtime impact.

[0050] Using the dynamic evaluation model of the present application can conveniently and accurately identify key devices. Based on the dynamic evaluation mechanism of multi-dimensional weight coefficients, it overcomes the disadvantages of the traditional static threshold determination method that cannot adapt to the dynamic adjustment of production plans, and the problem of inaccurate identification of key devices caused by relying on manual experience, improves the accuracy of key device identification, and enhances the overall monitoring effect.

[0051] In one embodiment, the key equipment of the production line / business line identified by the above-mentioned dynamic evaluation model based on the service weight coefficient is marked on the physical topology map completed in step S3 and / or the communication link map completed in step S4, and these marks are retained in subsequent steps. Such an operation ensures that when the communication monitoring map is merged for all production lines / business lines within the entire OT domain, the problem of distorted identification of key equipment caused by different requirements and business priorities of each production line is avoided.

[0052] S6. Integrate the communication link map with the physical topology map to establish a communication monitoring map: Integrate the physical topology map completed in step S3 above and the communication link map completed in step S4 to form a communication monitoring map of production line 1 as shown in Figure 3 Figure. This map distinguishes the physical topology connection relationship and the communication link connection relationship through different line types, and is a two-dimensional modeling communication monitoring map with a physical-communication dual view.

[0053] S7. Repeat S2 - S6 for all production lines / business lines within the OT domain: Repeatedly use S2 - S6 for all production lines / business lines within the OT domain to produce communication monitoring maps with a physical-communication dual view for all production lines / business lines.

[0054] S8. Integrate the communication link maps and physical topology maps of each production line within the OT domain to form a global communication monitoring map: In one embodiment, as shown in Figure 6 Figure, integrate the communication link maps and physical topology maps of production line 1 and production line 2 within the OT domain, and adopt a progressive construction method from the production line level to the global level to form a global communication monitoring map. It can be seen from the constructed global communication monitoring map that the physical topology relationship between devices cannot accurately reflect all the communication relationships between devices. When the number of devices within the scope of the business to be monitored is extremely large (for example, in the OT domain of a medium-sized factory, there are usually 150 to 250 production devices), it is difficult to form a monitoring of the business chain and the core business scope when simply monitoring the physical topology map or conducting single-device status monitoring and analysis. It is easy to cause problems such as omission or inaccurate handling of event priorities, thereby expanding the losses brought by single-device problems to the overall situation. However, the method of the present application focuses on the overall situation, constructs a set of communication relationships between devices, and conducts analysis. It can control the overall situation in a complete "link" or even "network" manner. On the basis of single-device monitoring and analysis, it expands the perspective, increases the dimension, and distinguishes priorities, so as to be responsible for the overall interests and provide an idea for the global analysis of the OT domain.

[0055] S9. When a communication interruption occurs between the key devices, perform an impact scope analysis based on the communication connection relationships between devices shown in the communication monitoring map: In an embodiment of the present application, the industrial protocol supported by the OT domain device monitoring system for industrial scenarios includes industrial protocols such as OPC UA, Modbus TCP, and Profinet.

[0056] In an embodiment of the present application, the communication monitoring map includes a monitoring and analysis module and a visualization interaction interface. The monitoring and analysis module is used to perform data analysis and status monitoring of the OT domain device monitoring system. The main functions completed include, but are not limited to: using the dynamic evaluation model based on service weight coefficients described above to identify the key devices of the production line / service line, and completing the annotation of key devices on the above physical topology map and / or communication link map; performing an impact scope analysis based on the communication connection relationships between devices shown in the communication monitoring map; combining the physical topology relationships between devices to perform communication fault location and traceability analysis, etc.

[0057] In one embodiment, when the monitoring and analysis module detects a communication interruption between the key devices, it performs an impact scope analysis based on the communication connection relationships between devices shown in the communication monitoring map. Specifically, it combines the communication connection relationships between devices shown in the communication monitoring map and the impact propagation model based on a directed graph to perform the impact scope analysis, as follows: In this embodiment, the impact scope analysis uses an impact propagation model based on a directed graph to establish a device impact association matrix: M ij = 1 indicates that a failure of device i will affect device j, M ij = 0 indicates that a failure of device i does not affect device j.

[0058] Suppose there are 4 devices (A, B, C, D) in a power system, and the fault propagation relationships between devices are as follows (directed graph): A failure of device A will directly affect B and C.

[0059] A failure of device B will directly affect C and D.

[0060] A failure of device C will directly affect D.

[0061] A failure of device D does not affect other devices.

[0062] Step S901: Construct a device impact association matrix M; According to the definition: M ij = 1 indicates that a failure of device i will affect device j, M ij = 0 means that the failure of device i does not affect device j.

[0063] Construct a matrix as Figure 9 shown in the device impact correlation matrix M (the rows represent the faulty source device i, and the columns represent the affected device j): Explanation: The first row (device A): The failure of A directly affects B (M AB = 1) and C (M AC = 1). The second row (device B): The failure of B directly affects C (M BC = 1) and D (M BD = 1). The third row (device C): The failure of C directly affects D (M CD = 1). The fourth row (device D): The failure of D does not affect any device, all are 0.

[0064] Step S902: Analyze the fault propagation range: Assume that device A fails. According to matrix M: 1. Direct impact: A → B and A → C (reflected by M AB = 1 and M AC = 1).

[0065] 2. Indirect impact (cascading propagation): After B is affected, B → C and B → D (reflected by M BC = 1 and M BD = 1).

[0066] After C is affected, C → D (reflected by M CD = 1).

[0067] 3. Final impact range: A → B → C → D and A → C → D, all devices are affected.

[0068] Step S903: Mathematically represent the propagation path: Through matrix power operation M K the k-step indirect impact can be analyzed: Direct impact: Represented by M itself.

[0069] Indirect impact: Represented by M 2 .M 3 ......

[0070] For example, calculate M 2(Two-step Propagation) As shown in Figure 10 : M 2 AD = 2 means starting from A, D can be affected within two steps through two paths (A → B → D and A → C → D).

[0071] The following problems can be solved through the incidence matrix M: 1) Identification of key equipment in the influence scope: If a device has many non-zero elements in its row, its fault influence scope is large (such as devices A and B).

[0072] 2) Maintenance priority: Protect high-influence devices first (such as repairing A can avoid cascading failures of B, C, and D).

[0073] 3) Fault isolation strategy: If D fails, by tracing back according to M, it can be found that the source may be B or C.

[0074] Therefore, the incidence matrix M of this embodiment intuitively depicts the direct relationship of fault propagation between devices, and through matrix operations (such as power operations, etc.), the cascading influence scope can be quantified, which is suitable for automatically and accurately estimating the influence scope through automation, improving the response speed and accuracy of influence scope warning. By using the communication connection relationship between devices shown in the communication monitoring graph, analyzing the influence scope of communication interruption between key devices can make perfect use of the natural advantage that the nodes / hop points of the network in the communication monitoring graph can be conveniently retrieved from the network connection log, quickly establish the above-mentioned incidence matrix, and the accurate and comprehensive influence scope analysis of key devices can make the response to communication interruption more targeted and efficient.

[0075] S10. Combine the physical topology relationship between devices to perform communication fault location and traceability analysis: In one embodiment, a reverse tracing algorithm is used to reverse-search for potential fault points along the physical topology path and generate a traceability path report containing fault probabilities.

[0076] The specific steps are as follows: Step S1001: Fault triggering and feature extraction: For example, when the 3# motor drive of a certain rolling mill production line frequently reports communication timeout faults: The PLC detects that the Profinet bus cycle is extended to 15 ms (standard value 4 ms ± 1 ms), Through the physical topology relationship between devices: Lock the fault influence path as: Motor drive ACS880-03: M12 interface → Fieldbus Segment-C (length 82 m, including 3 T-joints): Port P5 → Industrial switch SW-07: Fiber optic ring network → Master PLC S7-1500-01.

[0077] Step S1002: Physical path reverse troubleshooting: Construct a reverse detection path: Fault symptom layer: ACS880-03 drive (communication timeout rate 38%) ← Field layer: Bus Segment-C (measured impedance fluctuation ±8 Ω) ← Network layer: SW-07 switch (packet error rate at port P5 0.7%) ← Control layer: S7-1500-01 PLC (CPU load rate 92%).

[0078] Step S1003: Multi-dimensional fault probability assessment: Calculate the fault possibility according to the characteristics of industrial equipment. The results are shown in Table 1: Table 1 Analysis table of calculating fault possibility based on industrial equipment characteristics Troubleshooting node Evaluation index Weight Probability Bus Segment-C connector oxidation + length exceeding standard 0.55 78% SW-07 Switch port buffer overflow + abnormal heat dissipation 0.30 65% PLC controller CPU overload + memory occupancy exceeding limit CPU overload 0.15 42% Step S1004: Generate an industrial-level traceability report: Output a special report containing the following elements: 1) Topology heat map: Mark the high-risk path of bus Segment-C → switch P5 in red; 2) Repair priority: Urgently replace the T-joints of bus Segment-C (probability 78%), upgrade the heat dissipation system of SW-07 switch (probability 65%); 3) And calculate the influence range according to Step S9: Associated with 12 devices and 3 production line control signals.

[0079] It can be seen that this example solves the problem of low fault location efficiency in the current OT domain device monitoring by establishing a reverse tracking path from the physical layer to the service layer and through a cross-layer traceability positioning mechanism. When monitoring OT domain devices, only network traffic is collected through the SNMP protocol, without considering the physical connection relationship of the devices. It realizes multi-dimensional positioning of the root cause of the fault and greatly reduces the average fault location time.

[0080] S11. Implement panoramic monitoring of the communication status of OT domain devices based on the global communication monitoring map: In an embodiment of the present application, the global communication monitoring map uses a visual interaction interface to perform panoramic monitoring of the communication status of OT global devices. For example: Push the device status in real time through WebSocket, and in order to highlight the status attention of key devices, the system also supports filtering alarms according to the above device key value K (devices with K>0.7 are highlighted). OrFigure 4 As shown, in the communication monitoring map, a fault icon is used to indicate that a device has an abnormality.

[0081] In another embodiment, the system further includes a pre - plan recommendation module, which is used to, when a communication interruption is detected, be able to match similar cases based on the historical fault library and timely recommend an emergency plan.

[0082] For example, the historical fault library stores structured case data, and the included field categories include "fault characteristics", "disposal solutions", "spare part information", "list of affected devices", "historical disposal records", etc., which are multiple pieces of information representing fault information and corresponding coping strategies.

[0083] The pre - plan recommendation module further includes an emergency plan recommendation engine. The emergency plan recommendation engine analyzes the encountered fault information and automatically matches it with the "fault characteristics" in the historical fault library in multiple dimensions, so that the fault is accurately matched with the historical fault library, thereby giving an accurate plan. In one embodiment, the dimensions include: Device type matching: The historical fault disposal plan of the same model PLC (weight 60%); Environmental parameter correlation: The case of connector oxidation caused by high humidity (>80%RH) (weight 25%); Communication protocol characteristics: The comparison of abnormal characteristics between Profinet and Modbus TCP (weight 15%).

[0084] Based on the same technical concept, as Figure 7 shown, the embodiment of the present application also provides an industrial scenario OT - domain device communication monitoring system. For the sake of convenience of description, only the parts related to the embodiment of the present application are shown. The device includes: Physical topology map construction module: used to construct a production - line - level physical topology map based on the physical connection relationship between devices; Communication link map construction module: based on the communication connection relationship between the key devices, construct a communication link between the key devices and generate a communication link map; Monitoring and analysis module: adopt a dynamic evaluation model based on service weight coefficients to identify the key devices of the production line / business line, and mark them on the above - mentioned physical topology map and / or communication link map. When a communication interruption occurs between the key devices, perform an impact range analysis based on the communication connection relationship shown in the communication monitoring map; combine the physical topology relationship between the devices to perform communication fault location and traceability analysis; Global map database: used to store the global communication monitoring map formed by integrating the communication link maps and physical topology maps of each production line in the OT domain; Visual interaction interface: Provide a visual interaction interface for realizing panoramic monitoring of the communication status of OT domain devices based on the above-mentioned global communication monitoring map.

[0085] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not imply the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0086] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units, due to being based on the same concept as the method embodiments of the present application, for their specific functions and the technical effects brought, reference can be specifically made to the method embodiment part, and details will not be elaborated here.

[0087] Those skilled in the art can clearly understand that for the convenience and conciseness of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiments, and details will not be elaborated here.

[0088] As Figure 8 shown, the embodiments of the present application also provide an industrial control host. The industrial control host device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the industrial control host program, it implements the steps in any of the above method embodiments.

[0089] The embodiments of the present application provide a computer program product. When the computer program product runs on a terminal, it enables the terminal to execute and implement the steps in any of the above method embodiments.

[0090] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-mentioned embodiment methods of this application, a computer program can be used to instruct relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the photographing device / target device, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.

[0091] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0092] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0093] In the embodiments provided in this application, it should be understood that the disclosed device / network device and method can be implemented in other ways. For example, the device / network device embodiments described above are only illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0094] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0095] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A method for monitoring OT domain equipment in industrial scenarios, characterized in that: The monitoring method comprises the following steps: S1. Identify each production line / business line in the OT domain of the industrial scenario; S2. Determine the equipment set included in each production line / business line; S3, constructing a production line-level physical topology map based on the physical connection relationship between devices; S4. Building a communication chain between devices based on the communication connection relationship between devices and generating a communication link diagram; S5. Using a dynamic evaluation model based on business weight coefficients, identify key equipment of the production line / business line and mark them on the above-mentioned physical topology diagram and / or communication link diagram; S6, integrating the communication link map with the physical topology map to establish a communication monitoring map; S7. Repeat S2-S6 to process all production lines / business lines in the OT domain; S8. Integrate the communication link diagram and physical topology diagram of each production line in the OT domain to form a global communication monitoring map; S9. When it is detected that communication interruption occurs between the key devices, an impact range analysis is performed based on the communication connection relationship between the devices shown in the communication monitoring graph; S10, performing communication fault location and source tracing analysis based on the physical topology relationship between the devices; S11. Based on the global communication monitoring map, a panoramic monitoring of the communication status of OT domain devices is realized.

2. The industrial scene OT domain equipment monitoring method according to claim 1 is characterized in that: The physical topology map in step S3 includes device type, physical interface parameters, transmission medium type and topology structure characteristic parameters.

3. The industrial scene OT domain equipment monitoring method according to claim 1 is characterized in that: The key equipment identification in step S5 adopts a dynamic evaluation model based on business weight coefficients, and the weight coefficients include three dimensions: equipment downtime impact index, business relevance, and data throughput.

4. The industrial scene OT domain equipment monitoring method according to claim 3 is characterized in that: The dynamic evaluation model uses the following formula to calculate the key value K of the equipment: K = α × P + β × R + γ × T, Among them, P is the downtime impact index, R is the service relevance, T is the data throughput, and α, β, and γ are dynamic adjustment coefficients.

5. The industrial scene OT domain equipment monitoring method according to claim 1 is characterized in that: The communication monitoring map includes a monitoring and analysis module and a visual interactive interface, and supports protocols including OPC UA, Modbus TCP, and Profinet industrial protocols.

6. The industrial scene OT domain equipment monitoring method according to claim 1 is characterized in that: The impact range analysis in step S9 adopts an impact propagation model based on a directed graph to establish a device impact association matrix: M ij =1 means that the failure of device i will affect device j, M ij =0 means that the failure of device i does not affect device j.

7. The industrial scene OT domain equipment monitoring method according to claim 1 is characterized in that: The traceability analysis in step S10 adopts a reverse tracing algorithm to reversely search for potential fault points along the physical topology path and generate a traceability path report including the fault probability.

8. The industrial scene OT domain equipment monitoring method according to claim 1 is characterized in that: The method also includes an emergency plan recommendation module, which, when a communication interruption is detected, matches similar cases based on a historical fault library and recommends an emergency plan, the plan including spare parts information, disposal procedures, and a list of affected equipment.

9. An industrial scene OT domain equipment communication monitoring system, characterized in that: include: Physical topology construction module: used to construct a production line-level physical topology based on the physical connection relationship between devices; Communication link diagram construction module: constructs the communication chain between key devices based on the communication connection relationship between key devices and generates a communication link diagram; Monitoring and analysis module: adopts a dynamic evaluation model based on business weight coefficients to identify the key equipment of the production line / business line and mark them on the above-mentioned physical topology map and / or communication link map. When a communication interruption is detected between the key equipment, the impact range is analyzed based on the communication connection relationship between the equipment shown in the communication monitoring map; communication fault location and source tracing analysis are performed in combination with the physical topology relationship between the equipment; Global map database: used to store the global communication monitoring map formed by integrating the communication link map and physical topology map of each production line in the OT domain; Visual interactive interface: provides a visual interactive interface for panoramic monitoring of the communication status of OT domain devices based on the global communication monitoring map.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Sliding revolver pistol

    WO2020023001A1

  • Fault diagnosis method based on multi-device collaborative complex system

    CN116781500A

  • Communication complex network service end-to-end intelligent diagnostic analysis method

    CN116800587A

  • Industrial internet cross-domain zero-trust abnormal flow detection system and method

    CN118631533A

  • Transmission network fault positioning method, system and device and storage medium

    CN119520248A