Industrial control equipment intelligent operation and maintenance method and system based on large model
By constructing a distributed monitoring architecture based on a large model and cross-channel collaborative analysis, the problem of insufficient early warning capability in the industrial control equipment monitoring system was solved, realizing real-time status tracking and automated operation and maintenance management of industrial control equipment, and improving safety response speed and operation and maintenance efficiency.
Patent Information
- Application Number
- CN202511493029.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-20
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-10-20
AI Technical Summary
Existing industrial control equipment monitoring systems rely on traditional rules and static thresholds, resulting in insufficient early warning capabilities for complex security events, inability to respond to potential threats in a timely manner, and impact on the safe operation of the system.
A distributed monitoring architecture based on a large model is constructed. Time-series data is collected in real time through network situation monitoring nodes. Cross-channel collaborative analysis is carried out using long and short time-series feature prediction channels and network security event identification channels to generate a network security event risk index. When the threshold is reached, operation and maintenance instructions are automatically issued.
It enables real-time, continuous status tracking and automated operation and maintenance management of industrial control equipment, improving response speed and operation and maintenance efficiency, reducing manual intervention, and ensuring the stability and security of equipment and networks.
Smart Images

Figure CN120956542A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of network monitoring technology, specifically to a method and system for intelligent operation and maintenance of industrial control equipment based on a large model. Background Technology
[0002] In modern industrial environments, an increasing number of industrial control devices (ICS) are connected to enterprise networks. The security of these ICS is paramount, as any security attack on an ICS can severely impact production processes, data flow, and the stability of the overall system. However, existing ICS monitoring systems typically rely on traditional rule-based monitoring and static threshold settings, such as setting alarms for excessive network traffic or latency. This approach, however, depends on manually setting thresholds and struggles to effectively capture the interrelationships between devices and complex security events. The limitations of current technology result in insufficient early warning capabilities for complex security events, leading to a failure to respond promptly to potential threats to ICS. This can cause equipment malfunctions, the spread of network attacks, and even disruption to the entire production system. Summary of the Invention
[0003] This application provides a method and system for intelligent operation and maintenance of industrial control equipment based on a large model. It aims to solve the technical problem that existing industrial control equipment monitoring usually relies on traditional rule-based monitoring and static threshold settings, which leads to insufficient early warning capabilities for complex security events. Consequently, industrial control equipment cannot respond in a timely manner when encountering potential threats, affecting the safe operation of the entire system.
[0004] The first aspect disclosed in this application provides an intelligent operation and maintenance method for industrial control equipment based on a large model. The method includes: deploying monitoring nodes according to the distribution information of the industrial control equipment to obtain a distributed industrial control equipment monitoring architecture, wherein the distributed industrial control equipment monitoring architecture includes P network status monitoring nodes for P industrial control equipment, where P is a positive integer; during the network communication process of the P industrial control equipment, collecting P network status awareness time-series data according to a preset monitoring window based on the P network status monitoring nodes; inputting the P network status awareness time-series data into an industrial control equipment network security event prediction model for cross-channel collaborative analysis to generate P industrial control equipment network security event prediction results, wherein the industrial control equipment network security event prediction model includes a long and short time-series feature prediction channel and a network security event identification channel, and each industrial control equipment network security event prediction result includes a network security event risk index; when the network security event risk index of any industrial control equipment reaches a network security event risk index threshold, issuing an operation and maintenance instruction, and executing the operation and maintenance management of the arbitrary industrial control equipment based on the operation and maintenance instruction.
[0005] The second aspect of this application discloses an intelligent operation and maintenance system for industrial control equipment based on a large model. The system is used in the aforementioned intelligent operation and maintenance method for industrial control equipment based on a large model. The system includes: a monitoring architecture establishment module, used to deploy monitoring nodes according to the distribution information of the industrial control equipment, and obtain a distributed industrial control equipment monitoring architecture, wherein the distributed industrial control equipment monitoring architecture includes P network status monitoring nodes for P industrial control equipment, where P is a positive integer; and a perception data acquisition module, used to collect P network status perception time-series data according to a preset monitoring window based on the P network status monitoring nodes during the network communication process of the P industrial control equipment. The cross-channel collaborative analysis module is used to input the P network situational awareness time-series data into the industrial control equipment network security event prediction model for cross-channel collaborative analysis, generating P industrial control equipment network security event prediction results. The industrial control equipment network security event prediction model includes long and short time-series feature prediction channels and network security event identification channels. Each industrial control equipment network security event prediction result includes a network security event risk index. The operation and maintenance management module is used to issue operation and maintenance instructions when the network security event risk index of any industrial control equipment reaches the network security event risk index threshold, and to execute the operation and maintenance management of the arbitrary industrial control equipment based on the operation and maintenance instructions.
[0006] One or more technical solutions provided in this application have at least the following beneficial effects: By deploying monitoring nodes according to the distribution information of industrial control equipment (ICS), a distributed monitoring architecture is constructed. This architecture has strong scalability, allowing for flexible addition of monitoring nodes based on actual needs. It comprehensively covers all ICS connected to the enterprise network, with each monitoring node independently responsible for monitoring one device or a group of devices, collecting and processing relevant data in real time. By collecting network situational awareness time-series data from P ICS devices based on the monitoring nodes, the status of each ICS device can be tracked continuously in real time. This time-series data reflects the real-time operating status of the devices and provides important information on network security situation. The collected time-series data is input into the ICS network security event prediction model for cross-channel collaboration. Through analysis, the system extracts the long and short time-series characteristics and security event identification features of the devices, generating a network security event risk index for each device. This process utilizes long short-term memory networks and recurrent neural networks to achieve in-depth analysis and prediction of device status. When the network security event risk index of a certain industrial control device reaches a preset risk index threshold, an automatic maintenance command is issued to execute corresponding maintenance management operations to repair device malfunctions or mitigate security risks. This automated maintenance management mechanism significantly reduces the need for manual intervention, improves response speed and maintenance efficiency, and enables rapid response when devices malfunction, preventing security events from spreading to other devices or systems and ensuring the stability and security of devices and networks.
[0007] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description
[0008] Figure 1 A schematic diagram of the intelligent operation and maintenance method for industrial control equipment based on a large model provided in this application embodiment.
[0009] Figure 2 This is a schematic diagram of the structure of an intelligent operation and maintenance system for industrial control equipment based on a large model, provided in an embodiment of this application.
[0010] Figure labeling: Monitoring architecture establishment module 10, perception data acquisition module 20, cross-channel collaborative analysis module 30, operation and maintenance management module 40. Detailed Implementation
[0011] This application provides an intelligent operation and maintenance method and system for industrial control equipment based on a large model. It solves the technical problem that the monitoring of industrial control equipment in the prior art usually relies on traditional rule-based monitoring and static threshold settings, which leads to insufficient early warning capability for complex security events. Consequently, the industrial control equipment cannot respond in time when it encounters potential threats, affecting the safe operation of the entire system.
[0012] After introducing the basic principles of this application, various non-limiting embodiments of this application will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit this application.
[0013] Example 1, as Figure 1 As shown in the embodiments of this application, an intelligent operation and maintenance method for industrial control equipment based on a large model is provided, the method including: Based on the distribution information of industrial control equipment, monitoring nodes are deployed to obtain a distributed industrial control equipment monitoring architecture, which includes P network status monitoring nodes for P industrial control equipment, where P is a positive integer.
[0014] The deployment of monitoring nodes is based on the geographical distribution and operational requirements of industrial control equipment. The distribution information of industrial control equipment refers to the location, connection method, and functional requirements of industrial control equipment in the entire industrial environment. These industrial control equipment are located in different areas or working environments. Therefore, when deploying monitoring nodes, it is necessary to comprehensively consider factors such as the interconnection, communication paths, and data flow between industrial control equipment. In this process, the monitoring needs of each industrial control equipment are evaluated, and effective monitoring node deployment is carried out according to its specific tasks, determining the number, location, and communication paths between monitoring nodes.
[0015] The distributed industrial control equipment monitoring architecture is a networked system containing multiple monitoring nodes to continuously track and monitor the status of industrial control equipment. This distributed architecture means it consists of multiple independent but collaborative nodes connected via a network, sharing data and information. Each node is responsible for monitoring different devices or different areas. The architecture includes P network situational awareness nodes for P industrial control devices, where P is a positive integer representing the number of devices to be monitored. Each industrial control device obtains real-time network situational awareness data, such as device status, data traffic, and anomaly detection, through the monitoring nodes to analyze and monitor the overall network.
[0016] During the networking communication process of the P industrial control devices, P network situation awareness time-series data are collected and acquired based on the P network situation monitoring nodes according to a preset monitoring window.
[0017] When these industrial control devices communicate via a network, they exchange a large amount of data, including device operating status, warning signals, and abnormal data. During networking, each industrial control device exchanges data with monitoring nodes. Each monitoring node continuously collects network status data according to a preset monitoring window. The preset monitoring window refers to a time window, measured in seconds, minutes, or hours, used to set the frequency of data collection. Network status awareness time-series data refers to the network status data of industrial control devices captured over time, including parameters such as device operating status, network latency, traffic changes, error warnings, and load. This data helps to monitor the health status of devices in real time and identify potential network security risks. During data collection, monitoring nodes monitor the network data of each industrial control device in real time and capture various time-series data related to the device. This data is continuous in the time dimension, thus reflecting the dynamic characteristics of device operation changes and system status.
[0018] The P network situational awareness time-series data are input into the industrial control equipment network security event prediction model for cross-channel collaborative analysis to generate P industrial control equipment network security event prediction results. The industrial control equipment network security event prediction model includes long and short time-series feature prediction channels and network security event identification channels. Each industrial control equipment network security event prediction result includes a network security event risk index.
[0019] The collected P network situational awareness time-series data are used as input data and fed into the industrial control equipment network security event prediction model. The purpose of this prediction model is to analyze these time-series data and predict possible network security events, especially potential risks to equipment. The long and short time-series feature prediction channel aims to capture the behavioral patterns and state changes of equipment by analyzing the long and short-term characteristics of the data. These long and short-term characteristics include traffic change trends and load fluctuations of equipment within different time windows. By using time series analysis methods, the model can identify potential patterns in equipment behavior. The network security event identification channel focuses on identifying network security events. It detects potential security events by analyzing abnormal patterns in the data. For example, sharp fluctuations in network traffic and abnormal communication between devices can serve as indicators of network security events. This channel uses machine learning algorithms to identify abnormal events and assigns a corresponding label to each security event.
[0020] Based on the above analysis, the model generates cybersecurity incident prediction results for each industrial control device. These results include a cybersecurity incident risk index for each device, which is a quantitative value representing the likelihood and severity of a security incident affecting the device. For example, if a device has a high risk index, it means that the device is in a potentially risky state and may face threats such as cyberattacks, data breaches, or other security incidents.
[0021] When the network security event risk index of any industrial control device reaches the network security event risk index threshold, an operation and maintenance instruction is issued, and the operation and maintenance management of the arbitrary industrial control device is executed based on the operation and maintenance instruction.
[0022] Set a network security incident risk index threshold. If the risk index of an industrial control device exceeds the threshold, the device is considered to be facing a security incident and needs to be dealt with in a timely manner. For example, if the threshold is set to 0.8, it means that when the risk index of a device exceeds 0.8, the device needs further inspection or processing. The threshold setting is based on historical data, the security level of the device, and the overall security requirements of the network.
[0023] When the risk index of a device exceeds a threshold, an automatic maintenance command is issued. This command is a direct operational instruction for the device, such as closing a port, initiating a security remediation program, or adjusting device configuration. The issuance of these commands is automated, dynamically responding to changes in the risk index without human intervention, ensuring timely handling of potential security threats. Upon receiving the maintenance command, corresponding maintenance management operations are performed on the target industrial control equipment. Through this series of operations, the network security status of the industrial control equipment can be monitored in real time, and automated response measures can be taken when potential security threats are detected, ensuring the security of the entire industrial environment.
[0024] Furthermore, the method for constructing the network security event prediction model for the industrial control equipment includes: According to the preset monitoring window, network situational awareness data is backtracked for the P industrial control devices to obtain P historical network situational awareness time-series data sets; based on a predetermined long-short-term feature analysis scale, long-short-term features are captured from the P historical network situational awareness time-series data sets to obtain P long-short-term sample groups; supervised training is performed on the P long-short-term sample groups using a long short-term memory network, and when the training loss function converges to a predetermined threshold, the long-short-term feature prediction channel is generated; network security event sample groups and normal network status sample groups are extracted from the P long-short-term sample groups; supervised training is performed on the network security event sample groups and normal network status sample groups using a recurrent neural network, and when the training loss function converges to a predetermined threshold, the network security event identification channel is generated; multi-channel collaborative fusion is performed on the long-short-term feature prediction channel and the network security event identification channel to generate the industrial control device network security event prediction model.
[0025] Based on the preset monitoring window, historical network situational awareness data is traced back. The purpose of tracing back is to use historical data to understand the past operating status and network behavior of devices. This data includes device communication records, traffic data, system load, error warnings, etc. By tracing back this data, P sets of historical network situational awareness time series data are obtained. These sets of data contain the network status of various industrial control devices at different points in time, forming a complete time series dataset.
[0026] The predetermined long-term and short-term feature analysis scales refer to extracting features at different time scales. Short-term features reflect the instantaneous state changes of the equipment, while long-term features reflect the long-term operating trends and behavioral patterns of the equipment. These features can include short-term fluctuations in equipment load, long-term trends in flow rate changes, equipment response time, etc. Through this analysis, P long-term and short-term sample groups are obtained. Each sample group corresponds to the long-term and short-term features of an industrial control device in historical data. These sample groups provide the basic data for subsequent model training.
[0027] Long Short-Term Memory (LSTM) networks are used to process and predict time-series data. They can effectively capture long-term dependencies in time-series data and solve the gradient vanishing problem that may occur in ordinary neural networks during training on long-series data. In this step, P LTM sample groups are used for supervised training of the LTM network. Supervised training means training with labeled data that has been labeled with the past security status and risk situation of the device. The goal of training is to enable the model to learn the operating rules of the device and the probability of network security incidents under specific conditions through this historical data.
[0028] During training, the goal is to minimize the loss function, i.e., reduce the model's prediction error. The loss function is calculated based on the difference between the predicted result and the actual label. Training continues until the value of the loss function converges to a predetermined threshold. This means that the model has learned the potential patterns in the data and the prediction error has reached an acceptable range. When the loss function converges, training is completed and a long and short time series feature prediction channel is generated. This channel is specifically used to predict the long and short time series features of devices from new time series data and to provide a basis for subsequent security event prediction.
[0029] The network security incident sample group contains time-series data of industrial control equipment (ICS) during network security incidents. This data includes abnormal patterns, attack behaviors (such as DDoS attacks, malware, unauthorized access, etc.), and their manifestation in the time-series data. The normal network status sample group contains time-series data of ICS during normal operation, reflecting the normal behavior of the equipment and the network status when no security incidents occur. By classifying these sample groups, it is possible to identify which time-series data corresponds to network security incidents and which represents a normal network status.
[0030] Recurrent Neural Networks (RNNs) are a type of neural network architecture suitable for time-series data. They can capture dependencies within time series data and learn patterns that change over time. RNNs are particularly well-suited for processing network status data from industrial control equipment, which is time-dependent. Supervised training involves inputting network security incident samples and normal network status samples into the RNN. Supervised training means training the model using labeled data (i.e., labels for security incidents and normal states). The goal is for the model to learn to distinguish between these two types of data. During training, the RNN continuously optimizes its weights and minimizes the loss function. The loss function measures the difference between the model's predictions and the actual labels. When the loss function converges to a predetermined threshold, it means the model has learned the patterns in the data and can accurately classify network security incidents from normal states. After training, a network security incident identification channel is generated. This channel can extract features of network security incidents from the input time-series data and perform classification predictions, such as identifying whether there is an attack, equipment failure, or other abnormal behavior.
[0031] Multi-channel collaborative fusion of long and short time series feature prediction channels and network security event identification channels means combining the outputs of the two channels to form a comprehensive prediction model. This model can simultaneously extract long and short time series features and identify network security events. After collaboratively fusing the outputs of these two channels, a network security event prediction model for industrial control equipment is generated. This model can predict the network security status of equipment based on the input time series data and can provide accurate risk indices and prediction results for operation and maintenance decisions.
[0032] Furthermore, the network security incident sample group includes a first set of long and short-term features of industrial control equipment when a network security incident occurs, the first set of long and short-term features having a network security incident label, and the network status normal sample group includes a second set of long and short-term features of industrial control equipment when the network status is normal, the second set of long and short-term features having a network status normal label.
[0033] The network security incident sample set is extracted from time-series data of industrial control equipment during network security incidents. This data contains network posture characteristics of the equipment when it encounters attacks or other security problems. The first set of short- and long-term features captures the dynamic behavior of the equipment during network security incidents. By analyzing this data, abnormal patterns can be identified, such as sudden surges or decreases in network traffic, delays or abnormal fluctuations in equipment response time, etc. These features are extracted from time-series data and include both short-term features (such as instantaneous traffic changes and equipment status fluctuations) and long-term features (such as long-term patterns of equipment behavior). For example, when a device is attacked, it may experience sudden abnormal fluctuations, but it may also exhibit a gradual decline over a long period of time.
[0034] The first set of long and short time features has a cybersecurity event label, which is manually or automatically labeled to indicate whether the data sample belongs to a specific cybersecurity event, such as a cyberattack, data breach, or other anomaly.
[0035] The normal network status sample set is extracted from time-series data of industrial control equipment under normal operating conditions. This data includes the normal operating status of the equipment when no network security incidents occur. The second set of long and short-term features mainly describes the behavioral patterns of the equipment when there are no security incidents, such as the normal fluctuation range of network traffic, normal equipment response time, and load. The second set of long and short-term features has a "normal network status" label, which indicates that this time-series data belongs to the normal state of the equipment when there are no network security incidents. The purpose of the label is to tell the model that this data does not contain any security threats and is normal behavior of the equipment. This normal state data sample helps the model learn the normal behavior of the equipment in a healthy state and distinguish it from the characteristics of abnormal events.
[0036] Furthermore, generating the network security event identification channel also includes: Extract the set of network security event types of industrial control equipment when a network security event occurs from the network security event sample group; perform incremental training on the network security event identification channel based on the set of network security event types and the first long and short time feature set.
[0037] The set of network security event types specifically refers to the different security event types extracted from the sample group. These types can be: attack types, such as DDoS attacks, virus propagation, Trojan intrusion, etc.; device failure types, such as hardware failure, network connection interruption, etc.; and abnormal behavior types, such as data leakage, identity theft, unauthorized access, etc. Each network security event sample is labeled to indicate that the event belongs to a specific security event type. Therefore, the set of network security event types is a set of categories extracted from these labeled data to represent different types of network security events.
[0038] Incremental training refers to continuously updating and optimizing an existing model by adding new features, rather than retraining it from scratch. This allows the model to maintain high accuracy when facing new event types or different security scenarios. In this process, a set of cybersecurity event types and a first set of long and short-term features are used together for incremental training of the cybersecurity event recognition channel. During incremental training, the model continuously adjusts and optimizes its internal parameters based on new cybersecurity event types and corresponding long and short-term feature information. The training process teaches the model how to better identify new security events and improves recognition accuracy by learning different event patterns. Incremental training enables the model to continuously learn these new event types and distinguish them from existing security event types, thereby improving the model's recognition capabilities.
[0039] Furthermore, the method includes: Based on the P sets of historical network situational awareness time-series data, network security events are extracted and aggregated into similar categories to obtain Q sets of similar network security event data, where Q is a positive integer; based on the Q sets of similar network security event data, the duration of similar network security events is extracted to obtain Q sets of duration of similar network security events; duration distribution analysis is performed on the Q sets of duration of similar network security events, and the predetermined long and short-term feature analysis scale is generated based on the analysis results.
[0040] Network security events are extracted from P sets of historical network situational awareness time-series data. These time-series data contain the status information of each industrial control device at different points in time. By analyzing these data, security events can be identified. Similar security events are grouped together to form a set of similar network security event data. This allows for further analysis of the common characteristics of these network security events, avoids data being too scattered, and improves the effectiveness of event analysis. Through aggregation, Q sets of similar network security event data are finally obtained, where Q is a positive integer representing the number of event categories. Each category contains multiple similar security event samples.
[0041] Extract the duration of each event category from a dataset of Q similar network security events. Duration refers to the time interval between the start and end of the event. For example, for device failure, duration is the time between the start of the failure and the device returning to normal.
[0042] A duration distribution analysis is performed on a set of Q similar cybersecurity incidents. Duration distribution analysis examines the temporal distribution of the duration of different incident types. Statistical methods, such as histograms, probability distributions, and quantile analysis, can be used to obtain the distribution characteristics of duration. These analyses can identify common patterns in the duration of different types of events. The predetermined duration feature analysis scale refers to the time window set based on the characteristics of the duration distribution during the analysis process. This window is used for subsequent feature extraction and model training. This scale helps in modeling and predicting events based on their duration.
[0043] Furthermore, the step of collecting P network situational awareness time-series data according to a preset monitoring window also includes: Anomaly detection data identification is performed on the P network situational awareness time series data, wherein anomaly detection includes data missing, data noise, and data duplication; anomaly pattern analysis is performed on the anomaly detection data, and when the anomaly pattern analysis result is a repairable anomaly, the anomaly detection data is corrected; when the anomaly pattern analysis result is an unrepairable anomaly, the anomaly detection data is removed, thus completing the data cleaning of the P network situational awareness time series data.
[0044] Anomaly detection and data identification involves recognizing various anomalies in time-series data, such as missing data, data noise, data duplication, and abnormal data fluctuations. These anomalies can affect the accuracy and effectiveness of subsequent analysis, thus requiring identification and handling. Missing data refers to the absence of valid data at certain points in time or data segments, usually due to equipment failure, network interruption, or data transmission loss. Data noise refers to random fluctuations or errors in the data; these fluctuations typically do not represent actual network status changes but are caused by external interference, sensor errors, or unstable factors during data acquisition. Data duplication refers to the occurrence of identical values or data points at the same time or within a time period, usually caused by errors in the data acquisition system or network synchronization problems.
[0045] Anomaly pattern analysis is performed on the anomalous sensing data to determine whether the anomalous data can be restored to a normal state through correction methods, or whether it needs to be directly removed. Some anomalies are caused by minor errors or temporary problems, such as data noise or temporary network fluctuations. For these anomalies, interpolation, smoothing algorithms, or other correction methods can usually be used to handle them. Common repair methods include interpolation, mean substitution, and data imputation. The repaired data will be restored to a state consistent with the original data as much as possible. Some anomalies may be caused by serious failures, attacks, or data corruption. This anomalous data cannot be repaired and may mislead the analysis results. Therefore, it needs to be removed from the dataset to ensure the quality of the remaining data and the accuracy of the analysis. Data cleaning refers to correcting repairable anomalous data and removing irreparable anomalous data to obtain P clean and reliable network situational awareness time series data. The cleaned data can be better used for subsequent model training, prediction, and analysis, avoiding the negative impact of anomalous data on the results.
[0046] Furthermore, the method also includes: When the network security incident risk index of multiple industrial control devices reaches the network security incident risk index threshold, multi-device joint analysis is performed to generate a linkage response strategy, and global operation and maintenance management is carried out based on the linkage response strategy.
[0047] When the cybersecurity incident risk indices of multiple industrial control devices all reach the cybersecurity incident risk index threshold, a multi-device joint analysis is conducted. The purpose of this analysis is to examine the synergistic effects between devices from a global perspective, identifying potential mutual influences or synergistic effects. For example, security incidents on some devices may affect the operating status of other devices, or simultaneous security attacks on multiple devices may lead to the collapse or serious failure of the entire industrial system. The multi-device joint analysis is based on the interrelationships between devices, including network connections, shared resources and collaborative tasks, and temporal correlations. Based on the results of the multi-device joint analysis, a coordinated response strategy is generated. The goal is to reduce overall cybersecurity risks and ensure device stability and security by coordinating the operational behavior of multiple devices. For example, if multiple devices are subjected to the same type of attack or abnormal event, the repair processes of these devices are coordinated to ensure normal operation is restored in the shortest possible time. Global operational management is then executed according to this coordinated response strategy. Global operational management involves coordinating the operational work of multiple devices throughout the industrial control system to ensure that each device receives a timely and effective response when a cybersecurity incident occurs.
[0048] Furthermore, the step of executing the operation and maintenance management of the arbitrary industrial control equipment based on the operation and maintenance instructions also includes: Based on the operation and maintenance instructions, the operation and maintenance plan of the basic industrial control equipment is retrieved, and the operation and maintenance management of the arbitrary industrial control equipment is carried out. The system performs periodic tracking and evaluation based on the preset monitoring window to obtain a set of operation and maintenance effect indicators for the arbitrary industrial control equipment. Based on the set of operation and maintenance effect indicators for the arbitrary industrial control equipment, the operation and maintenance plan of the basic industrial control equipment is optimized to generate a target industrial control equipment operation and maintenance plan.
[0049] During operation and maintenance (O&M) management, the system first retrieves the basic industrial control equipment (ICS) O&M plan based on O&M commands. This plan includes pre-defined response measures for different types of equipment, fault modes, or security events, encompassing equipment repair, status monitoring, and troubleshooting. For example, for attacked equipment, network isolation and communication restrictions are used to isolate it from the system and prevent the attack from spreading. Based on the basic O&M plan, any ICS device is managed and operated. During this process, pre-defined monitoring windows are used to analyze various post-O&M indicators, including equipment recovery time, fault repair effectiveness, and system performance improvement. At the end of each monitoring window, O&M effectiveness indicators are collected to measure the success of the O&M, including recovery time and fault detection rate.
[0050] The feedback optimization process means adjusting and optimizing the existing basic operation and maintenance plan based on the collected operation and maintenance performance data. If some operation and maintenance measures are not effective, improvements are made based on the indicator results. For example, if the repair steps in the basic operation and maintenance plan are not efficient enough, or the equipment recovery time is too long, adjustments are made and new repair strategies or improvement measures are proposed. The target industrial control equipment operation and maintenance plan is a plan optimized through feedback, which aims to provide more efficient and accurate operation and maintenance. This plan is personalized according to the specific needs and operating environment of each device to ensure the best operation and maintenance effect.
[0051] Furthermore, the method includes: Within any preset monitoring window, based on the operation and maintenance effect evaluation function, the network situational awareness data after operation and maintenance is evaluated to generate operation and maintenance effect indicators for any industrial control equipment. The operation and maintenance effect evaluation includes network status stability evaluation, network anomaly detection rate evaluation, and operation response time evaluation. When the operation and maintenance effect indicators of any industrial control equipment do not meet the operation and maintenance effect constraints, the scheme feedback optimization for this cycle is performed. When the operation and maintenance effect indicators of any industrial control equipment meet the operation and maintenance effect constraints, the tracking evaluation for the next cycle begins.
[0052] Within any preset monitoring window, the network situational awareness data after maintenance is evaluated based on the maintenance effect evaluation function. This function measures the recovery status of equipment and the network security status after maintenance operations. Its purpose is to quantify the degree of improvement of equipment during maintenance and provide a basis for subsequent decision-making. Based on the evaluation results, industrial control equipment maintenance effect indicators are generated. These indicators measure whether the equipment has returned to normal working condition after maintenance or whether the maintenance measures are effective. Specifically, network stability evaluation measures the network stability of the equipment after maintenance, such as whether the equipment can return to normal working condition, whether the network can operate stably, and whether there are network fluctuations or disconnections. Network anomaly detection rate refers to whether all abnormal behaviors or potential threats in the network can be successfully detected after maintenance, such as whether network attacks, equipment failures, or other security events can be identified in a timely manner. Operation response time reflects the ability to quickly take countermeasures when problems or anomalies occur. This indicator assesses how quickly the system can initiate repair and protection measures after discovering a problem. A good maintenance solution should be able to respond quickly, reducing system downtime or losses.
[0053] If any industrial control equipment's operation and maintenance performance indicators do not meet the operation and maintenance performance constraints, that is, the equipment's performance does not meet the set standards, then the solution feedback optimization will be initiated. The operation and maintenance performance constraints may be: network stability must meet a certain standard, such as latency time being lower than a certain value, packet loss rate being lower than a certain percentage; anomaly detection rate should be higher than a certain threshold; response time must meet the requirements for rapid recovery, for example, response time less than 5 minutes.
[0054] During the feedback optimization process, the current operation and maintenance plan is optimized based on the evaluation results. The purpose of feedback optimization is to improve the repair and recovery capabilities of equipment by adjusting and improving the operation and maintenance strategy. For example, if the equipment recovery time is too long, the repair strategy is adjusted and more effective repair methods or tools are adopted; if the anomaly detection rate is insufficient, more monitoring indicators are added or existing monitoring rules are adjusted to improve detection accuracy; if the operation response time is too long, resource allocation is optimized, such as speeding up the response process and adding more resources.
[0055] If any industrial control equipment's operation and maintenance performance indicators meet the operation and maintenance performance constraints, it indicates that the operation and maintenance measures have successfully restored the equipment to normal operation and reached the predetermined standards. In this case, the next cycle of tracking and evaluation begins. This means that the equipment has been successfully restored and can enter the next stage of regular tracking to continue monitoring the equipment status and operation and maintenance performance. Periodic evaluation is the monitoring of the equipment's continuous health status to ensure that the equipment can operate stably at different times and to promptly identify potential problems.
[0056] Example 2 is based on the same inventive concept as the intelligent operation and maintenance method for industrial control equipment based on a large model in the previous examples, such as... Figure 2 As shown in the embodiment of this application, an intelligent operation and maintenance system for industrial control equipment based on a large model is provided. The system includes: The monitoring architecture establishment module 10 is used to deploy monitoring nodes according to the distribution information of industrial control equipment and obtain a distributed industrial control equipment monitoring architecture, which includes P network status monitoring nodes for P industrial control equipment, where P is a positive integer; the perception data acquisition module 20 is used to collect P network status perception time-series data according to a preset monitoring window based on the P network status monitoring nodes during the network communication process of the P industrial control equipment; the cross-channel collaborative analysis module 30 is used to input the P network status perception time-series data into the industrial control equipment network security event prediction model for cross-channel collaborative analysis and generate P industrial control equipment network security event prediction results, which include long and short time-series feature prediction channels and network security event identification channels, and each industrial control equipment network security event prediction result includes a network security event risk index; the operation and maintenance management module 40 is used to issue operation and maintenance instructions when the network security event risk index of any industrial control equipment reaches the network security event risk index threshold, and to execute the operation and maintenance management of the arbitrary industrial control equipment based on the operation and maintenance instructions.
[0057] Furthermore, the cross-channel collaborative analysis module 30 is used to perform the following operation steps: According to the preset monitoring window, network situational awareness data is backtracked for the P industrial control devices to obtain P historical network situational awareness time-series data sets; based on a predetermined long-short-term feature analysis scale, long-short-term features are captured from the P historical network situational awareness time-series data sets to obtain P long-short-term sample groups; supervised training is performed on the P long-short-term sample groups using a long short-term memory network, and when the training loss function converges to a predetermined threshold, the long-short-term feature prediction channel is generated; network security event sample groups and normal network status sample groups are extracted from the P long-short-term sample groups; supervised training is performed on the network security event sample groups and normal network status sample groups using a recurrent neural network, and when the training loss function converges to a predetermined threshold, the network security event identification channel is generated; multi-channel collaborative fusion is performed on the long-short-term feature prediction channel and the network security event identification channel to generate the industrial control device network security event prediction model.
[0058] Furthermore, the network security incident sample group includes a first set of long and short-term features of industrial control equipment when a network security incident occurs, the first set of long and short-term features having a network security incident label, and the network status normal sample group includes a second set of long and short-term features of industrial control equipment when the network status is normal, the second set of long and short-term features having a network status normal label.
[0059] Furthermore, the cross-channel collaborative analysis module 30 is used to perform the following operation steps: Extract the set of network security event types of industrial control equipment when a network security event occurs from the network security event sample group; perform incremental training on the network security event identification channel based on the set of network security event types and the first long and short time feature set.
[0060] Furthermore, the cross-channel collaborative analysis module 30 is used to perform the following operation steps: Based on the P sets of historical network situational awareness time-series data, network security events are extracted and aggregated into similar categories to obtain Q sets of similar network security event data, where Q is a positive integer; based on the Q sets of similar network security event data, the duration of similar network security events is extracted to obtain Q sets of duration of similar network security events; duration distribution analysis is performed on the Q sets of duration of similar network security events, and the predetermined long and short-term feature analysis scale is generated based on the analysis results.
[0061] Furthermore, the sensing data acquisition module 20 is used to perform the following operation steps: Anomaly detection data identification is performed on the P network situational awareness time series data, wherein anomaly detection includes data missing, data noise, and data duplication; anomaly pattern analysis is performed on the anomaly detection data, and when the anomaly pattern analysis result is a repairable anomaly, the anomaly detection data is corrected; when the anomaly pattern analysis result is an unrepairable anomaly, the anomaly detection data is removed, thus completing the data cleaning of the P network situational awareness time series data.
[0062] Furthermore, the operation and maintenance management module 40 is used to perform the following operation steps: When the network security incident risk index of multiple industrial control devices reaches the network security incident risk index threshold, multi-device joint analysis is performed to generate a linkage response strategy, and global operation and maintenance management is carried out based on the linkage response strategy.
[0063] Furthermore, the operation and maintenance management module 40 is used to perform the following operation steps: Based on the operation and maintenance instructions, the operation and maintenance plan of the basic industrial control equipment is retrieved, and the operation and maintenance management of the arbitrary industrial control equipment is carried out. The system performs periodic tracking and evaluation based on the preset monitoring window to obtain a set of operation and maintenance effect indicators for the arbitrary industrial control equipment. Based on the set of operation and maintenance effect indicators for the arbitrary industrial control equipment, the operation and maintenance plan of the basic industrial control equipment is optimized to generate a target industrial control equipment operation and maintenance plan.
[0064] Furthermore, the operation and maintenance management module 40 is used to perform the following operation steps: Within any preset monitoring window, based on the operation and maintenance effect evaluation function, the network situational awareness data after operation and maintenance is evaluated to generate operation and maintenance effect indicators for any industrial control equipment. The operation and maintenance effect evaluation includes network status stability evaluation, network anomaly detection rate evaluation, and operation response time evaluation. When the operation and maintenance effect indicators of any industrial control equipment do not meet the operation and maintenance effect constraints, the scheme feedback optimization for this cycle is performed. When the operation and maintenance effect indicators of any industrial control equipment meet the operation and maintenance effect constraints, the tracking evaluation for the next cycle begins.
[0065] Through the foregoing detailed description of the intelligent operation and maintenance method for industrial control equipment based on a large model, those skilled in the art can clearly understand the intelligent operation and maintenance system for industrial control equipment based on a large model in this embodiment. Since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and relevant parts can be referred to the method section.
[0066] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for intelligent operation and maintenance of industrial control equipment based on a large model, characterized in that, The method includes: Based on the distribution information of industrial control equipment, monitoring nodes are deployed to obtain a distributed industrial control equipment monitoring architecture, which includes P network status monitoring nodes for P industrial control equipment, where P is a positive integer. During the networking communication process of the P industrial control devices, P network situation awareness time-series data are collected and acquired according to a preset monitoring window based on the P network situation monitoring nodes. The P network situational awareness time-series data are input into the industrial control equipment network security event prediction model for cross-channel collaborative analysis to generate P industrial control equipment network security event prediction results. The industrial control equipment network security event prediction model includes a long and short time-series feature prediction channel and a network security event identification channel. Each industrial control equipment network security event prediction result includes a network security event risk index. When the network security event risk index of any industrial control device reaches the network security event risk index threshold, an operation and maintenance instruction is issued, and the operation and maintenance management of the arbitrary industrial control device is executed based on the operation and maintenance instruction.
2. The intelligent operation and maintenance method for industrial control equipment based on a large model as described in claim 1, characterized in that, The method for constructing the network security incident prediction model for the industrial control equipment includes: According to the preset monitoring window, network situational awareness data backtracking is performed on the P industrial control devices to obtain P sets of historical network situational awareness time series data. Based on a predetermined long and short time feature analysis scale, long and short time features are captured on the P sets of historical network situational awareness time series data to obtain P long and short time sample groups. The P long and short time sample groups are trained under supervision using a long short time memory network. When the training loss function converges to a predetermined threshold, the long and short time series feature prediction channel is generated. Based on the P long and short time sample groups, extract network security event sample groups and normal network status sample groups; The network security event sample group and the network normal status sample group are trained under supervision using a recurrent neural network. When the training loss function converges to a predetermined threshold, the network security event identification channel is generated. The long and short time-series feature prediction channels and the network security event identification channels are fused together to generate the network security event prediction model for the industrial control equipment.
3. The intelligent operation and maintenance method for industrial control equipment based on a large model as described in claim 2, characterized in that, The network security incident sample group includes a first set of long and short-term features of industrial control equipment when a network security incident occurs, and the first set of long and short-term features has a network security incident label. The network status normal sample group includes a second set of long and short-term features of industrial control equipment when the network status is normal, and the second set of long and short-term features has a network status normal label.
4. The intelligent operation and maintenance method for industrial control equipment based on a large model as described in claim 3, characterized in that, Generating the network security event identification channel also includes: Extract the set of network security event types for industrial control equipment in the network security event sample group when a network security event occurs; Incremental training is performed on the network security event identification channel based on the set of network security event types and the first set of long and short time features.
5. The intelligent operation and maintenance method for industrial control equipment based on a large model as described in claim 2, characterized in that, The method includes: Based on the P sets of historical network situational awareness time-series data, network security events are extracted and aggregated to obtain Q sets of network security event data of the same type, where Q is a positive integer. Based on the Q sets of similar network security event data, the duration of similar network security events is extracted to obtain a set of Q durations of similar network security events. A duration distribution analysis is performed on the set of Q similar network security events, and the predetermined short-term characteristic analysis scale is generated based on the analysis results.
6. The intelligent operation and maintenance method for industrial control equipment based on a large model as described in claim 1, characterized in that, The step of collecting P network situational awareness time-series data according to a preset monitoring window also includes: Anomaly detection data identification is performed on the P network situational awareness time-series data, wherein anomaly detection includes data missing, data noise, and data duplication; Anomaly pattern analysis is performed on the abnormal perception data. When the anomaly pattern analysis result is a repairable anomaly, the abnormal perception data is corrected. When the anomaly pattern analysis result is an unrepairable anomaly, the abnormal perception data is removed, thus completing the data cleaning of the P network situational awareness time series data.
7. The intelligent operation and maintenance method for industrial control equipment based on a large model as described in claim 1, characterized in that, The method further includes: When the network security incident risk index of multiple industrial control devices reaches the network security incident risk index threshold, multi-device joint analysis is performed to generate a linkage response strategy, and global operation and maintenance management is carried out based on the linkage response strategy.
8. The intelligent operation and maintenance method for industrial control equipment based on a large model as described in claim 1, characterized in that, The step of executing the operation and maintenance management of the arbitrary industrial control equipment based on the operation and maintenance instructions also includes: Based on the operation and maintenance instructions, the operation and maintenance plan of the basic industrial control equipment is invoked, and the operation and maintenance management of the arbitrary industrial control equipment is carried out. Based on the preset monitoring window, periodic tracking and evaluation are performed to obtain a set of operation and maintenance effect indicators for the arbitrary industrial control equipment. Based on the set of operation and maintenance performance indicators for any industrial control equipment, the operation and maintenance plan for the basic industrial control equipment is optimized through feedback, and the operation and maintenance plan for the target industrial control equipment is generated.
9. The intelligent operation and maintenance method for industrial control equipment based on a large model as described in claim 8, characterized in that, The method includes: Within any preset monitoring window, based on the operation and maintenance effect evaluation function, the operation and maintenance effect evaluation is performed on the network situation awareness data after operation and maintenance, and an operation and maintenance effect index for any industrial control equipment is generated. The operation and maintenance effect evaluation includes network status stability evaluation, network anomaly detection rate evaluation, and operation response time evaluation. When the operation and maintenance performance indicators of any industrial control equipment do not meet the operation and maintenance performance constraints, the solution feedback optimization for this cycle will be performed. When the operation and maintenance performance indicators of any industrial control equipment meet the operation and maintenance performance constraints, the next cycle of tracking and evaluation will begin.
10. An intelligent operation and maintenance system for industrial control equipment based on a large model, characterized in that: The system is used to implement the intelligent operation and maintenance method for industrial control equipment based on a large model as described in any one of claims 1-9, the system comprising: The monitoring architecture establishment module is used to deploy monitoring nodes according to the distribution information of industrial control equipment and obtain a distributed industrial control equipment monitoring architecture. The distributed industrial control equipment monitoring architecture includes P network status monitoring nodes for P industrial control equipment, where P is a positive integer. The sensing data acquisition module is used to acquire P network situation sensing time-series data based on the P network situation monitoring nodes according to a preset monitoring window during the network communication process of the P industrial control devices. The cross-channel collaborative analysis module is used to input the P network situational awareness time-series data into the industrial control equipment network security event prediction model for cross-channel collaborative analysis, and generate P industrial control equipment network security event prediction results. The industrial control equipment network security event prediction model includes a long and short time-series feature prediction channel and a network security event identification channel. Each industrial control equipment network security event prediction result includes a network security event risk index. The operation and maintenance management module is used to issue operation and maintenance instructions when the network security event risk index of any industrial control device reaches the network security event risk index threshold, and to perform operation and maintenance management of the arbitrary industrial control device based on the operation and maintenance instructions.
Citation Information
Patent Citations
Comprehensive device monitoring system architecture
CN106487585A
Vehicle-mounted network CAN bus node monitoring system and method
CN112083710A
Network security dynamic early warning method and system based on knowledge graph
CN119788344A
Power monitoring network security detection method and system
CN120110712A
Network security analysis method and system based on big data
CN120415850A