Security risk detection method, device and equipment of pipe network and storage medium
By obtaining the associated data of the pipeline security detection and using cluster analysis and Bayesian network model for risk prediction, the problem that traditional systems cannot cope with new attacks is solved, and the efficiency and accuracy of pipeline security detection is improved.
Patent Information
- Application Number
- CN202510707396.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-08-08
AI Technical Summary
Traditional network security monitoring systems cannot effectively deal with new and unknown attack methods, resulting in a reduced efficiency of pipeline security detection.
By obtaining the security detection correlation data of the pipeline network, using clustering analysis method to filter out abnormal data, and input it into the pre-trained Bayesian network model for risk prediction, and generating alarm prompt information.
The accuracy and efficiency of pipeline safety inspection have been improved, ensuring that managers understand potential risks in a timely manner, take measures to prevent and deal with them, and ensure the safe and stable operation of pipeline safety and stability.
Smart Images

Figure CN120455122A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data security technology, and in particular to a method, device, equipment and storage medium for detecting safety risks in a pipeline network. Background Art
[0002] With the rapid development of information technology, particularly the Internet of Things (IoT), industrial control systems, and the construction of large-scale network infrastructure, network security has become a crucial component of all systems. The secure operation of oil and gas pipeline networks increasingly relies on advanced technologies to monitor, analyze, and defend against network attacks and system failures. Against this backdrop, network security situational awareness technology has emerged. Its primary purpose is to promptly detect potential security threats, attacks, or system failures through real-time data collection and analysis, thereby ensuring the stability and security of network systems.
[0003] However, traditional network security monitoring systems are mostly based on fixed-point detection and rule-matching technologies. These systems generate alerts by detecting known attack patterns or abnormal behavior. However, as attacks become increasingly complex and diverse, traditional defenses are gradually exposing their limitations and are unable to effectively address new and unknown attack methods, resulting in reduced security monitoring efficiency for pipeline networks. Summary of the Invention
[0004] The present invention provides a method, device, equipment and storage medium for detecting safety risks of a pipeline network, so as to solve the problem of reduced safety detection efficiency of the pipeline network.
[0005] According to one aspect of the present invention, a method for detecting safety risks in a pipeline network is provided, the method comprising:
[0006] Obtaining safety detection related data of the pipe network, and determining whether abnormal data indicating behavioral attacks exists in the safety detection related data;
[0007] Performing clustering processing on the abnormal data based on a cluster analysis method to obtain at least one set of characterization data;
[0008] Inputting at least one set of the characterization data into a risk prediction model, and determining a predicted risk probability corresponding to the characterization data based on a model output, wherein the risk prediction model is a Bayesian network model pre-trained based on sample abnormal data;
[0009] When the predicted risk probability is not less than a preset risk probability threshold, an alarm prompt message is generated and an alarm prompt is issued.
[0010] According to another aspect of the present invention, a device for detecting safety risks in a pipeline network is provided, the device comprising:
[0011] An abnormal data determination module is used to obtain safety detection related data of the pipeline network and determine whether there is abnormal data of behavioral attack in the safety detection related data;
[0012] A cluster analysis module, configured to perform clustering processing on the abnormal data based on a cluster analysis method to obtain at least one set of characterization data;
[0013] a risk prediction module, configured to input at least one set of the characterization data into a risk prediction model, and determine a predicted risk probability corresponding to the characterization data based on the model output, wherein the risk prediction model is a Bayesian network model pre-trained based on sample abnormal data;
[0014] The alarm prompt module is used to generate alarm prompt information and issue an alarm prompt when the predicted risk probability is not less than a preset risk probability threshold.
[0015] According to another aspect of the present invention, an electronic device is provided, comprising:
[0016] at least one processor; and
[0017] a memory communicatively connected to the at least one processor; wherein,
[0018] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the safety risk detection method for the pipeline network described in any embodiment of the present invention.
[0019] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the safety risk detection method for a pipeline network described in any embodiment of the present invention when executed.
[0020] The technical solution of the embodiment of the present invention determines the presence of abnormal data indicating behavioral attacks in the safety detection associated data by acquiring safety detection associated data of the pipeline network; it is capable of quickly and accurately screening out data with potential threats from the detection data. Then, the abnormal data is clustered based on a cluster analysis method to obtain at least one set of characterization data; more representative feature patterns are extracted from the numerous abnormal data, which helps to more clearly understand the characteristics and patterns of different types of abnormal behaviors and provide more in-depth information support for subsequent risk assessment and response. Then, at least one set of the characterization data is input into a risk prediction model, and the predicted risk probability corresponding to the characterization data is determined based on the model output results, wherein the risk prediction model is a Bayesian network model pre-trained based on sample abnormal data; the Bayesian network model uses a large amount of sample abnormal data to learn the probabilistic relationship and dependency structure between risk factors, and can accurately predict the risk probability of new characterization data, thereby improving the accuracy of risk prediction. Finally, if the predicted risk probability is not less than a preset risk probability threshold, an alarm prompt information is generated and an alarm prompt is issued. This ensures that managers are immediately aware of potential risks in the pipeline network and can take appropriate preventive and treatment measures in a timely manner, effectively avoiding risk events or reducing losses caused by risks, and ensuring the safe and stable operation of the pipeline network. This solves the problem of reduced safety inspection efficiency in the pipeline network, achieving the beneficial effect of improving the accuracy of safety inspections and the overall level of pipeline network safety management.
[0021] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0023] Figure 1 This is a flow chart of a pipeline network safety risk detection method provided in accordance with the first embodiment of the present invention;
[0024] Figure 2 This is a flow chart of a pipeline network safety risk detection method provided in accordance with the second embodiment of the present invention;
[0025] Figure 3 This is a schematic structural diagram of a safety risk detection device for a pipeline network provided according to a third embodiment of the present invention;
[0026] Figure 4 It is a structural diagram of an electronic device for implementing the safety risk detection method for a pipe network according to an embodiment of the present invention. DETAILED DESCRIPTION
[0027] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0028] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0029] Example 1
[0030] Figure 1 A flow chart of a method for detecting safety risks in a pipe network is provided for the first embodiment of the present invention. This embodiment is applicable to the detection of safety risks in a pipe network. The method can be executed by a pipe network safety risk detection device. The pipe network safety risk detection device can be implemented in the form of hardware and / or software. The pipe network safety risk detection device can be configured in an electronic device. Figure 1 As shown, the method includes:
[0031] S110: Obtain security detection related data of the pipe network, and determine whether abnormal data indicating behavioral attacks exists in the security detection related data.
[0032] Among them, security detection-related data can be understood as a multidimensional data set related to pipeline network security detection. Security detection-related data can include network traffic data, control command logs, and sensor data. The existence of behavioral attacks can be understood as malicious operations or interference behaviors targeting pipeline network systems in the field of pipeline network security, intended to disrupt the normal operation of the pipeline network, steal sensitive information, or cause other security hazards. Abnormal data can be understood as data in security detection-related data that significantly deviates from normal data patterns. Abnormal data can reflect abnormal conditions in the pipeline network system, such as equipment failures and behavioral attacks.
[0033] Specifically, various safety monitoring data are collected from the various monitoring devices, sensors, and related management systems within the pipeline network. This data may be scattered across different systems and devices. It needs to be collected through data interfaces, network transmission, and other methods and integrated into a unified data storage platform for subsequent analysis and processing. The real-time safety monitoring data is then compared and analyzed against a normal data model. Any data that deviates from the range or characteristics set by the normal model is considered abnormal.
[0034] For example, network traffic data is collected by monitoring data packets flowing in and out of the network. Control command logs record the execution of network devices, systems, or control instructions, analyzing the legitimacy of operations and detecting possible malicious actions. Sensor data is real-time monitoring data from physical devices, such as temperature, pressure, and current, helping to detect hardware failures or abnormal operations. This data often comes from different hardware and software platforms and may be subject to format differences, noise interference, and missing values.
[0035] Optionally, after obtaining the safety inspection related data of the pipeline network, the following is also included:
[0036] Preprocessing is performed on the security detection associated data, and the security detection associated data is updated based on the preprocessed data, wherein the preprocessing includes denoising, interpolation, and normalization.
[0037] Specifically, the network traffic data is denoised based on Kalman filtering; the network traffic data includes packet size and transmission rate; linear interpolation is used to supplement missing values in the sensor data; the sensor data includes pressure, temperature and flow; robust normalization is used to normalize the network traffic data after noise removal, control command logs and sensor data after missing value supplementation to obtain preprocessed data.
[0038] Exemplarily, the formula for the robust normalization is:
[0039]
[0040] Wherein, x' represents the standardized data, x represents each data point in the safety detection associated data, median(x) represents the median in the safety detection associated data, and IQR(x) represents the range of the middle 50% of the data in the safety detection associated data.
[0041] Specifically, the Kalman filter is a recursive data processing algorithm used in signal processing and estimation problems. Its core is to exploit the relationship between dynamic models and measured data to derive optimal estimates. Using the Kalman filter, current data values can be inferred from historical data, thereby removing noise. Network traffic data (such as packet size and transmission rate) is typically time-series data. The Kalman filter estimates the true value of network traffic and removes noise caused by network jitter, external interference, or accidental anomalies. This makes network traffic data more stable and reliable, providing a clear data foundation for subsequent analysis. Sensor data (such as temperature, pressure, and flow) is an important data source reflecting system operating status. However, sensor data may be missing due to device failure, signal loss, or other reasons. In such cases, filling missing values is a crucial step in ensuring data integrity. For missing sensor data (such as pressure, temperature, and flow), linear interpolation can infer the missing data from known data at adjacent moments, thereby filling in the missing values and ensuring data continuity and integrity. Normalization converts data to a unified scale, giving different types of data equal weight, facilitating subsequent processing. Robust normalization is applied to noise-removed network traffic data, control command log data, and sensor data with missing values supplemented. This ensures that data from different sources are compared at the same scale, reduces the impact of extreme values on the model, and improves data quality.
[0042] In this embodiment of the present invention, Kalman filtering removes noise from network traffic data, ensuring its stability. Linear interpolation supplements missing values in sensor data, filling information gaps caused by device issues or data loss. Robust normalization eliminates the interference of outliers on the data, ensuring that multidimensional security detection-related data has a unified dimension and is comparable. This improves data validity and consistency.
[0043] Optionally, determining whether abnormal data indicating behavioral attacks exist in the security detection associated data includes: inputting the security detection associated data into the abnormal data identification model, and determining a model prediction value based on a decision function of the abnormal data identification model; and determining whether abnormal data indicating behavioral attacks exist in the security detection associated data based on the model prediction value and a preset abnormal threshold.
[0044] The predicted abnormality threshold may be preset based on experience, and this embodiment does not limit it.
[0045] Specifically, the security detection associated data is input into the abnormal data identification model, and a model prediction value is determined based on the abnormal data identification model and the model output result. If the model prediction value is less than the abnormal threshold, the security detection associated data is determined to be abnormal data indicating a behavioral attack.
[0046] Optionally, the model prediction value is determined based on the decision function of the abnormal data identification model using the following formula:
[0047]
[0048] Where f(x) represents the model prediction value, N represents the number of historical security event data, and α i is the Lagrange multiplier for each support vector, representing the weight of the support vector, K(x i , x) represents the kernel function, and calculates the support vector x i and the similarity between the test sample x, b represents the bias term, y i Represents the historical authentication status of the support vector.
[0049] Specifically, abnormal data is screened out according to the model prediction value of the abnormal data identification model.
[0050] Exemplarily, when f(x) is less than 0, the corresponding safety detection-related data is listed as abnormal data; when f(x) is greater than or equal to 0, the corresponding safety detection-related data is listed as non-abnormal data.
[0051] Optionally, before determining the risk number of behavioral attacks in the security detection associated data, it also includes: obtaining a preset number of historical security event data, and marking the historical authentication status corresponding to each of the historical security event data; wherein the historical authentication status includes the presence of a behavioral attack or the absence of a behavioral attack; constructing a first data set based on the marked historical security event data, and dividing the first data set into a first training set and a first test set; performing a preset number of iterative training on the initial support vector machine model based on the first training set, and testing the trained model based on the first test set, and when the model reaches a preset attack recall rate and the false alarm rate is less than a preset false alarm rate, using the model as an abnormal data recognition model.
[0052] Historical security event data can be understood as records of various security-related events that occurred in the past. The historical authentication status can be understood as a labeling result for each piece of historical security event data, indicating whether the event is related to a behavioral attack. It has two possible values: "Behavioral attack present" and "Behavioral attack not present." The first data set can be understood as a collection of historical security event data labeled with the historical authentication status. The preset attack recall rate and preset false alarm rate can be pre-set based on experience and are not limited in this embodiment.
[0053] Specifically, when establishing the SVM (Support Vector Machines) model, historical security event data is trained. Each historical data point consists of its features and labels, and the label indicates whether the data belongs to an attack behavior (for example: 1 means normal, -1 means attack). The SVM model is trained through the following steps: SVM constructs a decision boundary by calculating the similarity between data points. Commonly used kernel functions include linear kernel, radial basis kernel (RBF) and polynomial kernel. The RBF kernel function is selected in this embodiment. SVM solves an optimization problem to find an optimal hyperplane to separate data of different categories to the greatest extent. The optimal decision hyperplane is found by solving a convex optimization problem. The goal is to maximize the interval from the sample point to the hyperplane (ie, "maximum interval classification") while satisfying the correct classification constraints. By solving this optimization problem, the optimal Lagrange multiplier α can be obtained. i and the bias term b. After training, the SVM obtains the optimal decision function and can calculate its predicted value f(x) for a new data point x to be analyzed. If the predicted value is greater than or equal to 0, the sample is classified as normal data; if the predicted value is less than 0, the sample is considered abnormal data (i.e., potential attack behavior).
[0054] As you can see, the support vectors and their weights derived through training accurately delineate the boundaries between normal and abnormal behavior, demonstrating its superior ability to identify nonlinear security threats. Filtering data based on SVM predictions allows for simple and efficient labeling of abnormal data, reducing manual intervention and excessive rule-setting. This enables rapid response to potential security threats and timely identification of new attack behaviors.
[0055] S120: Perform clustering processing on the abnormal data based on a cluster analysis method to obtain at least one set of characterization data.
[0056] Among them, the characterization data can be understood as the representative feature data of each type of abnormal behavior.
[0057] Specifically, unsupervised learning methods, such as the K-means algorithm, are used to cluster the filtered abnormal data, grouping similar abnormal behaviors together. After cluster analysis is complete, representative data is extracted from each cluster, representing the representative features of each abnormal behavior category. These features help more accurately identify different types of attacks and reduce the complexity of subsequent analysis.
[0058] S130. Input at least one set of the characterization data into a risk prediction model, and determine the predicted risk probability corresponding to the characterization data based on the model output result, wherein the risk prediction model is a Bayesian network model pre-trained based on sample abnormal data.
[0059] Specifically, the Bayesian network is used to perform reasoning and analysis on the extracted abnormal data or characterization data to output the current security risk probability. The specific process is as follows: The Bayesian network models various variables (such as attacks, network traffic anomalies, etc.) by describing the causal relationship between events. Using Bayes' theorem, the probabilities of different security risk events (such as network attacks, equipment failures) are inferred based on the currently input abnormal data or characterization data. Through Bayesian network reasoning, the probability of the current security risk is calculated. For example, it is inferred that the probability of a certain type of attack behavior is 80%, thereby judging that the current security risk of the system is high. When the security risk probability exceeds a certain threshold, the early warning unit will trigger an alarm and send a warning to the security management personnel or the automatic response system, reminding them to take appropriate security protection measures.
[0060] Specifically, when it is determined that there is abnormal data of the same type, the characterization data is input into the Bayesian network model and the predicted risk probability is output; when the judgment unit determines that there is no abnormal data of the same type, the abnormal data is input into the Bayesian network model and the predicted risk probability is output.
[0061] For example, the risk probability is predicted based on the Bayesian network model by the following formula:
[0062]
[0063] Among them, P(Y|X1,X2,X3) represents the predicted risk probability of event Y occurring given the evidence X1,X2,X3, P(X1,X2,X3|Y) represents the predicted risk probability of observing Y occurring under the condition that events X1,X2,X3 occur, P(X1,X2,X3) represents the total probability of evidence X1,X2,X3 occurring, and P(Y) represents the predicted risk probability of event Y occurring in the absence of any evidence.
[0064] In this embodiment, the Bayesian network is used to infer the predicted risk probability based on the input data (i.e., abnormal data or characterization data). The Bayesian network is a graphical model used to represent the conditional dependency relationship between a set of variables, with an explicit probability distribution. Each node represents a random variable, and each directed edge represents the dependency relationship between variables. The Bayesian network can infer the probability of other unknown events through known evidence. When the judgment unit determines that there are abnormal data of the same type, the characterization data represents a specific type of attack behavior or abnormal behavior pattern, so these characterization data are input into the Bayesian network, and the Bayesian network derives the predicted risk probability through reasoning. When the judgment unit determines that there are no abnormal data of the same type, it means that the current abnormal data is composed of a group of isolated attack behaviors or system failures, and these abnormal data are input into the Bayesian network to output the corresponding predicted risk probability.
[0065] It is understandable that by introducing Bayesian networks as a reasoning tool, combined with conditional probability and evidential reasoning, the accuracy and flexibility of security risk situational awareness are improved. Bayesian networks can comprehensively consider the dependencies and uncertainties of multiple data sources, perform probabilistic reasoning based on multidimensional data, and assess current security risks. By selecting appropriate data (characteristic data or abnormal data) as input, risk assessments can be made for different attack patterns or abnormal behaviors. In addition, the dynamic update and reasoning capabilities of Bayesian networks ensure that the system can quickly respond to new attacks or emergencies, providing timely and effective decision-making support for network security management. This improves the warning capabilities and accuracy when facing complex and unknown threats.
[0066] Optionally, before inputting at least one set of the risk characterization data into the risk prediction model, it also includes: constructing a second data set, dividing the second data set into a second training set and a second test set; wherein, the second data set includes a preset number of sample abnormal data and the sample predicted risk probability corresponding to the sample abnormal data; calculating the posterior distribution of the model parameters through the second training set by the Bayesian estimation method, and updating the conditional probability table; testing the trained model on the second test set by log-likelihood and classification indicators to obtain at least one classification indicator, and determining the risk prediction model based on at least one of the model classifications.
[0067] The second dataset is the data set used to construct the risk prediction model. Sample anomaly data can be understood as data identified as having abnormal characteristics in the pipeline network or other systems. Abnormal characteristics may manifest as abnormal fluctuations in data values, inconsistencies with normal patterns, etc. Sample anomaly data is an important input for building risk prediction models. Sample predicted risk probability can be understood as each sample anomaly data having a corresponding sample predicted risk probability. This probability is calculated based on historical data, expert knowledge, or other relevant factors using certain methods (such as statistical analysis, machine learning, etc.). Bayesian estimation can be understood as a statistical method based on Bayes' theorem, used to estimate the posterior distribution of model parameters. In risk prediction models, conditional probability tables are used to describe the conditional probability relationships between different risk factors. Log-likelihood can be understood as an indicator used to measure how well the model fits the data. Classification metrics are a series of indicators used to evaluate the performance of classification models.
[0068] Specifically, a dataset specifically designed for risk prediction model construction, namely the second dataset, is collected and organized. This dataset contains a preset number of sample anomaly data, which reflect possible anomalies in the system. The constructed second dataset is divided into two parts: a second training set and a second test set. The second training set is used for model training, where the model adjusts its parameters by learning patterns and regularities from this data. The second test set, independent of the training set, is used for model evaluation, verifying the model's performance on unseen data. Using Bayesian estimation, the posterior distribution of the model parameters is calculated based on the data in the second training set. Bayesian estimation combines prior knowledge with observed data, applying Bayes' theorem to update the probability distribution of the model parameters. After obtaining the posterior distribution of the model parameters, these parameters are used to update the conditional probability table. The conditional probability table describes the conditional probability relationships between different risk factors. The updated conditional probability table more accurately reflects the actual relationships between risk factors, thereby improving the model's predictive accuracy. The trained model is tested using the second test set. During testing, the model's log-likelihood and classification metrics on the test set are calculated. The log-likelihood is used to measure the model's fit to the data, while classification metrics (such as precision and recall) are used to evaluate the model's classification performance. The final risk prediction model is determined based on the model's classification metric performance on the second test set.
[0069] S140: Generate alarm information and issue an alarm if the predicted risk probability is not less than a preset risk probability threshold.
[0070] The preset risk probability threshold may be pre-set based on experience, and this embodiment does not impose any specific restrictions thereon.
[0071] Specifically, the predicted risk probability is compared with the preset risk probability threshold. If the predicted risk probability is not less than the preset risk probability threshold, it means that the risk has reached a level that requires attention, and the system enters the alarm process. If the predicted risk probability is less than the preset risk probability threshold, it means that the risk is within an acceptable range, and the system continues to monitor without issuing an alarm. The system generates alarm prompt information containing key information based on preset templates and rules. The system sends alarm prompt information to relevant personnel through a variety of preset channels (such as emails, text messages, and other software that can accept alarm prompt information, etc., which are not specifically limited in this embodiment) to remind them to pay attention to risks and take appropriate measures.
[0072] For example, assume a pipeline safety monitoring system has a preset risk probability threshold of 60%. After analyzing the monitoring data for a particular section of the pipeline, the risk prediction model predicts a 65% probability of a leak in that section. The system then determines that 65% is not less than 60% and triggers an alarm process. The system generates an alarm message, perhaps stating, "Leakage risk exists in pipeline section XX, with a predicted risk probability of 65%." The system then sends an audible alarm and a text message to the monitoring personnel's mobile phone, enabling them to take timely measures for inspection and repair.
[0073] It's clear that the combination of multi-dimensional security detection correlation data fusion, SVM anomaly detection, cluster analysis, and Bayesian network reasoning improves the accuracy, robustness, and intelligence of security risk perception. Real-time collection and preprocessing of multi-dimensional data provides a more comprehensive reflection of the system's security status, avoiding false positives and false negatives associated with a single data source. The SVM model accurately identifies potential anomalous behavior, particularly novel attacks. Cluster analysis aggregates similar anomaly data, reducing data processing complexity and improving system efficiency. Leveraging the causal reasoning capabilities of Bayesian networks, the specific security risk probability is inferred from anomaly data and its characterization features, triggering precise early warnings. This improves the system's ability to identify and respond to complex security threats, ensuring the network security of critical infrastructure such as oil and gas pipelines.
[0074] The technical solution of the embodiment of the present invention determines the presence of abnormal data indicating behavioral attacks in the safety detection associated data by acquiring safety detection associated data of the pipeline network; it is capable of quickly and accurately screening out data with potential threats from the detection data. Then, the abnormal data is clustered based on a cluster analysis method to obtain at least one set of characterization data; more representative feature patterns are extracted from the numerous abnormal data, which helps to more clearly understand the characteristics and patterns of different types of abnormal behaviors and provide more in-depth information support for subsequent risk assessment and response. Then, at least one set of the characterization data is input into a risk prediction model, and the predicted risk probability corresponding to the characterization data is determined based on the model output results, wherein the risk prediction model is a Bayesian network model pre-trained based on sample abnormal data; the Bayesian network model uses a large amount of sample abnormal data to learn the probabilistic relationship and dependency structure between risk factors, and can accurately predict the risk probability of new characterization data, thereby improving the accuracy of risk prediction. Finally, if the predicted risk probability is not less than a preset risk probability threshold, an alarm prompt information is generated and an alarm prompt is issued. This ensures that managers are immediately aware of potential risks in the pipeline network and can take appropriate preventive and treatment measures in a timely manner, effectively avoiding risk events or reducing losses caused by risks, and ensuring the safe and stable operation of the pipeline network. This solves the problem of reduced safety inspection efficiency in the pipeline network, achieving the beneficial effect of improving the accuracy of safety inspections and the overall level of pipeline network safety management.
[0075] Example 2
[0076] Figure 2 A flowchart of a pipeline safety risk detection method provided in Example 2 of the present invention is provided. This embodiment further optimizes how to cluster the abnormal data based on the cluster analysis method in the above embodiment to obtain at least one set of risk characterization data. Optionally, the clustering of the abnormal data based on the cluster analysis method to obtain at least one set of risk characterization data includes: clustering the abnormal data based on the cluster analysis method to obtain at least two clusters; when there are at least two abnormal data in the same cluster, determining the existence of abnormal data of the same type; and extracting feature data of the abnormal data of the same type to obtain at least one set of risk characterization data.
[0077] like Figure 2 As shown, the method includes:
[0078] S210: Obtain security detection related data of the pipe network, and determine whether abnormal data indicating behavioral attacks exists in the security detection related data.
[0079] S220: Cluster the abnormal data based on a cluster analysis method to obtain at least two clusters.
[0080] Specifically, K centroids are initialized from all the abnormal data, the remaining abnormal data are assigned to the nearest centroids to form K clusters, and the centroid of each cluster is recalculated. The process of assigning the remaining abnormal data to the nearest centroid and calculating the centroid of each cluster is repeated until the centroid no longer changes.
[0081] For example, the centroid calculation formula is as follows:
[0082]
[0083] Where Mk represents the centroid of the kth cluster, |Nk| represents the number of abnormal data in the kth cluster, and Ti represents the feature set of the i-th abnormal data.
[0084] S230: When at least two abnormal data exist in the same cluster, determine that abnormal data of the same type exist.
[0085] Among them, abnormal data of the same type can be understood as abnormal data reflecting the same type of risk events or problems.
[0086] Specifically, when judging whether there is abnormal data of the same type in the abnormal data, it includes: when there are at least two abnormal data in the same cluster in the clustering result, it is determined that there are abnormal data of the same type; when any cluster in the clustering result contains only one abnormal data, it is determined that there is no abnormal data of the same type.
[0087] S240 , extracting characteristic data of the abnormal data of the same type respectively to obtain at least one set of characterization data.
[0088] Specifically, the distance between each abnormal data in each cluster and the center of the cluster is calculated respectively, and the abnormal data corresponding to the maximum distance is selected as the characterization data.
[0089] Specifically, cluster analysis is an unsupervised learning method whose goal is to divide the data points in the data set into several groups so that the similarity between data points in the same group is as high as possible, while the similarity between different groups is as low as possible. This embodiment adopts the K-means clustering algorithm, which divides the data into K clusters (each cluster has a centroid) in an iterative manner until the clustering result is stable. If a cluster contains at least two abnormal data, it means that this cluster represents a similar abnormal behavior pattern, and it is determined that there are abnormal data of the same type. If there is only one abnormal data in all clusters, it means that the data point is significantly different from the other data points and may be an isolated abnormal behavior, and it is determined that there are no abnormal data of the same type. In order to further reduce the difficulty of data processing in the subsequent Bayesian network causal inference process, the most representative abnormal data can be extracted from each cluster. For the abnormal data in each cluster, its Euclidean distance to the centroid of the cluster is calculated. For each cluster, the data point with the largest distance to the centroid is selected as the characterization data of the cluster. This data point represents the typical characteristics of this type of abnormal behavior.
[0090] It's easy to understand that cluster analysis can be used to classify and filter abnormal data, effectively identifying similar abnormal behaviors and simplifying subsequent causal reasoning by extracting the most representative abnormal data. Cluster analysis divides complex abnormal data sets into subsets, each representing a specific type of attack or abnormal behavior, enabling rapid identification of potential risk areas. Extracting representative data from each cluster reduces the interference of irrelevant or isolated data in the causal reasoning process, improving the accuracy and efficiency of the early warning system.
[0091] S250. Input at least one set of the characterization data into a risk prediction model, and determine the predicted risk probability corresponding to the characterization data based on the model output result, wherein the risk prediction model is a Bayesian network model pre-trained based on sample abnormal data.
[0092] S260: Generate alarm information and issue an alarm if the predicted risk probability is not less than a preset risk probability threshold.
[0093] The technical solution of the embodiment of the present invention clusters the abnormal data based on cluster analysis to obtain at least two clusters; when at least two abnormal data items exist in the same cluster, the presence of abnormal data of the same type is determined; and feature data of the abnormal data of the same type are extracted to obtain at least one set of risk characterization data. Cluster analysis can group risk data with similar characteristics into one category, thereby converting complex risk data into more representative risk characterization data. This risk characterization data can more clearly reflect the characteristics and patterns of different types of risks, facilitate a deeper understanding of the nature of risk, and provide more valuable information for subsequent risk prediction. This effectively refines risk characteristics.
[0094] Example 3
[0095] Figure 3 This is a schematic diagram of the structure of a pipe network safety risk detection device provided by the third embodiment of the present invention. Figure 3 As shown, the device includes: an abnormal data determination module 310, a cluster analysis module 320, a risk prediction module 330 and an alarm prompt module 340.
[0096] Among them, the abnormal data determination module 310 is used to obtain the safety detection related data of the pipeline network and determine whether there is abnormal data of behavioral attacks in the safety detection related data; the cluster analysis module 320 is used to cluster the abnormal data based on the cluster analysis method to obtain at least one set of characterization data; the risk prediction module 330 is used to input at least one set of the characterization data into the risk prediction model, and determine the predicted risk probability corresponding to the characterization data based on the model output result, wherein the risk prediction model is a Bayesian network model pre-trained based on sample abnormal data; the alarm prompt module 340 is used to generate alarm prompt information and issue an alarm prompt when the predicted risk probability is not less than the preset risk probability threshold.
[0097] The technical solution of the embodiment of the present invention determines the presence of abnormal data indicating behavioral attacks in the safety detection associated data by acquiring safety detection associated data of the pipeline network; it is capable of quickly and accurately screening out data with potential threats from the detection data. Then, the abnormal data is clustered based on a cluster analysis method to obtain at least one set of characterization data; more representative feature patterns are extracted from the numerous abnormal data, which helps to more clearly understand the characteristics and patterns of different types of abnormal behaviors and provide more in-depth information support for subsequent risk assessment and response. Then, at least one set of the characterization data is input into a risk prediction model, and the predicted risk probability corresponding to the characterization data is determined based on the model output results, wherein the risk prediction model is a Bayesian network model pre-trained based on sample abnormal data; the Bayesian network model uses a large amount of sample abnormal data to learn the probabilistic relationship and dependency structure between risk factors, and can accurately predict the risk probability of new characterization data, thereby improving the accuracy of risk prediction. Finally, if the predicted risk probability is not less than a preset risk probability threshold, an alarm prompt information is generated and an alarm prompt is issued. This ensures that managers are immediately aware of potential risks in the pipeline network and can take appropriate preventive and treatment measures in a timely manner, effectively avoiding risk events or reducing losses caused by risks, and ensuring the safe and stable operation of the pipeline network. This solves the problem of reduced safety inspection efficiency in the pipeline network, achieving the beneficial effect of improving the accuracy of safety inspections and the overall level of pipeline network safety management.
[0098] Optionally, the device further includes: a data annotation module, configured to obtain a preset number of historical security event data before determining the risk number of the presence of a behavioral attack in the security detection associated data, and annotate a historical authentication status corresponding to each historical security event data; wherein the historical authentication status includes whether a behavioral attack exists or does not exist;
[0099] A first data set construction module is used to construct a first data set based on the annotated historical security event data, and divide the first data set into a first training set and a first test set;
[0100] The first model training module is used to perform a preset number of iterative training on the initial support vector machine model based on the first training set, and test the trained model based on the first test set. When the model reaches a preset attack recall rate and the false alarm rate is less than a preset false alarm rate, the model is used as an abnormal data recognition model.
[0101] Optionally, the abnormal data determination module includes:
[0102] a model recognition unit, configured to input the security detection associated data into the abnormal data recognition model, and determine a model prediction value based on a decision function of the abnormal data recognition model;
[0103] The abnormal data determining unit is used to determine whether abnormal data indicating a behavioral attack exists in the security detection associated data based on the model prediction value and a preset abnormal threshold.
[0104] Optionally, the model recognition unit is specifically configured to:
[0105] The model prediction value is determined based on the decision function of the abnormal data identification model using the following formula:
[0106]
[0107] Where f(x) represents the model prediction value, N represents the number of historical security event data, and α i is the Lagrange multiplier for each support vector, representing the weight of the support vector, K(x i , x) represents the kernel function, and calculates the support vector x i and the similarity between the test sample x, b represents the bias term, y i Represents the historical authentication status of the support vector.
[0108] Optionally, the cluster analysis module includes:
[0109] a cluster analysis unit, configured to perform clustering processing on the abnormal data based on a cluster analysis method to obtain at least two clusters;
[0110] a same-type data determining unit, configured to determine the presence of abnormal data of the same type when at least two abnormal data are contained in the same cluster;
[0111] The characterization data extraction unit is used to respectively extract the characteristic data of the abnormal data of the same type to obtain at least one set of characterization data.
[0112] Optionally, the device further includes:
[0113] A second data set construction module is configured to construct a second data set before inputting at least one set of the characterization data into the risk prediction model, and divide the second data set into a second training set and a second test set; wherein the second data set includes a preset number of sample abnormal data and sample predicted risk probabilities corresponding to the sample abnormal data;
[0114] A second model training module is used to calculate the posterior distribution of model parameters using the second training set by using a Bayesian estimation method, and update the conditional probability table;
[0115] The risk prediction model determination module is used to test the trained model on the second test set through log-likelihood and classification indicators to obtain at least one classification indicator, and determine the risk prediction model based on at least one of the model classifications.
[0116] Optionally, the device further includes:
[0117] The preprocessing module is used to preprocess the safety detection associated data of the pipeline network after obtaining the safety detection associated data, and update the safety detection associated data based on the preprocessed data, wherein the preprocessing includes denoising, interpolation and normalization.
[0118] The pipeline network safety risk detection device provided in the embodiment of the present invention can execute the pipeline network safety risk detection method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0119] Example 4
[0120] Figure 4 A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0121] like Figure 4 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11. The memory stores a computer program that can be executed by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0122] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0123] The processor 11 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the method for detecting safety risks in a pipeline network.
[0124] In some embodiments, the security risk detection of the method pipeline network can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the security risk detection of the method pipeline network described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to perform security risk detection of the method pipeline network in any other appropriate manner (e.g., by means of firmware).
[0125] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0126] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0127] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0128] To provide interaction with a service acquirer, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the service acquirer; and a keyboard and pointing device (e.g., a mouse or trackball), through which the service acquirer can provide input to the electronic device. Other types of devices can also be used to provide interaction with the service acquirer; for example, the feedback provided to the service acquirer can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the service acquirer can be received in any form (including acoustic input, voice input, or tactile input).
[0129] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a service acquirer computer having a graphical service acquirer interface or a web browser through which a service acquirer can interact with embodiments of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0130] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.
[0131] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.
[0132] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.
Claims
1. A method for detecting safety risks of a pipe network, characterized in that: include: Obtaining safety detection related data of the pipe network, and determining whether abnormal data indicating behavioral attacks exists in the safety detection related data; Performing clustering processing on the abnormal data based on a cluster analysis method to obtain at least one set of characterization data; Inputting at least one set of the characterization data into a risk prediction model, and determining a predicted risk probability corresponding to the characterization data based on a model output, wherein the risk prediction model is a Bayesian network model pre-trained based on sample abnormal data; When the predicted risk probability is not less than a preset risk probability threshold, an alarm prompt message is generated and an alarm prompt is issued.
2. The method according to claim 1, characterized in that Before determining the risk number of behavioral attacks in the security detection associated data, the method further includes: Acquire a preset number of historical security event data and mark the historical authentication status corresponding to each of the historical security event data; wherein the historical authentication status includes whether a behavioral attack exists or does not exist; Constructing a first data set based on the annotated historical security event data, and dividing the first data set into a first training set and a first test set; The initial support vector machine model is iteratively trained a preset number of times based on the first training set, and the trained model is tested based on the first test set. When the model reaches a preset attack recall rate and the false alarm rate is less than a preset false alarm rate, the model is used as an abnormal data recognition model.
3. The method according to claim 2, characterized in that The determining whether abnormal data indicating a behavioral attack exists in the security detection associated data includes: Inputting the security detection associated data into the abnormal data recognition model, and determining a model prediction value based on a decision function of the abnormal data recognition model; Based on the model prediction value and a preset abnormality threshold, it is determined that abnormal data indicating behavioral attacks exists in the security detection associated data.
4. The method according to claim 3, characterized in that The model prediction value is determined based on the decision function of the abnormal data identification model using the following formula: Where f(x) represents the model prediction value, N represents the number of historical security event data, and α i is the Lagrange multiplier for each support vector, representing the weight of the support vector, K(x i , x) represents the kernel function, and calculates the support vector x i and the similarity between the test sample x, b represents the bias term, y i Represents the historical authentication status of the support vector.
5. The method according to claim 1, wherein The clustering process of the abnormal data based on the cluster analysis method to obtain at least one set of characterization data includes: Performing clustering processing on the abnormal data based on a cluster analysis method to obtain at least two clusters; In the case where there are at least two abnormal data in the same cluster, it is determined that there are abnormal data of the same type; Feature data of the abnormal data of the same type are extracted respectively to obtain at least one set of characterization data.
6. The method according to claim 1, characterized in that Before inputting at least one set of the characterization data into the risk prediction model, the method further comprises: Constructing a second data set, and dividing the second data set into a second training set and a second test set; wherein the second data set includes a preset number of sample abnormal data and the predicted risk probabilities of samples corresponding to the sample abnormal data; Calculating the posterior distribution of model parameters using the second training set by a Bayesian estimation method, and updating the conditional probability table; The trained model is tested on the second test set using log-likelihood and classification indicators to obtain at least one classification indicator, and the risk prediction model is determined based on at least one of the model classifications.
7. The method according to claim 1, characterized in that After obtaining the safety inspection related data of the pipeline network, it also includes: Preprocessing is performed on the security detection associated data, and the security detection associated data is updated based on the preprocessed data, wherein the preprocessing includes denoising, interpolation, and normalization.
8. A safety risk detection device for a pipe network, characterized in that: include: An abnormal data determination module is used to obtain safety detection related data of the pipeline network and determine whether there is abnormal data of behavioral attack in the safety detection related data; A cluster analysis module, configured to perform clustering processing on the abnormal data based on a cluster analysis method to obtain at least one set of characterization data; a risk prediction module, configured to input at least one set of the characterization data into a risk prediction model, and determine a predicted risk probability corresponding to the characterization data based on the model output, wherein the risk prediction model is a Bayesian network model pre-trained based on sample abnormal data; The alarm prompt module is used to generate alarm prompt information and issue an alarm prompt when the predicted risk probability is not less than a preset risk probability threshold.
9. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the safety risk detection method for the pipeline network according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the safety risk detection method for a pipeline network according to any one of claims 1 to 7 when executed.