Petrochemical engineering emergency early warning method and device based on large model data distillation

By using large-scale model data distillation technology, the problems of data redundancy and noise in petrochemical production have been solved, the speed and flexibility of emergency response have been improved, and efficient emergency early warning and dynamic data analysis have been achieved.

CN122087440APending Publication Date: 2026-05-26CHENGDU GREATECH ELECTRONIC TECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHENGDU GREATECH ELECTRONIC TECHNOLOGY CO LTD
Filing Date
2026-04-22
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently process massive amounts of sensor data in petrochemical production, leading to data redundancy and noise issues. Furthermore, they are slow to respond to emergencies, lack flexibility, and fail to meet real-time requirements.

Method used

A method based on large model data distillation is adopted, which extracts key features through data preprocessing, feature decomposition, feature extraction and deep learning models, and combines clustering algorithms to detect anomalies and trigger early warnings.

Benefits of technology

It improved data processing efficiency, optimized emergency response speed and flexibility, and enabled dynamic real-time data analysis and emergency dispatch, significantly enhancing the timeliness and accuracy of emergency response.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122087440A_ABST
    Figure CN122087440A_ABST
Patent Text Reader

Abstract

The invention relates to a petrochemical engineering emergency early warning method and device based on large model data distillation. The method comprises the steps of obtaining original data, preprocessing and standardizing the original data, and constructing a data set. The method comprises the following steps: mapping a data set from a high-dimensional space to a low-dimensional space through linear transformation, and performing eigenvalue decomposition on a covariance matrix of data in the data set to obtain eigenvalues and eigenvectors; and sorting the feature vectors according to the feature values from large to small, selecting the feature vectors of which the feature value ranking is not lower than a first threshold value to form a new feature space, and extracting important features from the feature space through LASSO regression. And calling a CNN deep learning model to extract key features from the important features, distilling representative features in the original data, and outputting a feature set after distillation. And performing anomaly detection on the petrochemical engineering system through a clustering algorithm according to the feature set after distillation, and triggering early warning when a detection result does not meet a preset condition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence and petrochemical automation technology, and in particular to a petrochemical emergency early warning method and device based on large model data distillation. Background Technology

[0002] As the petrochemical industry becomes increasingly complex and modernized, and the level of intelligence in production environments and equipment continues to rise, enterprises are facing ever-more urgent needs for safety management, fault prediction, and emergency response. Petrochemical production processes involve a vast amount of sensor data, equipment operating status, and environmental monitoring information; the rapid processing and efficient management of this data are crucial. However, this data is often highly complex and updated in real time. Extracting key features from massive amounts of data, making effective predictions, and enabling dynamic decision-making have become significant challenges for enterprises.

[0003] Currently, existing technologies rely on traditional data processing methods and human decision-making, which leads to data redundancy and noise problems, low processing efficiency, and difficulty in extracting valuable information in a timely manner. In addition, existing models are slow when processing large-scale complex data, making it difficult to meet the real-time requirements of emergency decision-making. Furthermore, emergency responses are based on static rules, lacking flexibility and making it difficult to respond to emergencies in a timely manner. Summary of the Invention

[0004] Therefore, it is necessary to provide a petrochemical emergency early warning method and device based on large model data distillation, which has high data processing efficiency and good timeliness and flexibility in emergency response, in order to address the above-mentioned technical problems.

[0005] This invention provides a petrochemical emergency early warning method based on large model data distillation, the method comprising:

[0006] Obtain raw data and preprocess and standardize the raw data to construct a dataset based on the preprocessed and standardized raw data;

[0007] The dataset is mapped from a high-dimensional space to a low-dimensional space by a linear transformation, so as to perform eigenvalue decomposition on the covariance matrix of the data in the dataset to obtain eigenvalues ​​and eigenvectors.

[0008] The feature vectors are sorted according to their feature values ​​from largest to smallest, and feature vectors with feature values ​​ranking no lower than a first threshold are selected to form a new feature space. Important features are then extracted from the feature space using LASSO regression.

[0009] The CNN deep learning model is invoked to extract key features from the important features, so as to distill the representative features in the original data and output the distilled feature set.

[0010] Anomalies in the petrochemical system are detected by clustering algorithm based on the feature set after distillation, and an early warning is triggered when the detection result does not meet the preset conditions.

[0011] The raw data includes real-time data on temperature, pressure, and flow rate collected from multiple sensors and monitoring devices. The conditions for triggering an early warning include an anomaly density value being lower than a second threshold, the number of anomalies exceeding a predetermined proportion, the frequency of anomaly patterns exceeding a third threshold, and the deviation between the detection results and historical normal data exceeding a fourth threshold.

[0012] This invention also provides a petrochemical emergency early warning device based on large model data distillation, the device comprising:

[0013] A data preprocessing module is used to acquire raw data and preprocess and standardize the raw data to construct a dataset based on the preprocessed and standardized raw data.

[0014] The feature decomposition module is used to map the dataset from a high-dimensional space to a low-dimensional space through linear transformation, so as to perform eigenvalue decomposition on the covariance matrix of the data in the dataset to obtain eigenvalues ​​and eigenvectors.

[0015] The feature extraction module is used to sort the feature vectors according to the feature values ​​from largest to smallest, select feature vectors whose feature values ​​are ranked not lower than a first threshold, form a new feature space, and extract important features from the feature space through LASSO regression.

[0016] The feature distillation module is used to call a CNN deep learning model to extract key features from the important features, so as to distill representative features from the original data and output the distilled feature set.

[0017] The early warning triggering module is used to perform anomaly detection on the petrochemical system based on the feature set after distillation using a clustering algorithm, and to trigger an early warning when the detection result does not meet the preset conditions.

[0018] The raw data includes real-time data on temperature, pressure, and flow rate collected from multiple sensors and monitoring devices. The conditions for triggering an early warning include an anomaly density value being lower than a second threshold, the number of anomalies exceeding a predetermined proportion, the frequency of anomaly patterns exceeding a third threshold, and the deviation between the detection results and historical normal data exceeding a fourth threshold.

[0019] The present invention also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the petrochemical emergency early warning method based on large model data distillation as described above.

[0020] The present invention also provides a computer storage medium storing a computer program, which, when executed by a processor, implements the petrochemical emergency early warning method based on large model data distillation as described above.

[0021] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the petrochemical emergency early warning method based on large model data distillation as described above.

[0022] The aforementioned petrochemical emergency early warning method and device based on large-scale model data distillation collects raw data on temperature, pressure, and flow rate from multiple sensors and monitoring devices. This raw data is preprocessed and standardized to construct a dataset. Then, a linear transformation maps the dataset from a high-dimensional space to a low-dimensional space, performing eigenvalue decomposition on the covariance matrix to obtain eigenvalues ​​and eigenvectors. The eigenvectors are then sorted according to their eigenvalues ​​from largest to smallest, selecting those with eigenvalues ​​ranking above a set threshold to form a new feature space. Important features are extracted from this feature space using LASSO regression. Subsequently, a CNN deep learning model is invoked to extract key features from the extracted important features, distilling representative features from the raw data and outputting a distilled feature set. Finally, a clustering algorithm is used to detect anomalies in the petrochemical system based on the distilled feature set, triggering an early warning when the detection results do not meet preset conditions. This method improves the response speed and accuracy of emergency decision-making by increasing data distillation efficiency to remove redundancy and noise, accurately extracting key features, and optimizing the data processing process. It also enables dynamic real-time data analysis and emergency dispatch, significantly enhancing the flexibility and timeliness of emergency response. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0024] Figure 1 A flowchart illustrating the petrochemical emergency early warning method based on large model data distillation provided by this invention;

[0025] Figure 2 This is a schematic diagram of the overall process of the petrochemical emergency early warning method based on large model data distillation in a specific embodiment of the present invention;

[0026] Figure 3A schematic diagram of the data acquisition and preliminary processing flow of the petrochemical emergency early warning method based on large model data distillation in a specific embodiment of the present invention;

[0027] Figure 4 A schematic diagram of the data dimensionality reduction and feature selection process for a petrochemical emergency early warning method based on large model data distillation, provided in a specific embodiment of the present invention;

[0028] Figure 5 A schematic diagram of the feature extraction and data distillation process of the deep learning model in the petrochemical emergency early warning method based on large model data distillation in a specific embodiment of the present invention;

[0029] Figure 6 A schematic diagram of the real-time monitoring and early warning process of the petrochemical emergency early warning method based on large model data distillation in a specific embodiment of the present invention;

[0030] Figure 7 A schematic diagram of the structure of the petrochemical emergency early warning device based on large model data distillation provided by the present invention;

[0031] Figure 8 This is an internal structural diagram of a computer device according to one embodiment. Detailed Implementation

[0032] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0033] The following is combined with Figures 1 to 8 This invention describes a petrochemical emergency early warning method and apparatus based on large model data distillation.

[0034] like Figure 1 As shown, in one embodiment, a petrochemical emergency early warning method based on large model data distillation includes the following steps:

[0035] Step S110: Obtain the raw data and preprocess and standardize the raw data to construct a dataset based on the preprocessed and standardized raw data.

[0036] In some embodiments, the petrochemical emergency early warning method based on large model data distillation provided by the present invention acquires raw data, preprocesses and standardizes the raw data, and constructs a dataset based on the preprocessed and standardized raw data, specifically including the following steps:

[0037] Step S111: Impute missing values ​​in the original data using linear interpolation. The interpolation formula is as follows:

[0038]

[0039] In the formula, For missing data at any given time point, and They are respectively with The next time point and The original data from the previous time point.

[0040] Step S112: Standardize the original data after missing value imputation to eliminate dimensional differences between different features. The standardization formula is:

[0041]

[0042] In the formula, and These are the mean and standard deviation of the i-th original data point, respectively. For the i-th original data, This is the standardized result of the i-th original data.

[0043] Step S120: The dataset is mapped from a high-dimensional space to a low-dimensional space through a linear transformation to perform eigenvalue decomposition on the covariance matrix of the data in the dataset, thereby obtaining eigenvalues ​​and eigenvectors.

[0044] In some embodiments, the petrochemical emergency early warning method based on large model data distillation provided by the present invention maps the dataset from a high-dimensional space to a low-dimensional space through linear transformation, and performs eigenvalue decomposition on the covariance matrix of the data in the dataset to obtain eigenvalues ​​and eigenvectors. Specifically, it includes the following steps:

[0045] Step S121: Calculate the covariance matrix of the standard data in the dataset, and call the PCA principal component analysis algorithm to perform eigenvalue decomposition on the covariance matrix to select principal components and obtain eigenvalues ​​and eigenvectors.

[0046] The formula for calculating the covariance matrix is ​​as follows:

[0047]

[0048]

[0049] In the formula, This is the mean vector of the standard data in the dataset. The number of standard data in the dataset. The standardized result of the i-th original data. Let be the covariance matrix.

[0050] Step S130: Sort the feature vectors according to the feature values ​​from largest to smallest, select the feature vectors whose feature values ​​are ranked not lower than the first threshold, form a new feature space, and extract important features from the feature space through LASSO regression.

[0051] In some embodiments, the petrochemical emergency early warning method based on large model data distillation provided by the present invention uses LASSO regression to minimize an objective function, the expression of which is:

[0052]

[0053] In the formula, To minimize the regression coefficient, For the target variable, For the standard data after dimensionality reduction, For regression coefficients, This is a regularization parameter used to control the intensity of feature selection.

[0054] Step S140: Call the CNN deep learning model to extract key features from important features, so as to distill the representative features in the original data and output the distilled feature set.

[0055] In some embodiments, the petrochemical emergency early warning method based on large model data distillation provided by the present invention uses a CNN deep learning model comprising convolutional layers, activation layers, pooling layers, and fully connected layers. The method involves calling the CNN deep learning model to extract key features from important features, distilling representative features from the original data, and outputting the distilled feature set. Specifically, this includes the following steps:

[0056] Step S141: The important features with non-zero regression coefficients are minimized as input to the convolutional layer. The important features are weighted and summed by sliding convolution kernels to extract local features and output a feature map composed of local features.

[0057] Step S142: Call the nonlinear activation function to map each feature value in the feature map to the non-negative domain, so as to reference the nonlinear feature and activate the feature map.

[0058] In some embodiments, the petrochemical emergency early warning method based on large model data distillation provided by the present invention uses a CNN deep learning model comprising convolutional layers, activation layers, pooling layers, and fully connected layers. The method involves calling the CNN deep learning model to extract key features from important features, distilling representative features from the original data, and outputting the distilled feature set. Specifically, the method further includes the following steps:

[0059] Step S143: Select the maximum value in the pooling window through max pooling, and downsample the feature map to reduce the spatial dimension of the feature map and retain the important features corresponding to the maximum value in each pooling window.

[0060] Step S144: Combine the features extracted by the convolutional layer and pooling layer through a fully connected layer to output the final prediction result, and construct the distilled feature set based on the final prediction result.

[0061] Step S150: Anomaly detection of the petrochemical system is performed using a clustering algorithm based on the feature set after distillation, and an early warning is triggered when the detection results do not meet preset conditions.

[0062] The conditions for triggering an alert include an anomaly density value being lower than the second threshold, the number of anomalies exceeding a predetermined proportion, the frequency of anomaly patterns exceeding the third threshold, and the deviation between the detection results and historical normal data exceeding the fourth threshold.

[0063] It should be noted that the first threshold, the second threshold, the third threshold, and the fourth threshold are all different preset thresholds that can be set as needed.

[0064] In some embodiments, the petrochemical emergency early warning method based on large model data distillation provided by the present invention uses a clustering algorithm to detect anomalies in the petrochemical system based on the feature set after distillation, and triggers an early warning when the detection result does not meet preset conditions. Specifically, it also includes the following steps:

[0065] Step S151: The similarity between data points in the feature set is calculated using a clustering algorithm, and potential risks or anomalies in the petrochemical system are identified based on the calculated similarity. When potential risks or anomalies exist in the petrochemical system, an alarm command is triggered and emergency response measures are updated through real-time data analysis.

[0066] Step S152: In response to the alarm command, the first data point is marked as an anomaly when the density value of the first data point is less than the second threshold, and an alarm is triggered when the number of detected anomalies exceeds the preset proportion in the dataset, and an alarm is triggered when the frequency of anomaly patterns exceeds the third threshold, and an alarm is triggered when the deviation between the current anomaly detection result and the historical normal data pattern exceeds the fourth threshold.

[0067] Among these, potential risks or abnormal situations include equipment malfunctions, production process abnormalities, and potential accident risks.

[0068] Combination Figures 2 to 6 As shown in the specific embodiment, the petrochemical emergency early warning method based on large model data distillation provided by the present invention includes steps 1 to 4:

[0069] Step 1: Data acquisition and preliminary preprocessing.

[0070] Raw data This includes real-time data such as temperature, pressure, and flow rate collected from various sensors and monitoring devices. Data Dimensions The number of features (i.e., the number of features) can be very high and contain noise and redundancy.

[0071] Preprocessing objectives: Remove redundancy, fill in missing data, and standardize data to prepare for subsequent feature extraction.

[0072] Missing value imputation: Missing values ​​are imputed using linear interpolation. Assume data from a specific point in time. Missing, interpolation calculation is as follows:

[0073]

[0074] in, and To and The next time point and The original data from the previous time point.

[0075] Standardization: To eliminate dimensional differences between different features, the data is standardized so that the mean of each feature is 0 and the variance is 1.

[0076]

[0077] in, and They are the first The mean and standard deviation of each feature.

[0078] Output: The dataset after standardization and interpolation. .

[0079] Step 2: Data dimensionality reduction and feature selection.

[0080] The dataset after standardization and interpolation To reduce the dimensionality of the data, the PCA (Principal Component Analysis) algorithm is primarily used for dimensionality reduction. Its purpose is to map data from a high-dimensional space to a low-dimensional space through linear transformation, while preserving as much variance as possible. The input data is... ,in It is the first One sample, with One characteristic.

[0081] Principal Component Analysis (PCA) is a type of principal component analysis algorithm that selects the most important principal components by performing eigenvalue decomposition on the covariance matrix of the data. First, the covariance matrix of the dataset is calculated. :

[0082]

[0083]

[0084] in, This is the mean vector of the data.

[0085] Then, eigenvalue decomposition is performed on the covariance matrix to obtain eigenvalues ​​and eigenvectors. The first eigenvalues ​​are then selected. The eigenvector with the largest eigenvalue forms a new feature space:

[0086]

[0087] in, For the front A matrix composed of eigenvectors It is the data after dimensionality reduction.

[0088] Feature selection: The goal of feature selection is to choose the features that have the greatest impact on the model based on the relevance and predictive power of the data. After dimensionality reduction, LASSO regression is used to further select the most important features. The optimization objective of LASSO regression is to minimize the following objective function:

[0089]

[0090] in, It is the target variable (e.g., the prediction result of petrochemical safety incidents). It is the data after dimensionality reduction. These are regression coefficients, representing the impact of each feature on the model. It is a regularization parameter that controls the intensity of feature selection.

[0091] In this embodiment, a set of sparse regression coefficients will be obtained through LASSO regression. Some coefficients are zero, indicating that the corresponding features were deleted. The final selected features correspond to those with non-zero coefficients. The final feature set From the data after dimensionality reduction The dataset consists of the most important features selected from the dataset.

[0092] In this embodiment, LASSO automatically selects the features that contribute most to the target variable by adjusting the regression coefficients. LASSO forces the regression coefficients of certain features to zero, removing these unimportant features with zero regression coefficients, ultimately obtaining a concise and effective feature set. It contains the key information needed for the prediction task.

[0093] Step 3: The deep learning model performs feature extraction and data distillation.

[0094] Specifically, the CNN deep learning model is used to further extract key features, distilling out the most representative information from the data to facilitate subsequent analysis and decision-making.

[0095] Convolutional Neural Network (CNN) Feature Extraction:

[0096] Convolutional Layer: The convolution operation is the most crucial operation in a CNN, primarily used to extract local features from the input data. The input data is a matrix. (For example, images, time-series data, etc.), the convolution kernel is .

[0097] Input data: ,in It is the sample size. This refers to the feature dimensions of each sample (e.g., the height, width, and number of channels of an image). The convolution operation is performed by sliding the convolution kernel. (its size is) ) Input data The above extracts local features. The formula for calculating the convolution operation is:

[0098]

[0099] in, It is the feature map (or activation map) after convolution. It is a convolution kernel. This is a bias term, typically added to each convolution result. The convolution kernel slides across the input data, performing weighted summations to obtain a new feature map. The feature map after the convolution operation. , which represents a combination of local features in the data. The input data matrix is ​​typically the data after feature selection, and its shape may be... ,in It is the number of samples. It is the characteristic number. The convolution kernel (filter) is typically a small matrix of size . It is used to extract local features from the input data, and the size and weight of each convolutional kernel are obtained through learning. The bias term is a constant value added to each convolution result during the convolution operation to help adjust the output of the convolution operation. This represents the position index in the convolution operation, referring to the position at the convolution output. Each position in it. Represents the convolution kernel The index is used to iterate through each element of the convolution kernel and perform calculations. The output feature map after convolution at position The value represents the convolution result at that position. Represents the input data matrix In the middle, convolution kernel The element applied to a local region of the input data matrix, at position . . convolution kernel The element in, at position .

[0100] Convolution process: Convolution kernel In the input data Slide the graph upwards, multiplying the data in the local region by the corresponding elements of the convolution kernel, and summing the results to obtain a single value. Then, add the bias term. This yields a feature value in the convolution output.

[0101] Output: The output feature map after the convolution operation. It is a new matrix that contains feature information of the input data in a local region.

[0102] Activation Layer: After the convolutional layer, a non-linear activation function, such as ReLU (Rectified Linear Unit), is usually added. The formula for the ReLU function is:

[0103]

[0104] It maps each feature value output by the convolutional layer to a non-negative domain, thereby introducing non-linear features that enable the network to learn and represent more complex patterns, resulting in activated feature maps. It is obtained by processing the convolution output using the ReLU function.

[0105] Pooling Layer: Pooling operations are used for downsampling, reducing the spatial dimensionality of the feature map while retaining the most important features. This patent uses max pooling. A 2x2 max pooling operation is used, and the pooling formula is:

[0106]

[0107] in, It is the feature map after convolution and activation. This is the feature map after pooling. Pooling layers reduce computation and prevent overfitting by decreasing the size of the feature map. (Feature map after pooling) It contains the most representative features. This represents the data input to the pooling layer, typically a feature map output from a convolutional layer, with a size of [size missing]. . This represents the output feature map after pooling. Pooling reduces the dimensionality of the data and extracts the most important features. Represents the output feature map in the pooling operation. The index position in the pooled feature map represents each position in the pooled feature map. Other related These indices indicate that the pooling operation occurred on the input feature map. The selected 2x2 region has a pooling window size of 2x2, so the maximum value is selected within each 2x2 region (for max pooling).

[0108] Pooling operation: The max pooling method is used to select the maximum value in the pooling window. The maximum value selection part in the formula is: Here, the pooling operation affects the input feature map. Perform downsampling, selecting the largest value within each 2x2 region as the pooling output.

[0109] Output: Feature map after pooling The feature map is downsampled in the spatial dimension, which preserves important feature information and reduces computation and storage space.

[0110] Convolution operation: Extract local features from the input data, perform weighted summation through convolution kernels, and obtain the convolution output.

[0111] Pooling operation: Reduces the spatial dimension of the feature map by downsampling, and selects the maximum value in each pooling window to retain the most important features.

[0112] Fully Connected Layer: After convolutional and pooling layers, one or more fully connected layers are typically added. Their role is to further combine the features extracted from the convolutional and pooling layers and output the final prediction result. The calculation formula for a fully connected layer is:

[0113]

[0114] in, This is the feature map after pooling. It is the weight matrix of the fully connected layer. It is a bias term.

[0115] Fully connected layers combine the features extracted from convolutional layers to generate output features.

[0116] Output: Features output by the fully connected layer This represents the final feature set after distillation through a deep learning model.

[0117] Step 4: Real-time monitoring and early warning.

[0118] Feature data after distillation It monitors potential risks and anomalies in the system in real time and issues early warnings.

[0119] Anomaly detection: Anomaly detection is performed using a clustering algorithm (DBSCAN), assuming the data points are... Anomaly detection is performed by calculating the density of each point:

[0120]

[0121] in, It is a point The neighborhood, It is a point in the neighborhood.

[0122] In this embodiment, after feature extraction and data distillation using a deep learning model, the obtained... This is processed, high-level feature data, typically containing key patterns in production or equipment status. This feature data serves as input to anomaly detection systems, which identify potential risks or anomalies by calculating the similarity between data points.

[0123] By monitoring and analyzing these characteristics in real time, the system can determine whether the following situations exist:

[0124] Equipment malfunctions: For example, abnormal fluctuations in equipment vibration, temperature, pressure, etc., may indicate equipment failure or potential safety hazards.

[0125] Abnormalities in the production process: If certain important variables in the production process (such as flow rate, pressure, temperature, etc.) deviate from the expected range, it may mean that the process is abnormal or the operation is improper.

[0126] Potential accident risks: If certain characteristics exhibit typical accident precursor patterns (such as equipment overload, high temperature, etc.), the system will be able to detect potential safety accidents in advance.

[0127] The system determines whether to trigger an alarm based on these characteristic data and dynamically updates emergency response measures through real-time data analysis.

[0128] formula:

[0129]

[0130] In the formula, It is the feature set obtained after distillation. each These are data samples processed by a deep learning model, representing the key features of production or equipment. It is a point The neighborhood of a specific point. In a data space, the set of all nearest neighbors can be defined by a radius. This determines that it includes all points within that radius. It is the first in the dataset There are 1 point, which belongs to point 1. neighborhood Points in the neighborhood are usually distances Points that are relatively close together are likely to be similar samples. It is a point and The distance metric between two samples is typically calculated using Euclidean distance, which represents the similarity between them. A smaller distance indicates that these points are closer in the feature space, meaning they may belong to the same category or pattern. It is a point neighborhood The number of dots in the middle represents How many neighboring points are there? This number can affect the density calculation. It is a calculation and all its neighboring points The sum of the distances between them. This sum can be used to assess... "Density" in the data space, i.e. The degree of concentration of surrounding points. This means that by normalizing the number of points in the neighborhood, the average density can be calculated, thus eliminating the influence of different neighborhood sizes on the density calculation.

[0131] Anomaly detection process: density calculation, by calculating the density of each point. Anomalies are identified by the density of a point (i.e., its similarity to other points in its neighborhood). Points with lower density are considered to be isolated from other points in their neighborhood and may be anomalies.

[0132] Conditions for triggering an alert:

[0133] The conditions for triggering an alert are usually based on the following aspects:

[0134] The density of outliers is lower than a set threshold: if a certain point The density value is less than the set threshold. If the condition is met, it indicates that the point is highly isolated in the feature space, potentially representing an anomaly. In this case, the system will trigger an alert and mark the point as an anomaly.

[0135]

[0136] in, It is a pre-set threshold, meaning that points with density values ​​below this threshold will be considered outliers.

[0137] Anomalies exceeding a predetermined percentage: When the detected anomalies exceed a certain percentage in the dataset (e.g., exceeding 5%), the system will automatically trigger an alert, indicating that a large-scale failure or anomaly may have occurred.

[0138] Frequency of abnormal patterns: The system can also monitor the frequency of abnormal patterns. When an abnormal pattern occurs frequently in a short period of time, it may mean that the system is experiencing a potential major failure, triggering an early warning.

[0139] Deviation from historical data: If the current anomaly detection results deviate significantly from historical normal data patterns, the system will trigger an alert based on these changes. For example, if current equipment operating data or production process characteristics deviate far from the mean or standard deviation of historical data, it may indicate that an anomaly is occurring.

[0140] The following describes the petrochemical emergency early warning device based on large model data distillation provided by the present invention. The petrochemical emergency early warning device based on large model data distillation described below can be referred to in correspondence with the petrochemical emergency early warning method based on large model data distillation described above.

[0141] like Figure 7As shown, in one embodiment, a petrochemical emergency early warning device based on large model data distillation includes a data preprocessing module 710, a feature decomposition module 720, a feature extraction module 730, a feature distillation module 740, and an early warning triggering module 750.

[0142] The data preprocessing module 710 is used to acquire raw data and preprocess and standardize the raw data to construct a dataset based on the preprocessed and standardized raw data.

[0143] The eigenvalue decomposition module 720 is used to map the dataset from a high-dimensional space to a low-dimensional space through linear transformation, so as to perform eigenvalue decomposition on the covariance matrix of the data in the dataset to obtain eigenvalues ​​and eigenvectors.

[0144] The feature extraction module 730 is used to sort the feature vectors according to the feature values ​​from largest to smallest, select feature vectors whose feature values ​​are ranked not lower than the first threshold, form a new feature space, and extract important features from the feature space through LASSO regression.

[0145] The feature distillation module 740 is used to call the CNN deep learning model to extract key features from important features, so as to distill representative features from the original data and output the distilled feature set.

[0146] The early warning triggering module 750 is used to perform anomaly detection on the petrochemical system based on the feature set after distillation using a clustering algorithm, and to trigger an early warning when the detection result does not meet the preset conditions.

[0147] The raw data includes real-time data on temperature, pressure, and flow rate collected from multiple sensors and monitoring devices. The conditions for triggering an alert include an anomaly density value being lower than a second threshold, the number of anomalies exceeding a predetermined proportion, the frequency of anomaly patterns exceeding a third threshold, and the deviation between the detection results and historical normal data exceeding a fourth threshold.

[0148] In this embodiment, the petrochemical emergency early warning device based on large model data distillation provided by the present invention, the data preprocessing module 710 is specifically used for:

[0149] Missing values ​​are filled in the original data using linear interpolation. The interpolation formula is as follows:

[0150]

[0151] In the formula, For missing data at any given time point, and They are respectively with The next time point and The original data from the previous time point.

[0152] The original data after missing value imputation is standardized to eliminate the dimensional differences between different features. The standardization formula is as follows:

[0153]

[0154] In the formula, and These are the mean and standard deviation of the i-th original data point, respectively. For the i-th original data, This is the standardized result of the i-th original data.

[0155] In this embodiment, the petrochemical emergency early warning device based on large model data distillation provided by the present invention, the feature decomposition module 720 is specifically used for:

[0156] Calculate the covariance matrix of the standard data in the dataset, and call the PCA principal component analysis algorithm to perform eigenvalue decomposition on the covariance matrix to select principal components and obtain eigenvalues ​​and eigenvectors.

[0157] The formula for calculating the covariance matrix is ​​as follows:

[0158]

[0159]

[0160] In the formula, This is the mean vector of the standard data in the dataset. The number of standard data in the dataset. The standardized result of the i-th original data. Let be the covariance matrix.

[0161] In this embodiment, the petrochemical emergency early warning device based on large model data distillation provided by the present invention uses LASSO regression to minimize the objective function, the expression of which is:

[0162]

[0163] In the formula, To minimize the regression coefficient, For the target variable, For the standard data after dimensionality reduction, For regression coefficients, This is a regularization parameter used to control the intensity of feature selection.

[0164] In this embodiment, the petrochemical emergency early warning device based on large model data distillation provided by the present invention uses a CNN deep learning model that includes convolutional layers, activation layers, pooling layers, and fully connected layers.

[0165] The characteristic distillation module 740 is specifically used for:

[0166] The important features with non-zero regression coefficients are used as input to the convolutional layer. The important features are weighted and summed by the sliding convolution kernel to extract local features and output a feature map composed of local features.

[0167] The nonlinear activation function is called to map each feature value in the feature map to the non-negative domain, thereby referencing the nonlinear features and activating the feature map.

[0168] In this embodiment, the petrochemical emergency early warning device based on large model data distillation provided by the present invention, specifically the characteristic distillation module 740 is further used for:

[0169] Max pooling is used to select the maximum value in the pooling window, and the feature map is downsampled to reduce the spatial dimension of the feature map while retaining the important features corresponding to the maximum value in each pooling window.

[0170] The features extracted by the convolutional and pooling layers are combined through a fully connected layer to output the final prediction result, and a distilled feature set is constructed based on the final prediction result.

[0171] In this embodiment, the petrochemical emergency early warning device based on large model data distillation provided by the present invention, the early warning triggering module 750 is specifically used for:

[0172] Clustering algorithms are used to calculate the similarity between data points in a feature set, and based on the calculated similarity, potential risks or anomalies in the petrochemical system are identified. When potential risks or anomalies exist in the petrochemical system, alarm commands are triggered and emergency response measures are updated through real-time data analysis.

[0173] In response to an alarm command, the first data point is marked as an anomaly when its density value is less than a second threshold, and an alarm is triggered when the number of detected anomalies exceeds a preset proportion in the dataset, when the frequency of anomaly patterns exceeds a third threshold, and when the deviation between the current anomaly detection result and the historical normal data pattern exceeds a fourth threshold.

[0174] Among these, potential risks or abnormal situations include equipment malfunctions, production process abnormalities, and potential accident risks.

[0175] Figure 8 This example illustrates a schematic diagram of the physical structure of an electronic device, which can be a smart terminal. Its internal structure diagram can be as follows: Figure 8As shown. The electronic device includes a processor, memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements a petrochemical emergency early warning method based on large model data distillation, which includes:

[0176] Obtain the raw data, and preprocess and standardize the raw data to construct a dataset based on the preprocessed and standardized raw data;

[0177] By using linear transformation to map the dataset from a high-dimensional space to a low-dimensional space, the covariance matrix of the data in the dataset is decomposed into eigenvalues ​​and eigenvectors.

[0178] The eigenvectors are sorted according to their eigenvalues ​​from largest to smallest. The eigenvectors with eigenvalues ​​ranked no lower than the first threshold are selected to form a new feature space. Important features are then extracted from the feature space using LASSO regression.

[0179] The CNN deep learning model is invoked to extract key features from important features, so as to distill representative features from the original data and output the distilled feature set.

[0180] Anomalies in the petrochemical system are detected by clustering algorithm based on the feature set after distillation, and an early warning is triggered when the detection results do not meet the preset conditions.

[0181] The raw data includes real-time data on temperature, pressure, and flow rate collected from multiple sensors and monitoring devices. The conditions for triggering an alert include an anomaly density value being lower than a second threshold, the number of anomalies exceeding a predetermined proportion, the frequency of anomaly patterns exceeding a third threshold, and the deviation between the detection results and historical normal data exceeding a fourth threshold.

[0182] Those skilled in the art will understand that Figure 8 The structure shown is merely a block diagram of a portion of the structure related to the present invention and does not constitute a limitation on the electronic device to which the present invention is applied. A specific electronic device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0183] On the other hand, the present invention also provides a computer storage medium storing a computer program, which, when executed by a processor, implements the above-mentioned petrochemical emergency early warning method based on large model data distillation.

[0184] On another front, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium, and when the processor executes the computer instructions, it implements the aforementioned petrochemical emergency early warning method based on large model data distillation.

[0185] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided by this invention can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory.

[0186] By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0187] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0188] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of this patent should be determined by the appended claims.

Claims

1. A petrochemical emergency early warning method based on large model data distillation, characterized in that, The method includes: Obtain raw data and preprocess and standardize the raw data to construct a dataset based on the preprocessed and standardized raw data; The dataset is mapped from a high-dimensional space to a low-dimensional space by a linear transformation, so as to perform eigenvalue decomposition on the covariance matrix of the data in the dataset to obtain eigenvalues ​​and eigenvectors. The feature vectors are sorted according to their feature values ​​from largest to smallest, and feature vectors with feature values ​​ranking no lower than a first threshold are selected to form a new feature space. Important features are then extracted from the feature space using LASSO regression. The CNN deep learning model is invoked to extract key features from the important features, so as to distill the representative features in the original data and output the distilled feature set. Anomalies in the petrochemical system are detected by clustering algorithm based on the feature set after distillation, and an early warning is triggered when the detection result does not meet the preset conditions. The raw data includes real-time data on temperature, pressure, and flow rate collected from multiple sensors and monitoring devices. The conditions for triggering an early warning include an anomaly density value being lower than a second threshold, the number of anomalies exceeding a predetermined proportion, the frequency of anomaly patterns exceeding a third threshold, and the deviation between the detection results and historical normal data exceeding a fourth threshold.

2. The petrochemical emergency early warning method based on large model data distillation according to claim 1, characterized in that, The process of acquiring raw data and preprocessing and standardizing the raw data to construct a dataset based on the preprocessed and standardized raw data includes: The original data is filled with missing values ​​using linear interpolation. The interpolation formula is as follows: , In the formula, For missing data at any given time point, and They are respectively with The next time point and The original data from the previous time point; The original data after missing value imputation is standardized to eliminate the dimensional differences between different features. The standardization formula is as follows: , In the formula, and These are the mean and standard deviation of the i-th original data point, respectively. For the i-th original data, This is the standardized result of the i-th original data.

3. The petrochemical emergency early warning method based on large model data distillation according to claim 1, characterized in that, The process of mapping the dataset from a high-dimensional space to a low-dimensional space through linear transformation, and performing eigenvalue decomposition on the covariance matrix of the data in the dataset to obtain eigenvalues ​​and eigenvectors, includes: Calculate the covariance matrix of the standard data in the dataset, and call the PCA principal component analysis algorithm to perform eigenvalue decomposition on the covariance matrix to select principal components and obtain the eigenvalues ​​and eigenvectors; The formula for calculating the covariance matrix is ​​as follows: , , In the formula, Let be the mean vector of the standard data in the dataset. The number of standard data in the dataset. The standardized result of the i-th original data. Let be the covariance matrix.

4. The petrochemical emergency early warning method based on large model data distillation according to claim 3, characterized in that, The optimization objective of the LASSO regression is to minimize the objective function, the expression of which is: , In the formula, To minimize the regression coefficient, For the target variable, For the standard data after dimensionality reduction, For regression coefficients, This is a regularization parameter used to control the intensity of feature selection.

5. The petrochemical emergency early warning method based on large model data distillation according to claim 4, characterized in that, The CNN deep learning model includes convolutional layers, activation layers, pooling layers, and fully connected layers; The step of calling a CNN deep learning model to extract key features from the important features, distilling out representative features from the original data, and outputting a distilled feature set includes: The important features with non-zero regression coefficients are used as input to the convolutional layer. The local features are extracted by weighted summation on the important features through the sliding convolution kernel, and a feature map composed of the local features is output. A nonlinear activation function is invoked to map each feature value in the feature map to a non-negative domain, thereby referencing the nonlinear feature and activating the feature map.

6. The petrochemical emergency early warning method based on large model data distillation according to claim 5, characterized in that, The step of calling a CNN deep learning model to extract key features from the important features, so as to distill representative features from the original data and output the distilled feature set, also includes: Max pooling is used to select the maximum value in the pooling window, and the feature map is downsampled to reduce the spatial dimension of the feature map while retaining the important features corresponding to the maximum value in each pooling window. The features extracted by the convolutional and pooling layers are combined through a fully connected layer to output the final prediction result, and the distilled feature set is constructed based on the final prediction result.

7. The petrochemical emergency early warning method based on large model data distillation according to any one of claims 1 to 6, characterized in that, The method of using a clustering algorithm to detect anomalies in the petrochemical system based on the feature set after distillation, and triggering an early warning when the detection results do not meet preset conditions, includes: The similarity between data points in the feature set is calculated by clustering algorithm, and potential risks or anomalies in the petrochemical system are identified based on the calculated similarity. When potential risks or anomalies exist in the petrochemical system, alarm commands are triggered and emergency response measures are updated through real-time data analysis. In response to the alarm command, the first data point is marked as an anomaly when the density value of the first data point is less than the second threshold, and an alarm is triggered when the number of detected anomalies exceeds a preset proportion in the dataset, and an alarm is triggered when the frequency of anomaly patterns exceeds a third threshold, and an alarm is triggered when the deviation between the current anomaly detection result and the historical normal data pattern exceeds a fourth threshold. The potential risks or abnormal situations include equipment malfunctions, production process abnormalities, and potential accident risks.

8. A petrochemical emergency early warning device based on large model data distillation, characterized in that, The device includes: A data preprocessing module is used to acquire raw data and preprocess and standardize the raw data to construct a dataset based on the preprocessed and standardized raw data. The feature decomposition module is used to map the dataset from a high-dimensional space to a low-dimensional space through linear transformation, so as to perform eigenvalue decomposition on the covariance matrix of the data in the dataset to obtain eigenvalues ​​and eigenvectors. The feature extraction module is used to sort the feature vectors according to the feature values ​​from largest to smallest, select feature vectors whose feature values ​​are ranked not lower than a first threshold, form a new feature space, and extract important features from the feature space through LASSO regression. The feature distillation module is used to call a CNN deep learning model to extract key features from the important features, so as to distill representative features from the original data and output the distilled feature set. The early warning triggering module is used to perform anomaly detection on the petrochemical system based on the feature set after distillation using a clustering algorithm, and to trigger an early warning when the detection result does not meet the preset conditions. The raw data includes real-time data on temperature, pressure, and flow rate collected from multiple sensors and monitoring devices. The conditions for triggering an early warning include an anomaly density value being lower than a second threshold, the number of anomalies exceeding a predetermined proportion, the frequency of anomaly patterns exceeding a third threshold, and the deviation between the detection results and historical normal data exceeding a fourth threshold.

9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the petrochemical emergency early warning method based on large model data distillation as described in any one of claims 1 to 7.

10. A computer storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the petrochemical emergency early warning method based on large model data distillation as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Abnormal data prediction method based on machine learning

    CN118277883A

  • Digital communication power supply health prediction method based on data collaboration

    CN118885890A

  • Regional interlocking cooperative early warning method, system and device, and storage medium

    CN120671023A