Industrial network data anomaly detection method and related equipment

By automatically generating a physical model and combining it with a temporal feature extraction network, a hybrid model is constructed for industrial network anomaly detection. This solves the problem of poor adaptability to complex systems in existing technologies and achieves high-precision and robust anomaly detection.

CN120856479AActive Publication Date: 2025-10-28PENG CHENG LAB

Patent Information

Application Number
CN202511359373.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-23
Publication Date
2025-10-28
Estimated Expiration
2045-09-23

AI Technical Summary

Technical Problem

Existing technologies for detecting anomalies in industrial networks based on physical models rely on specialized knowledge, making it difficult to accurately characterize the nonlinear dynamic characteristics of complex industrial systems. Furthermore, they are poorly adaptable to changes in system operating conditions and equipment aging, resulting in low detection accuracy.

Method used

By automatically generating physical models and combining them with a time-series feature extraction network, a hybrid physical model is constructed. The residual sequence is used to calculate cumulative statistics for anomaly detection, reducing reliance on domain expert knowledge and achieving automated discovery of physical laws and accurate fitting of nonlinear dynamic characteristics.

Benefits of technology

It significantly improves the accuracy and robustness of industrial network anomaly detection, enabling sensitive detection of low-amplitude covert attacks and gradual system failures, thereby enhancing overall prediction accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120856479A_ABST
    Figure CN120856479A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an industrial network data anomaly detection method and related equipment, and the method comprises the steps: firstly, obtaining industrial network data, and generating a physical model corresponding to the industrial network data; next, inputting the industrial network data into a time sequence feature extraction network for feature prediction to obtain a residual prediction value, and obtaining a mixed physical model based on the physical model and the residual prediction value; then, obtaining a mixed predicted value of the industrial network data based on the mixed physical model, and obtaining a residual sequence based on a difference value between the mixed predicted value and an actual observation value of the industrial network data; and finally, on the basis of the residual sequence, calculating to obtain a cumulative statistical value of the industrial network data, and on the basis of the cumulative statistical value, obtaining an anomaly detection result of the industrial network data, thereby remarkably improving the accuracy and robustness of the anomaly detection of the industrial network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of network data processing technology, and in particular to methods and equipment for detecting anomalies in industrial network data. Background Technology

[0002] For anomaly detection in industrial control networks (Industrial Internet), existing technologies are usually based on physical models or expert rules. In industrial control systems, physical processes inevitably follow engineering constraints and natural laws specific to the field. These constraints become important bases for anomaly detection. Therefore, anomaly detection methods based on physical model constraints are an effective protection path in the field of Industrial Internet security.

[0003] However, while this anomaly detection method based on physical model constraints has a certain degree of interpretability, the model construction process is highly dependent on professional knowledge, making it difficult to accurately characterize the nonlinear dynamic characteristics of complex industrial systems. Furthermore, it has poor adaptability to changes in system operating conditions and equipment aging, resulting in low accuracy of the model in detecting industrial network data. Summary of the Invention

[0004] This application provides an anomaly detection method and related equipment for industrial network data, which can improve the detection accuracy of anomalies in industrial network data.

[0005] To achieve the above objectives, a first aspect of this application proposes a method for anomaly detection in industrial network data, the method comprising: Acquire industrial network data and generate a physical model corresponding to the industrial network data; The industrial network data is input into a time-series feature extraction network for feature prediction to obtain residual prediction values. Based on the physical model and the residual prediction values, a hybrid physical model is obtained. Based on the hybrid physics model, a hybrid predicted value for the industrial network data is obtained, and based on the difference between the hybrid predicted value and the actual observed value of the industrial network data, a residual sequence is obtained. Based on the residual sequence, the cumulative statistical value of the industrial network data is calculated, and the anomaly detection result of the industrial network data is obtained based on the cumulative statistical value.

[0006] In some embodiments, generating a physical model corresponding to the industrial network data includes: Obtain multiple initial physical models of the industrial network data; The fitness is calculated based on the prediction accuracy and complexity of each initial physical model. Based on the fitness of each initial physical model, a target physical model is selected from multiple initial physical models, and cross-mutation is performed on multiple initial physical models to obtain an updated physical model. The updated physical model is then used as the new initial physical model, and multiple cross-mutation updates are performed. The target physical model is updated during the cross-mutation update process. The target physical model updated by the last crossover mutation is used as the physical model.

[0007] In some embodiments, the calculation of the corresponding fitness based on the prediction accuracy and complexity of each initial physical model includes: Substitute the industrial network data into each of the initial physical models to obtain initial physical prediction values; The prediction accuracy of each initial physical model is obtained based on the difference between the initial physical prediction value and the actual observation value of the industrial network data. Based on the expression composition parameters of each initial physical model, the corresponding complexity is determined, wherein the expression composition parameters include at least one of the following: number of symbol nodes, tree depth, and operator type. The fitness of each initial physical model is obtained based on a weighted sum of the prediction accuracy and the complexity.

[0008] In some embodiments, the training process of the temporal feature extraction network includes: Obtain the training physical model obtained from the training industrial network data, and determine the training prediction residual of the training physical model; The training industrial network data is input into the temporal feature extraction network for feature extraction to obtain multi-scale features, and the multi-scale features are integrated to obtain the training residual prediction value. The training prediction residual is used as the learning target for the training residual prediction value, and the network parameters of the temporal feature extraction network are adjusted accordingly.

[0009] In some embodiments, obtaining a hybrid physical model based on the physical model and the residual prediction values ​​includes: Obtain the first and second weights; The hybrid physical model is obtained by summing the product of the first weight and the physical model, the second weight and the residual prediction value.

[0010] In some embodiments, the residual sequence includes residual values ​​corresponding to multiple detection times, and the calculation of the cumulative statistical value of the industrial network data based on the residual sequence includes: Determine the mean and standard deviation of the residual sequence; The current residual ratio is obtained by dividing the difference between the residual value at the current detection time and the mean of the residual sequence by the standard deviation of the residual sequence. The cumulative statistical value of the industrial network data at the current detection time is obtained by subtracting the drift detection parameter from the sum of the cumulative statistical value of the previous detection time and the current residual ratio.

[0011] In some embodiments, obtaining the anomaly detection result of the industrial network data based on the cumulative statistical value includes: Obtain the anomaly detection range; When the cumulative statistical value is not within the anomaly detection range, an anomaly detection result is generated to characterize the industrial network data as abnormal; When the cumulative statistical value is within the anomaly detection range, an anomaly detection result is generated that indicates the industrial network data is normal.

[0012] To achieve the above objectives, a second aspect of this application provides an anomaly detection device for industrial network data, the device comprising: The physical model generation module is used to acquire industrial network data and generate a physical model corresponding to the industrial network data. The hybrid physical model generation module is used to input the industrial network data into a time-series feature extraction network for feature prediction, obtain residual prediction values, and obtain a hybrid physical model based on the physical model and the residual prediction values. The residual sequence generation module is used to obtain the mixed predicted value of the industrial network data based on the mixed physical model, and to obtain the residual sequence based on the difference between the mixed predicted value and the actual observed value of the industrial network data. An anomaly detection module is used to calculate the cumulative statistical value of the industrial network data based on the residual sequence, and to obtain the anomaly detection result of the industrial network data based on the cumulative statistical value.

[0013] To achieve the above objectives, a third aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the abnormal detection method for industrial network data as described in the first aspect.

[0014] To achieve the above objectives, a fourth aspect of the present application provides a storage medium, which is a computer-readable storage medium storing a computer program that, when executed by a processor, implements the industrial network data anomaly detection method described in the first aspect.

[0015] The present application proposes an anomaly detection method and related equipment for industrial network data. The method includes: first, acquiring industrial network data and generating a physical model corresponding to the industrial network data; next, inputting the industrial network data into a time-series feature extraction network for feature prediction to obtain residual prediction values, and obtaining a hybrid physical model based on the physical model and the residual prediction values; then, obtaining a hybrid prediction value of the industrial network data based on the hybrid physical model, and obtaining a residual sequence based on the difference between the hybrid prediction value and the actual observed value of the industrial network data; finally, calculating the cumulative statistical value of the industrial network data based on the residual sequence, and obtaining the anomaly detection result of the industrial network data based on the cumulative statistical value. This application's embodiments automatically generate physical models from industrial network data, reducing reliance on domain expert knowledge and achieving automated discovery of physical laws. This solves the problems of complexity and poor adaptability in traditional physical modeling processes. Furthermore, by constructing a hybrid physical model, a physically interpretable physical model is combined with a time-series feature extraction network capable of accurately fitting nonlinear dynamic characteristics. The time-series feature extraction network, which specifically learns and predicts the residuals of the physical model, achieves complementary advantages and greatly improves the overall prediction accuracy for complex industrial processes. Subsequently, by calculating the cumulative statistical value of the residual sequence of this high-precision hybrid model for anomaly detection, small but continuous anomalous signals can be effectively amplified. This enables sensitive detection of low-amplitude, covert attacks or gradual system failures that are difficult to detect using traditional methods, thereby significantly improving the accuracy and robustness of industrial network anomaly detection.

[0016] Other features and advantages of the present application will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present application. The purposes and other advantages of the present application can be achieved and obtained through the structures particularly pointed out in the description, claims and drawings. Attached Figure Description

[0017] Figure 1 This is a flowchart of an anomaly detection method for industrial network data provided in an embodiment of this application.

[0018] Figure 2 yes Figure 1 The flowchart for step 101.

[0019] Figure 3 yes Figure 2 The flowchart for step 202.

[0020] Figure 4 This is a flowchart of the training process of a temporal feature extraction network provided in an embodiment of this application.

[0021] Figure 5 yes Figure 1 The flowchart for step 102.

[0022] Figure 6 yes Figure 1 The flowchart for step 104.

[0023] Figure 7 yes Figure 1 Another flowchart for step 104.

[0024] Figure 8 This is a schematic diagram of the structure of an anomaly detection system that applies an anomaly detection method according to an embodiment of this application.

[0025] Figure 9 This is a schematic diagram of the structure of an industrial network data anomaly detection device provided in an embodiment of this application.

[0026] Figure 10 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0027] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0028] It should be noted that although functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart.

[0029] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0030] For anomaly detection in industrial control networks (Industrial Internet), existing technologies are usually based on physical models or expert rules. In industrial control systems, physical processes inevitably follow engineering constraints and natural laws specific to the field. These constraints become important bases for anomaly detection. Therefore, anomaly detection methods based on physical model constraints are an effective protection path in the field of Industrial Internet security.

[0031] Rule-based anomaly detection methods mainly include physical model-based methods and traditional methods based on expert experience rules. The core idea of ​​these methods is to utilize prior knowledge and theoretical foundations of the system to identify abnormal behavior. Physical model-based methods involve establishing mathematical description models based on the physical principles and engineering laws of industrial control systems, such as fluid dynamics models based on the laws of mass and energy conservation, or temperature distribution models based on heat transfer principles. These methods compare actual measured values ​​with the predicted values ​​of the physical model, and determine anomalies when the deviation exceeds a preset threshold. For example, in chemical process control, a physical model of the reactor can be established based on reaction kinetic equations, and anomalies can be detected by monitoring deviations between actual temperature, pressure, concentration, and other parameters and the model's predicted values. Expert experience rule-based methods are traditional methods that judge system anomalies through predefined expert rules and thresholds. They typically establish a series of if-then rules or boundary conditions based on the design specifications, operating manuals, and long-term operating experience of the industrial control system. While these methods have a certain degree of interpretability and theoretical basis, they have significant limitations: firstly, the model construction is highly complex and relies heavily on specialized knowledge. Accurate physical models require a deep understanding of the system's physical mechanisms, boundary conditions, and parameter characteristics. The model-building process necessitates extensive theoretical analysis and experimental verification, demanding a high level of expertise from modelers. For complex multiphysics coupled systems, accurate modeling often presents significant challenges. Secondly, models have limited adaptability and generalization capabilities. During long-term operation, industrial control systems are subject to changes in system characteristics due to factors such as equipment aging, process parameter drift, and environmental changes. Physical models based on fixed parameters struggle to adapt to these changes. When system configurations are updated or processes are adjusted, remodeling or significant parameter corrections are necessary. Thirdly, they struggle to handle high-dimensional complex systems. Modern industrial control systems often contain hundreds of monitoring points and control variables with complex nonlinear coupling relationships. Building an accurate model encompassing all variables based entirely on physical principles presents challenges in terms of computational complexity and practicality. Finally, they are sensitive to model errors and uncertainties. Parameter estimation errors, modeling assumption biases, and measurement noise in the physical model can all affect the accuracy of anomaly detection. Especially when facing sophisticated, covert attacks, the attack signal may be overwhelmed by model uncertainties, leading to poor detection results.

[0032] However, while this anomaly detection method based on physical model constraints has a certain degree of interpretability, the model construction process is highly dependent on professional knowledge, making it difficult to accurately characterize the nonlinear dynamic characteristics of complex industrial systems. Furthermore, it has poor adaptability to changes in system operating conditions and equipment aging, resulting in low accuracy of the model in detecting industrial network data.

[0033] To improve the accuracy of anomaly detection in industrial network data, this application automatically generates physical models from industrial network data, reducing reliance on domain expert knowledge and enabling automated discovery of physical laws. This solves the problems of complexity and poor adaptability in traditional physical modeling. Furthermore, by constructing a hybrid physical model, a physically interpretable physical model is combined with a time-series feature extraction network capable of accurately fitting nonlinear dynamic characteristics. This utilizes a time-series feature extraction network that specifically learns and predicts the residuals of the physical model, achieving complementary advantages and greatly improving the overall prediction accuracy for complex industrial processes. Subsequently, by calculating the cumulative statistical value of the residual sequence of this high-precision hybrid model for anomaly judgment, small but continuous anomalous signals can be effectively amplified. This enables sensitive detection of low-amplitude, covert attacks or gradual system failures that are difficult to detect using traditional methods, thereby significantly improving the accuracy and robustness of industrial network anomaly detection.

[0034] The following section will further describe the anomaly detection method and related equipment for industrial network data provided in this application.

[0035] The following section will first describe in detail the method for detecting anomalies in industrial network data in embodiments of this application. (Refer to...) Figure 1 This is an optional flowchart of the industrial network data anomaly detection method provided in the embodiments of this application. Figure 1 The method described may include, but is not limited to, steps 101 to 104. It is also understood that this embodiment... Figure 1 The order of steps 101 to 104 is not specifically limited; the order of steps can be adjusted or certain steps can be added or removed according to actual needs. The industrial network data anomaly detection method provided in this application can be applied to any server, processor, etc., connected to an industrial network.

[0036] Step 101: Acquire industrial network data and generate a physical model corresponding to the industrial network data.

[0037] Step 101 will be described in detail below.

[0038] In some embodiments, the data acquisition module first acquires multivariate time-series industrial network data from the industrial control network. This industrial network data includes system operation information such as various sensor data, controller status data, and network communication data. The data acquisition module includes a data interface unit, a data preprocessing unit, and a data storage unit. The data interface unit supports multiple industrial communication protocols (such as Modbus, OPC, Profinet, etc.); the data preprocessing unit cleans, denoises, and standardizes the raw data; and the data storage unit adopts a time-series database structure to ensure efficient storage and rapid retrieval of high-frequency data.

[0039] Subsequently, to establish a physically interpretable benchmark model, the obtained industrial network data was input into the symbolic regression physical modeling module for data processing. The symbolic regression physical modeling module is the core innovation of this application, comprising three sub-units: a genetic programming algorithm engine, a physical constraint library, and an expression optimizer. The genetic programming algorithm engine employs an improved genetic algorithm, searching for the optimal physical expression in a predefined function space through steps such as population initialization, fitness evaluation, selection operations, and crossover mutation. The physical constraint library pre-stores fundamental physical laws and constraints in the field of industrial control, guiding the search process and filtering candidate solutions that do not conform to physical laws. The expression optimizer employs a multi-objective optimization strategy, simultaneously considering model accuracy and complexity to ensure that the generated physical expression is both accurate and concise.

[0040] In the symbolic regression physical modeling module, the system employs automated modeling techniques such as symbolic regression to process the acquired industrial network data. Symbolic regression is a machine learning method that, without pre-setting a model structure, uses an evolutionary algorithm similar to genetic programming to automatically search for and discover explicit mathematical expressions describing the relationships between variables in the data. The resulting mathematical expression is the physical model corresponding to the industrial network data. This physical model reveals, in clear equation form, the inherent, physically-compliant relationships between system variables.

[0041] Reference Figure 2 Generate a physical model corresponding to the industrial network data, including the following steps 201 to 204.

[0042] Step 201: Obtain multiple initial physical models of industrial network data.

[0043] Step 202: Calculate the corresponding fitness based on the prediction accuracy and complexity of each initial physical model.

[0044] Steps 201 to 202 are described in detail below.

[0045] In some embodiments, after obtaining industrial network data, the symbolic regression physical modeling module utilizes a genetic programming algorithm engine to acquire multiple initial physical models of the industrial network data. This step is the initialization phase of the automated modeling process, and its core lies in constructing a diverse set of candidate solutions; in the context of genetic algorithms, this is called "population initialization." The system does not start from a single model hypothesis, but rather randomly generates a series of mathematical expressions with varying structures and parameters, each constituting an initial physical model. These models are represented in a tree structure, with leaf nodes representing variables or constants in the industrial network data and non-leaf nodes representing mathematical operators (such as addition, subtraction, multiplication, division, trigonometric functions, etc.). By generating multiple initial physical models with different structures, a rich genetic foundation is provided for subsequent evolution and optimization processes, preventing the algorithm from prematurely getting trapped in local optima.

[0046] Taking a typical industrial heating process as an example, suppose the goal is to automatically detect and describe the temperature of the storage tank. With heater power Material inflow rate and ambient temperature A physical model of the relationship between them. The gene programming algorithm engine first defines a set of "genes": basic variables (called "terminal nodes"), i.e. , , And some random constants; and a set of mathematical operations (called "function nodes"), such as addition, subtraction, multiplication, division, exponentiation, logarithms, etc. Then, the engine will randomly and recursively combine these terminal nodes and function nodes into multiple tree-structured mathematical expressions, just like randomly combining building blocks. These expressions are the initial physical model.

[0047] For example, during this initialization phase, the engine might randomly generate several initial physics models with vastly different structures as the first-generation "population": Model 1 is... Model 1 is a simple linear relationship; Model 2 is... A more complex nonlinear relationship; while Model 3 is This may be a physically implausible expression. The key to this process lies in "randomness" and "diversity." It does not presuppose any model form, but rather creates an initial solution space containing a large number of candidate models, ranging from simple to complex and from reasonable to unreasonable. This provides abundant raw materials for the subsequent selection and optimization of the best physical model through a process of "survival of the fittest."

[0048] Next, to select or generate the most suitable physical model from these initial physical models, the fitness of each initial physical model needs to be calculated based on its prediction accuracy and complexity. This step is a process of quantitatively evaluating each individual in the population (i.e., each initial physical model). Fitness is a key evaluation metric used to measure the quality of a model. It is not solely determined by prediction accuracy, which refers to the model's fit to industrial network data and is usually measured by metrics such as root mean square error. Fitness also considers complexity, which is a penalty for excessive model complexity, such as the number of nodes or tree depth, aiming to prevent the model from becoming overly bloated in order to fit the data, i.e., "overfitting." By combining prediction accuracy and complexity, a comprehensive fitness score is calculated, which guides the algorithm to evolve towards a model that is both accurate and concise. The following section further describes how to calculate the fitness of each initial physical model.

[0049] Reference Figure 3 The fitness is calculated based on the prediction accuracy and complexity of each initial physical model, including the following steps 301 to 304.

[0050] Step 301: Substitute the industrial network data into each initial physical model to obtain the initial physical prediction value.

[0051] Step 302: Based on the difference between the initial physical predictions and the actual observations of the industrial network data, obtain the prediction accuracy of each initial physical model.

[0052] Step 303: Based on the expression of each initial physical model, form parameters and determine the corresponding complexity.

[0053] Step 304: Based on the weighted sum of prediction accuracy and complexity, obtain the fitness of each initial physical model.

[0054] Steps 301 to 304 are described in detail below.

[0055] In some embodiments, to calculate the fitness of each initial physical model, industrial network data is first substituted into each initial physical model to obtain initial physical prediction values. This step is the execution phase of model evaluation, aimed at obtaining the specific predictive performance of each candidate model. For each initial physical model (i.e., each candidate mathematical expression) in the population, the system substitutes the industrial network data (e.g., historical or real-time sensor readings, control signals, etc.) as input into the expression for calculation. The output of this calculation process is the prediction output of that specific initial physical model under the given input data, i.e., the initial physical prediction value.

[0056] Then, based on the difference between the initial physical predictions and the actual observations from the industrial network data, the prediction accuracy of each initial physical model is obtained. The core of this step is quantifying the accuracy of each model's predictions. The system compares the initial physical predictions obtained in step 301 point-by-point with the corresponding, truly recorded actual observations from the industrial network data, calculating the prediction error index between them. Utilizing the inverse relationship between prediction accuracy and error, the prediction accuracy of each initial physical model is obtained. Generally, the smaller the error value, the higher the calculated prediction accuracy index. The prediction error index can employ one or more conventional error measurement methods, such as mean square error, root mean square error, and mean absolute error.

[0057] Then, the system aims to assess the complexity of the model structure itself to prevent overfitting due to excessive complexity. The system parses the mathematical expression structure of each initial physical model (i.e., analyzes the parameters that make up the expression) and calculates its complexity score based on one or more predefined parameters. The number of symbol nodes refers to the total number of variables, constants, and operators that make up the expression; a higher number indicates greater complexity. Tree depth refers to the maximum level of the expression in its tree-like representation; a greater depth indicates more nested computations and greater complexity. Operator type assigns different complexity weights to different types of mathematical operators. By quantifying these parameters, the system determines an objective complexity score for each model that represents its structural simplicity.

[0058] Next, the fitness of each initial physical model is obtained based on a weighted sum of prediction accuracy and complexity. This step is the core of the comprehensive evaluation, combining the model's performance and structure to form the final evaluation metric. The system calculates the final fitness of each initial physical model by performing a weighted sum of the obtained prediction accuracy and complexity. The weight coefficients can be preset according to specific application scenarios to adjust the relative importance of prediction accuracy and model simplicity in the final evaluation. For example, high prediction accuracy can be given a positive weight, while high complexity can be given a negative weight (i.e., as a penalty). The fitness obtained in this way is a comprehensive and balanced metric that can effectively evaluate whether a model has sufficient predictive power while also possessing good simplicity and generalization potential.

[0059] Through steps 301 to 304 above, a scientific and comprehensive model evaluation system is constructed. It not only evaluates the model's ability to fit existing data (i.e., prediction accuracy), but also innovatively introduces a quantitative consideration of the model's own structure (i.e., complexity). By calculating the fitness through a weighted sum of the two, the "overfitting" phenomenon can be effectively suppressed in the model selection process. This avoids the algorithm from choosing models that, although they can perfectly fit the training data, have an abnormally complex structure. As a result, the physical model selected in the end is not only accurate, but also more concise, easier to understand and interpret, thereby significantly improving the model's generalization ability and ensuring that it can maintain stable and reliable performance when facing new and unseen data.

[0060] Step 203: Select a target physical model from multiple initial physical models based on the fitness of each initial physical model, and perform cross-mutation on multiple initial physical models to obtain an updated physical model. Use the updated physical model as a new initial physical model and perform multiple cross-mutation updates. Update the target physical model during the cross-mutation update process.

[0061] Step 204: Use the target physical model updated by the last crossover mutation as the physical model.

[0062] Steps 201 to 204 are described in detail below.

[0063] After determining the fitness of each initial physical model, a target physical model is selected from multiple initial physical models based on their fitness. Crossover and mutation are then performed on these initial physical models to obtain updated physical models. These updated physical models are then used as the new initial physical models for multiple crossover and mutation updates, with the target physical model being updated during each crossover and mutation process. This is the core iterative step of the model evolution. First, based on the calculated fitness, the system selects models with higher fitness from the current population according to the principle of "survival of the fittest," with the model with the highest fitness being temporarily stored as the target physical model. Next, crossover and mutation are performed on the selected models: crossover involves exchanging the substructures (subtrees) of two models to generate new descendant models; mutation involves randomly changing a node (such as an operator or variable) in a single model. These operations generate entirely new updated physical models, which constitute a new generation of the population (i.e., new initial physical models). This process is repeated, i.e., multiple crossover and mutation updates are performed. After each generation update, the current optimal target physical model is re-evaluated and updated.

[0064] Finally, the target physical model updated after the last crossover mutation is used as the physical model. This step marks the termination of the entire evolutionary optimization process. After the set number of generations of evolution (i.e., multiple crossover mutation updates), the overall fitness of the model in the population will tend to stabilize, or the preset number of iterations will be reached. At this point, the algorithm stops evolving. The system uses the target physical model with the highest fitness recorded and saved throughout the entire iteration process as the final output. This final model is the optimal or suboptimal mathematical expression searched from a massive candidate solution space after balancing prediction accuracy and model simplicity. It is established as the physical model that can represent the inherent laws of the industrial network data. This is used in subsequent anomaly detection processes.

[0065] Through steps 201 to 204 above, an automated, data-driven physical model construction method is realized. By simulating the "survival of the fittest" and "genetic variation" mechanisms in biological evolution, it can automatically search for and discover potential, explicit mathematical relationships between variables in industrial network data, thereby generating a model with clear physical meaning. This solves the technical problems of traditional physical modeling methods, which heavily rely on domain expert knowledge, are time-consuming and labor-intensive, and are difficult to adapt to complex systems. At the same time, by introducing constraints on complexity into the fitness function, it ensures that the final generated physical model not only has high prediction accuracy but also has a simple structure and strong generalization ability, effectively avoiding overfitting problems. This provides a solid and interpretable foundation for building a highly robust anomaly detection system.

[0066] Step 102: Input the industrial network data into the time series feature extraction network for feature prediction to obtain the residual prediction value, and obtain the hybrid physical model based on the physical model and the residual prediction value.

[0067] Step 102 is described in detail below.

[0068] In some embodiments, after obtaining the physical model corresponding to the industrial network data After that, then... The hybrid physics model construction module is used for model fusion processing. This module is responsible for fusing the physical expressions (i.e., physical models) obtained from symbolic regression with deep learning models. This module includes a temporal feature extraction network unit, a model fusion unit, and a model training optimization unit. The temporal feature extraction network unit adopts an adaptive deep neural network architecture, which can automatically adjust the network structure according to the characteristics of specific industrial scenarios, effectively capturing the multi-timescale dynamic characteristics of industrial systems. The model fusion unit is responsible for organically combining the physical expressions from symbolic regression with the feature representations from deep learning, constructing a model like... A hybrid physics model, in which For industrial network data The corresponding physical model part, In time Industrial Network Data The corresponding residual prediction value part, This represents the random perturbation value of the model; the model training optimization unit employs adaptive learning rate adjustment and regularization techniques to prevent overfitting and improve generalization ability. In a preferred embodiment, this module can also integrate a physical constraint loss function, further enhancing the physical rationality of the model by adding a physical consistency constraint term during training.

[0069] Therefore, in order to construct a suitable hybrid physics model, it is first necessary to input industrial network data into a time-series feature extraction network for feature prediction to obtain residual prediction values. Then based on the physical model and residual prediction values This process yields a hybrid physical model. This step aims to compensate for complex nonlinear or dynamic characteristics that the generated physical model may not fully capture. The system feeds the same industrial network data into a temporal feature extraction network in parallel. This network typically employs deep learning architectures such as Long Short-Term Memory (LSTM) networks or Transformers, which excel at capturing complex dependencies in the data over time. The learning objective of this temporal feature extraction network is not the raw values ​​of the industrial network data, but rather the physical model. The prediction error is the difference between the physical model's predicted value and the actual observed value. Therefore, the feature prediction result output by this network is a prediction based on this error, i.e., the residual prediction value. Finally, by combining the deterministic output of the physical model with the residual predictions from the temporal feature extraction network, a hybrid physical model describing the system behavior is formed.

[0070] The following section first describes how to train the temporal feature extraction network to improve the output residual prediction value. It can effectively compensate for the physical model The prediction error.

[0071] Reference Figure 4 The training process of the temporal feature extraction network includes the following steps 401 to 403.

[0072] Step 401: Obtain the training physical model obtained from the training industrial network data, and determine the training prediction residuals of the training physical model.

[0073] Step 402: Input the training industrial network data into the time series feature extraction network for feature extraction to obtain multi-scale features, and integrate the multi-scale features to obtain the training residual prediction value.

[0074] Step 403: Use the training prediction residual as the learning target for the training residual prediction value, and adjust the network parameters of the temporal feature extraction network.

[0075] Steps 401 to 403 are described in detail below.

[0076] In some embodiments, the units of the temporal feature extraction network adopt an adaptive deep neural network architecture, including: an input layer that receives multi-dimensional temporal data and employs a multi-scale time window mechanism to automatically adjust the window size according to the time characteristics of different industrial scenarios; a feature extraction layer that adopts a network structure suitable for temporal data (such as a recurrent neural network, a long short-term memory network, or a Transformer architecture), adaptively selected according to data complexity and computational resources; and a fusion layer that integrates multi-scale temporal features and outputs residual prediction values. The adaptive adjustment mechanism dynamically adjusts the network depth, width, and connection method according to the characteristics of the specific industrial control system (such as the number of variables, time scale, degree of nonlinearity, etc.).

[0077] Based on this, in training the temporal feature extraction network, the first step is to obtain a training physical model from the training industrial network data and determine the training prediction residuals of the training physical model. This step aims to prepare "learning labels" and "ground values" for the training of the temporal feature extraction network. First, the system uses a portion of industrial network data specifically used for model training (i.e., training industrial network data) to construct a baseline physical model, which is referred to here as the training physical model. Subsequently, the training industrial network data is input into the pre-built training physical model for prediction. The model's prediction results are then compared with the actual observations in the training industrial network data, and the difference between the two is calculated. This sequence of differences is the training prediction residual of the training physical model, which precisely quantifies the shortcomings of the physical model in capturing the complex dynamics of the system. This training prediction residual serves as the learning objective of the deep learning model in the temporal feature extraction network.

[0078] Furthermore, to ensure the compatibility of the training physical model obtained from the symbolic regression output with the deep learning input of the temporal feature extraction network in terms of data format and dimensionality, feature alignment is required for the training prediction residuals corresponding to the training physical model. This includes: temporal dimension alignment, unifying the time sampling frequency and time window length of the symbolic regression physical expression and the deep learning model to ensure that both models process data from the same time period; variable dimension alignment, converting the physical expression output obtained from the symbolic regression into a data format compatible with the input layer of the deep learning network, including dimension matching and numerical range standardization; and data type unification, ensuring that the calculation results of the physical expression are consistent with the data type of the deep learning model to avoid data type conversion errors.

[0079] Next, the training industrial network data is input into a time-series feature extraction network for feature extraction, resulting in multi-scale features. These multi-scale features are then integrated to obtain the training residual prediction values. This step is the forward propagation process for the network's prediction. The system feeds the same training industrial network data into a time-series feature extraction network that is either not fully trained or is currently being trained. This network, through its internal recurrent layers (such as LSTM) or self-attention mechanisms (such as Transformer), captures dynamic patterns and dependencies at different time scales from the input time-series data; these extracted information constitute the multi-scale features. Subsequently, the network integrates and maps these multi-scale features through structures such as fully connected layers, ultimately outputting a sequence of prediction values. This sequence represents the network's estimate of the physical model residuals, i.e., the training residual prediction values.

[0080] Next, the training prediction residuals are used as the learning target for the training residual predictions, and the network parameters of the temporal feature extraction network are adjusted. This step is the core of network training, namely backpropagation and parameter update. The system uses the obtained "true" training prediction residuals as the learning target for the network's output training residual predictions. By calculating the loss function (e.g., mean squared error) between these two, the system can quantify the accuracy of the network's predictions. Then, using optimization algorithms such as gradient descent, all trainable network parameters (such as weights and biases) within the temporal feature extraction network are adjusted backward based on the loss value. This process is iterated repeatedly until the network's output (training residual predictions) can fit its learning target (training prediction residuals) very accurately, thus enabling the network to accurately predict the errors of the physical model.

[0081] During the training process, the model training optimization unit is responsible for the training process of the temporal feature extraction network, which includes: (1) Loss function design: the mean squared error is used as the basic loss function to measure the overall prediction accuracy of the hybrid model; (2) Training strategy formulation: a phased training strategy is adopted, first the physical expression is fixed to train the deep learning part, and then end-to-end joint optimization is performed (that is, the temporal feature extraction network is trained based on a fixed training physical model, and then the training physical model is updated before the temporal feature extraction network is trained); (3) Parameter update mechanism: an adaptive learning rate adjustment algorithm is used to dynamically adjust the learning rate according to the loss change during the training process; (4) Regularization processing: L1 / L2 regularization and Dropout technology are used to prevent overfitting and ensure the generalization ability of the model; (5) Early stopping strategy: monitor the performance of the validation set and stop training in time when the validation loss no longer decreases to avoid overfitting.

[0082] In one example, the training process of the temporal feature extraction network can also employ a physically constrained loss function. ,in, The standard mean square error loss, λ is an optional physical constraint term, and λ is the equilibrium parameter. Physical constraints can be various industrial physical constraints mentioned above, such as energy conservation, mass conservation, and temperature monotonicity, depending on the specific application scenario. This optional constraint mechanism further ensures the physical rationality of the model output; however, this constraint is not mandatory, and the system can still function effectively without using the physical constraint loss function.

[0083] Through steps 401 to 403, this invention designs an efficient and goal-oriented network training method. Instead of having the temporal feature extraction network learn the entire complex industrial process containing physical laws, it focuses the learning task on the training prediction residuals that the physical model cannot explain. This "differential learning" strategy greatly reduces the learning difficulty of the network, enabling it to focus on capturing the nonlinear and time-varying dynamic characteristics missed by the physical model. This not only makes the network training process easier to converge and more efficient, but also achieves a perfect complementarity between physical mechanisms and data-driven approaches in the final hybrid model. The physical model is responsible for explaining deterministic laws, while the network model accurately compensates for random and nonlinear errors, thereby achieving prediction accuracy and generalization ability far exceeding that of a single model.

[0084] The following section will further describe how to obtain a hybrid physical model based on the physical model and residual predictions.

[0085] Reference Figure 5 Based on the physical model and the residual prediction values, a hybrid physical model is obtained, including the following steps 501 to 502.

[0086] Step 501: Obtain the first weight and the second weight.

[0087] Step 502: Accumulate the product of the first weight and the physical model, and the second weight and the residual prediction value to obtain the hybrid physical model.

[0088] Steps 501 to 502 are described in detail below.

[0089] In some embodiments, the industrial network data to be detected is obtained. Corresponding physical model and residual prediction values Next, the first and second weights corresponding to these two factors are obtained. This step aims to determine the contribution of each component in the subsequent model fusion process. The first and second weights are pre-defined or learned numerical coefficients, which correspond to the physical model respectively. and residual prediction values The weights in the final hybrid model are determined by their proportions. These weights can be statically set based on an understanding of prior knowledge of the industrial system; for example, if the physical model is considered highly reliable under most operating conditions, a higher value can be assigned to the first weight. Alternatively, these weights can be used as hyperparameters, dynamically optimized during model training using an optimization algorithm to find the optimal combination that minimizes the final hybrid prediction error.

[0090] Then, the product of the first weight and the physical model, the second weight, and the residual prediction value are summed to obtain the hybrid physical model. This step involves the specific computational process for model fusion. The system multiplies the physical model's predicted output for industrial network data by a first weight, and simultaneously multiplies the residual predicted value from the time-series feature extraction network output by a second weight. Subsequently, these two products are summed point by point; the final output represents the comprehensive prediction result of the hybrid physical model. Through this weighted combination, the hybrid physical model essentially uses the physical model's prediction as a benchmark and refines it with weighted residual predicted values, thus forming a complete prediction model that has both a physical basis and data-driven corrections.

[0091] Through steps 501 to 502 above, a model fusion framework with a clear structure and significant effect is constructed. By introducing a first weight and a second weight, a flexible and adjustable mechanism is provided to balance the determinism of the physical mechanism and the fitting ability of the data-driven model, avoiding the instability that may be caused by the rigid combination of the two models. This weighted accumulation fusion method enables the physical model to provide a prediction baseline with strong interpretability, while the residual prediction model serves as its "correction term" to specifically compensate for the physical model's insufficient description of nonlinear and time-varying characteristics. The resulting hybrid physical model can therefore combine the advantages of both models, not only significantly improving the prediction accuracy of complex industrial processes, but also enhancing the model's robustness and adaptability to different working conditions.

[0092] Step 103: Obtain the mixed predicted values ​​of industrial network data based on the mixed physics model, and obtain the residual sequence based on the difference between the mixed predicted values ​​and the actual observed values ​​of industrial network data.

[0093] Step 103 will be described in detail below.

[0094] After obtaining the hybrid physics model Then, the current detection time will be... The corresponding industrial network data is input into the statistically enhanced anomaly detection module for anomaly detection processing to obtain anomaly detection results. The statistically enhanced anomaly detection module performs anomaly detection based on the output residual sequence of the hybrid physical model. This module includes a residual statistical analysis unit and a CUSUM control chart unit. The residual statistical analysis unit is responsible for establishing an accurate statistical model of the residual sequence, analyzing its distribution type, parameter characteristics, and time correlation; the CUSUM control chart unit achieves sensitive detection of small but persistent deviations through cumulative sum control charts. To amplify directional deviation signals. For the current detection time The cumulative statistics, For the previous detection time The cumulative statistics, This represents the residual value corresponding to the current detection time. It is the mean of the residual sequence (the residual sequence includes residual values ​​from multiple consecutive detection times). δ represents the standard deviation of the residual sequence, and δ is the drift detection parameter.

[0095] Based on this, a hybrid physics model was obtained. Next, the first thing to do is to set the current detection time. The corresponding industrial network data is input into this hybrid physics model to obtain the hybrid prediction value of the industrial network data. And based on mixed forecasts Actual observations of industrial network data The difference is used to obtain the residual value. And based on the residual values ​​of multiple consecutive detection times. The residual sequence is obtained. In this step, the system utilizes the hybrid physics model constructed in the previous step, which combines physical interpretability and high fitting capability, to make online or offline predictions on the input industrial network data, thereby generating a high-precision hybrid prediction value. This hybrid prediction value represents the model's best estimate of the ideal state of the system at each time step. Subsequently, this hybrid prediction value is compared point-by-point with the corresponding actual observations in the industrial network data, and the difference is calculated, thereby generating a time series that evolves over time, i.e., the residual sequence. Ideally, this residual sequence should appear as white noise with a mean of zero. It reflects the part of the system behavior that even a highly optimized hybrid physics model cannot explain, thus becoming an extremely sensitive indicator for detecting anomalies.

[0096] Step 104: Based on the residual sequence, calculate the cumulative statistical value of the industrial network data, and obtain the anomaly detection results of the industrial network data based on the cumulative statistical value.

[0097] Step 104 is described in detail below.

[0098] In some embodiments, when industrial network data is obtained at the current detection time Following the residual sequence, the system calculates the cumulative statistics of the industrial network data and uses these statistics to obtain anomaly detection results. To effectively detect subtle but persistent attacks or gradual faults with small amplitudes, this step avoids simple instantaneous threshold judgments on the residual sequence. Instead, it employs algorithms such as Cumulative Sum (CUSUM) from statistical process control to process the residual sequence. Specifically, the system continuously calculates the cumulative statistics of the residual sequence. This value amplifies small, unidirectional deviations over time, thus amplifying abnormal signals. The system pre-sets one or more control limits as thresholds. If the calculated cumulative statistics exceed these limits, the system determines that the industrial network data is abnormal and generates corresponding anomaly detection results, such as issuing an alarm signal. Conversely, if the cumulative statistics remain within the control limits, a detection result indicating normal system operation is generated.

[0099] The following section will further describe how to determine the cumulative statistics of industrial network data.

[0100] Reference Figure 6 Based on the residual sequence, the cumulative statistical value of the industrial network data is calculated, including the following steps 601 to 603.

[0101] Step 601: Determine the mean and standard deviation of the residual sequence.

[0102] Step 602: Based on the difference between the residual value at the current detection time and the mean of the residual sequence, divide it by the standard deviation of the residual sequence to obtain the current residual ratio.

[0103] Step 603: Based on the sum of the cumulative statistical value of the previous detection time and the current residual ratio, subtract the drift detection parameter to obtain the cumulative statistical value of the industrial network data at the current detection time.

[0104] Steps 601 to 603 are described in detail below.

[0105] In some embodiments, the residual mean and standard deviation are first determined. This step calibrates the baseline parameters for subsequent statistical analysis. The system first analyzes a residual sequence obtained after confirming that the system is in normal operating condition; this sequence is considered a reference for the "normal mode." Statistical calculations are then performed on this reference residual sequence to obtain its overall mean (i.e., the residual mean). Sum of standard deviations (i.e., the standard deviation of the residual series) These two statistics together characterize the central location and dispersion of residual fluctuations under normal operating conditions, providing a key benchmark for subsequent judgment on whether new residual values ​​deviate from the normal range.

[0106] Next, based on the residual value at the current detection time... With the mean of the residual sequence The difference is then divided by the standard deviation of the residual sequence. To obtain the current residual ratio This step aims to standardize each new residual value to eliminate the influence of dimensions and make them comparable. During online detection, after the system obtains the residual value at the current detection time, it first calculates the difference between it and the mean of the residual sequence. This difference reflects the degree to which the current residual deviates from the normal center. Subsequently, this difference is divided by the standard deviation of the residual sequence. This process is called "standardization" or "Z-score calculation," and the result is the current residual ratio. This ratio indicates how many "standard deviations" the current residual deviates from the normal center, and it is a dimensionless relative value that facilitates subsequent unified statistical accumulation.

[0107] In order to effectively detect small but continuous system changes, the cumulative statistics from the previous detection time are used. Ratio to current residual The accumulated value is then subtracted from a preset drift detection parameter. Finally, the cumulative statistical value of industrial network data at the current detection time is obtained. The cumulative statistic here is a core variable in sequential analysis techniques. It amplifies weak signals over time by continuously accumulating standardized residual information. The drift detection parameter is a key regulating factor that slightly "pulls back" the cumulative sum in each iteration to counteract the cumulative effect of purely random noise. This ensures that the cumulative statistic only increases significantly when the residuals are consistently and systematically biased in a certain direction. This iterative update mechanism allows the statistic to act like an "integrator," sensitively responding to long-term trend changes in the data (i.e., concept drift).

[0108] This invention constructs a dynamic cumulative monitoring mechanism for industrial network data through steps 601 to 603. By modeling the statistical characteristics of the residual sequence, standardizing the residuals at each time step to eliminate scale effects, and finally integrating historical and current information using a cumulative summation method with drift adjustment, the resulting cumulative statistics can not only effectively identify significant abrupt changes in the data, but also sensitively capture problems such as slow-changing system performance degradation or operating condition drift that are difficult to detect with traditional thresholding methods. This design makes the detection method more robust and sensitive in complex industrial network environments, providing strong technical support for ensuring the stable operation and predictive maintenance of industrial systems.

[0109] After obtaining the cumulative statistical values, they are input into the anomaly detection result processing module for anomaly matching and detection. The anomaly detection result processing module of this application is responsible for output processing and alarm management of detected anomalies. This module includes an anomaly scoring unit, a result output unit, and an alarm management unit. The anomaly scoring unit quantitatively evaluates the detected anomalies and generates anomaly confidence scores; the result output unit outputs the anomaly detection results in a standardized format for easy subsequent system integration and manual analysis; the alarm management unit generates alarm information and pushes it to relevant maintenance personnel through various means (email, SMS, system notifications, etc.). In a further embodiment, this module can also be expanded to include anomaly type classification and attack tracing analysis functions to provide more comprehensive anomaly diagnostic capabilities.

[0110] The following section will further describe how to use cumulative statistics to obtain anomaly detection results for industrial network data.

[0111] Reference Figure 7 The abnormal detection results of industrial network data are obtained based on the cumulative statistical values, including the following steps 701 to 703.

[0112] Step 701: Obtain the anomaly detection range.

[0113] Step 702: When the cumulative statistical value is not within the anomaly detection range, generate an anomaly detection result that characterizes the industrial network data as abnormal.

[0114] Step 703: When the cumulative statistical value is within the anomaly detection range, generate an anomaly detection result that characterizes the industrial network data as normal.

[0115] Steps 701 to 703 are described in detail below.

[0116] In some embodiments, in addition to obtaining the cumulative statistical value corresponding to the current detection time, it is also necessary to obtain the anomaly detection range. This step sets a clear and quantifiable standard for the final decision. The anomaly detection range typically consists of an upper control limit and a lower control limit, which together define a numerical interval. The setting of this interval is based on statistical principles and a deep understanding of the normal behavior of the system, aiming to define the reasonable limits of fluctuation that the cumulative statistical value should have under normal operating conditions. The process of obtaining this range may include statistical calculations based on historical normal data, or it may be set based on the requirements for system security (e.g., the expected false alarm rate and false negative rate), providing an objective and stable basis for subsequent judgments.

[0117] When the cumulative statistics of real-time detection If the data is outside the anomaly detection range, an anomaly detection result is generated to indicate that the industrial network data is abnormal. This step is the triggering and response mechanism for anomaly events. During continuous monitoring, the system continuously compares the cumulative statistical values ​​calculated in real time with the anomaly detection range obtained in the system. Once the cumulative statistical value exceeds the upper or lower limit of this range, it indicates that the cumulative deviation of the system residuals has exceeded the statistically acceptable normal fluctuations and reached a significantly abnormal level. At this time, the system immediately determines that the current state is abnormal and generates an anomaly detection result to indicate that the industrial network data is abnormal. This result can be manifested as issuing an alarm, logging, or triggering the corresponding safety response procedure.

[0118] When the cumulative statistics are within the anomaly detection range Within this process, anomaly detection results are generated to characterize the industrial network data as normal. This step describes the continuous verification process of the system under normal operating conditions. As long as the cumulative statistical values ​​calculated in real time fluctuate between the upper and lower limits defined by the anomaly detection range, the system considers the cumulative effect of the residuals not to pose a threat, and the system behavior conforms to its normal pattern. Therefore, the system will continuously generate anomaly detection results to characterize the industrial network data as normal. This step ensures that the detection system remains silent when no real anomalies occur, avoiding unnecessary interference, and is a routine state for maintaining normal system operation and monitoring.

[0119] This invention, through steps 701 to 703, establishes a clear, reliable, and automated decision-making process. It transforms a complex, continuous sequence of accumulated statistical values ​​into a simple, clear binary judgment result (normal or abnormal) by comparing it with a defined anomaly detection range. This decision-making mechanism based on statistical control limits has higher robustness and sensitivity compared to simple instantaneous threshold judgments because it operates on a signal amplified over time. This ensures that the final anomaly detection result is not only highly accurate but also effectively distinguishes between statistically random noise and persistent real system anomalies, thus providing timely, accurate, and actionable decision support for the security protection of industrial networks.

[0120] Reference Figure 8 This is a schematic diagram of the structure of an anomaly detection system applying an anomaly detection method, provided in an embodiment of this application. For example... Figure 8 As shown, the system first acquires real-time industrial network data from the industrial control network through the data acquisition module. This data is then fed into the symbolic regression physical modeling module and the hybrid physical model construction module. The core function of these modules is to construct a hybrid physical model capable of accurately predicting the normal behavior of the system. The difference between the predicted values ​​generated by this model and the actual acquired data constitutes the residual sequence described in the aforementioned technical solution. The statistically enhanced anomaly detection module in the figure is the core functional unit for calculating the aforementioned cumulative statistical value. It receives the hybrid physical model as a baseline and executes complete anomaly detection logic within this module: determining the statistical characteristics of the residual sequence, calculating the standardized current residual ratio, and finally iteratively updating the cumulative statistical value by combining the drift detection parameters, thereby generating the anomaly detection result. This result is ultimately sent to the anomaly detection result processing module for subsequent alarm or response handling, fully demonstrating a closed-loop processing process from establishing a dynamic baseline to implementing sensitive drift detection.

[0121] The following example uses an industrial scenario.

[0122] First, industrial network data needs to be acquired and a baseline behavior model established. In one scenario, an industrial control network connecting a programmable logic controller (PLC) and a remote terminal unit (RTU) is monitored. The request-response time intervals of the Modbus communication protocol between the two during a normal production cycle are collected as key industrial network data. The system uses this historical data on health status to train a predictive model (e.g., a time series prediction model), which can predict the normal time interval for the next detection moment based on the historical time interval sequence.

[0123] Afterward, the system enters the real-time monitoring phase and begins calculating the residual sequence. At each detection moment, the system records the actual request-response time interval (actual observation value), and simultaneously, the established prediction model outputs a predicted time interval (predicted value). The system calculates the difference between these two values ​​(actual value - predicted value) as a residual. Over time, these continuously calculated differences constitute the core analysis object of this scheme—the residual sequence. When the network is normal, this residual sequence should fluctuate randomly around the value of 0, like irregular white noise.

[0124] Then, the system performs the crucial cumulative analysis process. First, based on an initial, confirmed anomaly-free residual sequence, the system calculates its statistical characteristics: the residual mean (theoretically close to 0 milliseconds) and the residual standard deviation (e.g., 0.5 milliseconds). Next, assuming a slow, persistent performance degradation in the network, possibly due to equipment aging or minor network congestion, the response time is consistently and slightly longer than the model's predictions. For example, at a certain moment, with a residual value of +0.8 milliseconds, the system calculates the current residual ratio as (0.8 - 0) / 0.5 = 1.6. Immediately afterward, the system uses the cumulative statistic from the previous moment, adds this 1.6, and subtracts a small drift detection parameter (e.g., 0.1) to obtain the new cumulative statistic for the current moment. Because the performance problem is persistent, subsequent residual values ​​will remain positive, causing this cumulative statistic to continuously increase.

[0125] Finally, the system performs a final anomaly determination. An alarm threshold, such as 10.0, is preset within the system. At each detection point, the calculated cumulative statistics are compared to this threshold. In the early stages of performance degradation, although the residuals of a single instance are small, the cumulative statistics steadily increase over time, from 1.5 (0 + 1.6 - 0.1) to 2.9, then to 4.4… Finally, when this value exceeds 10.0, the system determines that a significant performance drift anomaly has occurred. At this point, the system triggers an alarm, notifying operations personnel of a potential, continuously deteriorating problem in the network. In this way, the solution successfully transforms a series of isolated, insignificant, small delays into a clear, quantifiable alarm signal, achieving effective detection of slow performance degradation or covert network attacks that are difficult to detect using traditional thresholding methods.

[0126] The present application proposes an anomaly detection method and related equipment for industrial network data. The method includes: first, acquiring industrial network data; acquiring multiple initial physical models of the industrial network data; substituting the industrial network data into each initial physical model to obtain initial physical prediction values; obtaining the prediction accuracy of each initial physical model based on the difference between the initial physical prediction values ​​and the actual observed values ​​of the industrial network data; determining the corresponding complexity based on the expression composition parameters of each initial physical model, where the expression composition parameters include at least one of the following: number of symbol nodes, tree depth, and operator type; obtaining the fitness of each initial physical model based on a weighted sum of prediction accuracy and complexity; selecting a target physical model from the multiple initial physical models based on the fitness of each initial physical model; performing cross-mutation on the multiple initial physical models to obtain an updated physical model; using the updated physical model as the new initial physical model; performing multiple cross-mutation updates; updating the target physical model during the cross-mutation update process; and finally, using the target physical model updated after the last cross-mutation update as... The physical model is used as follows: Next, industrial network data is input into a time-series feature extraction network for feature prediction, resulting in residual prediction values. First and second weights are then obtained, and the product of the first weight and the physical model, along with the second weight and the residual prediction values, is summed to obtain a hybrid physical model. Then, based on the hybrid physical model, a hybrid prediction value for the industrial network data is obtained, and based on the difference between the hybrid prediction value and the actual observed value of the industrial network data, a residual sequence is obtained. Finally, the residual sequence mean and residual sequence standard deviation are determined. Based on the difference between the residual value at the current detection time and the residual sequence mean, this is divided by the residual sequence standard deviation to obtain the current residual ratio. Based on the sum of the cumulative statistical value at the previous detection time and the current residual ratio, and subtracting the drift detection parameter, the cumulative statistical value of the industrial network data at the current detection time is obtained, and the anomaly detection range is acquired. When the cumulative statistical value is outside the anomaly detection range, an anomaly detection result indicating that the industrial network data is abnormal is generated; when the cumulative statistical value is within the anomaly detection range, an anomaly detection result indicating that the industrial network data is normal is generated.

[0127] This application's embodiments automatically generate physical models from industrial network data, reducing reliance on domain expert knowledge and achieving automated discovery of physical laws. This solves the problems of complexity and poor adaptability in traditional physical modeling processes. Furthermore, by constructing a hybrid physical model, it combines a physically interpretable physical model with a time-series feature extraction network capable of accurately fitting nonlinear dynamic characteristics. Utilizing a time-series feature extraction network that specifically learns and predicts the residuals of the physical model, it achieves complementary advantages, significantly improving the overall prediction accuracy for complex industrial processes. Next, by calculating cumulative statistics from the residual sequence of this high-precision hybrid model for anomaly detection, it effectively amplifies small but persistent anomalous signals, enabling sensitive detection of low-amplitude, covert attacks or gradual system failures that are difficult to detect using traditional methods. This significantly improves the accuracy and robustness of industrial network anomaly detection. In addition, a scientific and comprehensive model evaluation system is constructed, which not only assesses the model's fitting ability to existing data (i.e., prediction accuracy) but also innovatively introduces a quantitative consideration of the model's own structure (i.e., complexity). This is achieved by weighting the two factors together. Using fitness to calculate the model selection process effectively suppresses overfitting, preventing the algorithm from selecting models that perfectly fit the training data but have excessively complex structures. This results in a physical model that is not only accurate but also simpler, easier to understand and interpret, significantly improving its generalization ability and ensuring stable and reliable performance even with new and unseen data. Furthermore, an automated, data-driven physical model construction method is implemented. By simulating the "survival of the fittest" and "genetic variation" mechanisms in biological evolution, it can automatically search for and discover potential and explicit mathematical relationships between variables in industrial network data, generating models with clear physical meaning. This solves the technical problems of traditional physical modeling methods, which heavily rely on domain expert knowledge, are time-consuming and labor-intensive, and are difficult to adapt to complex systems. At the same time, by introducing constraints on complexity into the fitness function, it ensures that the final physical model not only has high prediction accuracy but also a simple structure and strong generalization ability, effectively avoiding overfitting and providing a solid and interpretable foundation for building a highly robust anomaly detection system.Furthermore, a well-structured and effective model fusion framework was constructed. By introducing first and second weights, a flexible and adjustable mechanism was provided to balance the determinism of the physical mechanism with the fitting ability of the data-driven model, avoiding the instability that might result from a rigid combination of the two models. This weighted accumulation fusion method provides a highly interpretable prediction baseline for the physical model, while the residual prediction model serves as its "correction term," specifically compensating for the physical model's shortcomings in describing nonlinear and time-varying characteristics. The resulting hybrid physical model thus combines the advantages of both models, significantly improving the prediction accuracy for complex industrial processes and enhancing the model's robustness and adaptability to different operating conditions. Additionally, a dynamic cumulative monitoring mechanism for industrial network data was constructed. This involves modeling the statistical characteristics of the residual sequence, then standardizing the residuals at each time step to eliminate scaling effects. By integrating historical and current information using a cumulative summation method with drift adjustment, the resulting cumulative statistics can not only effectively identify significant abrupt changes in the data, but also sensitively capture problems such as slow-changing system performance degradation or operational drift that are difficult to detect with traditional thresholding methods. This design makes the detection method more robust and sensitive in complex industrial network environments, providing strong technical support for ensuring the stable operation and predictive maintenance of industrial systems. Finally, a clear, reliable, and automated decision-making process is established, transforming complex, continuous cumulative statistical value sequences into simple and clear binary judgment results (normal or abnormal) by comparing them with a well-defined anomaly detection range. This decision-making mechanism based on statistical control limits has higher robustness and sensitivity than simple instantaneous threshold judgment because it operates on a signal amplified over time. This makes the final anomaly detection results not only highly accurate, but also able to effectively distinguish between statistically random noise and persistent real system anomalies, thus providing timely, accurate, and actionable decision support for the security protection of industrial networks.

[0128] This application also provides an anomaly detection device for industrial network data, which can implement the above-described anomaly detection method for industrial network data, as described above. Figure 9 The device 900 includes: The physical model generation module 910 is used to acquire industrial network data and generate a physical model corresponding to the industrial network data. The hybrid physical model generation module 920 is used to input industrial network data into a time series feature extraction network for feature prediction, obtain residual prediction values, and obtain a hybrid physical model based on the physical model and the residual prediction values. The residual sequence generation module 930 is used to obtain the mixed predicted values ​​of industrial network data based on the mixed physical model, and to obtain the residual sequence based on the difference between the mixed predicted values ​​and the actual observed values ​​of industrial network data. The anomaly detection module 940 is used to calculate the cumulative statistical value of industrial network data based on the residual sequence, and to obtain the anomaly detection result of industrial network data based on the cumulative statistical value.

[0129] In some embodiments, the physical model generation module 910 is further configured to: Multiple initial physical models for acquiring industrial network data; The fitness is calculated based on the prediction accuracy and complexity of each initial physical model. Based on the fitness of each initial physical model, a target physical model is selected from multiple initial physical models, and cross-mutation is performed on multiple initial physical models to obtain an updated physical model. The updated physical model is then used as a new initial physical model, and multiple cross-mutation updates are performed. The target physical model is updated during the cross-mutation update process. The target physical model updated by the last crossover mutation is used as the physical model.

[0130] In some embodiments, the physical model generation module 910 is further configured to: Substitute industrial network data into each initial physical model to obtain initial physical prediction values; The prediction accuracy of each initial physical model is obtained based on the difference between the initial physical predictions and the actual observations from the industrial network data. The complexity is determined based on the expression composition parameters of each initial physical model. The expression composition parameters include at least one of the following: number of symbol nodes, tree depth, and operator type. The fitness of each initial physical model is obtained by weighting the prediction accuracy and complexity.

[0131] In some embodiments, the hybrid physics model generation module 920 is further configured to: Obtain the training physical model obtained from the training industrial network data, and determine the training prediction residuals of the training physical model; The training industrial network data is input into the time series feature extraction network for feature extraction to obtain multi-scale features. The multi-scale features are then integrated to obtain the training residual prediction value. The training prediction residual is used as the learning objective for the training residual prediction value, and the network parameters of the temporal feature extraction network are adjusted accordingly.

[0132] In some embodiments, the hybrid physics model generation module 920 is further configured to: Obtain the first and second weights; The hybrid physical model is obtained by summing the product of the first weight and the physical model, the second weight, and the residual prediction value.

[0133] In some embodiments, the residual sequence generation module 930 is further configured to: Determine the mean and standard deviation of the residual sequence; The current residual ratio is obtained by dividing the difference between the residual value at the current detection time and the mean of the residual sequence by the standard deviation of the residual sequence. The cumulative statistical value of the industrial network data at the current detection time is obtained by subtracting the drift detection parameter from the sum of the cumulative statistical value of the previous detection time and the current residual ratio.

[0134] In some embodiments, the anomaly detection module 940 is further configured to: Obtain the anomaly detection range; When the cumulative statistical value is outside the anomaly detection range, an anomaly detection result is generated that characterizes the industrial network data as abnormal; When the cumulative statistical value is within the anomaly detection range, an anomaly detection result is generated to characterize the industrial network data as normal.

[0135] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, the specific implementation of the industrial network data anomaly detection device is basically the same as the specific implementation of the above industrial network data anomaly detection method, and will not be repeated here.

[0136] In this embodiment, the anomaly detection device for industrial network data automatically generates a physical model from industrial network data, reducing reliance on domain expert knowledge and achieving automated discovery of physical laws. This solves the problems of complexity and poor adaptability in traditional physical modeling processes. Furthermore, by constructing a hybrid physical model, a physically interpretable physical model is combined with a time-series feature extraction network capable of accurately fitting nonlinear dynamic characteristics. This utilizes a time-series feature extraction network that specifically learns and predicts the residuals of the physical model, achieving complementary advantages and significantly improving the overall prediction accuracy for complex industrial processes. Then, by calculating the cumulative statistical value of the residual sequence of this high-precision hybrid model for anomaly judgment, small but persistent anomaly signals can be effectively amplified, enabling sensitive detection of low-amplitude, covert attacks or gradual system failures that are difficult to detect using traditional methods. This significantly improves the accuracy and robustness of industrial network anomaly detection. In addition, a scientific and comprehensive model evaluation system is constructed, which not only assesses the model's fitting ability to existing data (i.e., prediction accuracy) but also innovatively introduces a quantitative consideration of the model's own structure (i.e., complexity). The fitness is calculated by weighted summation of the two factors, which effectively suppresses overfitting during model selection. This avoids the algorithm from selecting models that, while perfectly fitting the training data, have excessively complex structures. The resulting physical model is not only accurate but also simpler, easier to understand and interpret, thus significantly improving its generalization ability and ensuring stable and reliable performance when faced with new and unseen data. Furthermore, an automated, data-driven physical model construction method is implemented. By simulating the mechanisms of "survival of the fittest" and "genetic variation" in biological evolution, it can automatically search for and discover potential and explicit mathematical relationships between variables in industrial network data, thereby generating models with clear physical meaning. This solves the technical problems of traditional physical modeling methods, which heavily rely on domain expert knowledge, are time-consuming and labor-intensive, and are difficult to adapt to complex systems. At the same time, by introducing constraints on complexity into the fitness function, it ensures that the final physical model not only has high prediction accuracy but also a simple structure and strong generalization ability, effectively avoiding overfitting. This provides a solid and interpretable foundation for building a highly robust anomaly detection system.Furthermore, a well-structured and effective model fusion framework was constructed. By introducing first and second weights, a flexible and adjustable mechanism was provided to balance the determinism of the physical mechanism with the fitting ability of the data-driven model, avoiding the instability that might result from a rigid combination of the two models. This weighted accumulation fusion method provides a highly interpretable prediction baseline for the physical model, while the residual prediction model serves as its "correction term," specifically compensating for the physical model's shortcomings in describing nonlinear and time-varying characteristics. The resulting hybrid physical model thus combines the advantages of both models, significantly improving the prediction accuracy for complex industrial processes and enhancing the model's robustness and adaptability to different operating conditions. Additionally, a dynamic cumulative monitoring mechanism for industrial network data was constructed. This involves modeling the statistical characteristics of the residual sequence, then standardizing the residuals at each time step to eliminate scaling effects. By integrating historical and current information using a cumulative summation method with drift adjustment, the resulting cumulative statistics can not only effectively identify significant abrupt changes in the data, but also sensitively capture problems such as slow-changing system performance degradation or operating condition drift that are difficult to detect with traditional thresholding methods. This design makes the detection method more robust and sensitive in complex industrial network environments, providing strong technical support for ensuring the stable operation and predictive maintenance of industrial systems. Finally, a clear, reliable, and automated decision-making process is established, transforming the complex, continuous sequence of cumulative statistics into a simple and clear binary judgment result (normal or abnormal) by comparing it with a clear anomaly detection range. This decision-making mechanism based on statistical control limits has higher robustness and sensitivity than simple instantaneous threshold judgment because it acts on a signal that has been amplified over time. This makes the final anomaly detection result not only highly accurate, but also able to effectively distinguish between statistically random noise and persistent real system anomalies, thus providing timely, accurate, and actionable decision support for the security protection of industrial networks. This application also provides an electronic device, including: At least one memory; At least one processor; At least one program; The program is stored in a memory, and the processor executes the at least one program to implement the above-described method for detecting anomalies in industrial network data. The electronic device can be any smart terminal, including mobile phones, tablets, personal digital assistants (PDAs), and in-vehicle computers.

[0137] Please see Figure 10 , Figure 10The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes: The processor 1001 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application. The memory 1002 can be implemented in the form of ROM (Read-Only Memory), static storage device, dynamic storage device, or RAM (Random Access Memory). The memory 1002 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1002 and is called and executed by the processor 1001 to execute the industrial network data anomaly detection method of the embodiments of this application. Input / output interface 1003 is used to implement information input and output; The communication interface 1004 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 1005 transmits information between various components of the device (e.g., processor 1001, memory 1002, input / output interface 1003, and communication interface 1004); The processor 1001, memory 1002, input / output interface 1003 and communication interface 1004 are connected to each other within the device via bus 1005.

[0138] This application also provides a storage medium, which is a computer-readable storage medium, storing a computer program that, when executed by a processor, implements the above-described method for detecting anomalies in industrial network data.

[0139] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0140] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0141] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0142] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0143] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0144] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0145] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0146] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. The coupling or direct coupling or communication connection between the shown or discussed units may be through some interfaces, or indirect coupling or communication connection between the apparatus or units, and may be electrical, mechanical, or other forms.

[0147] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0148] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0149] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0150] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. A method for detecting anomalies in industrial network data, characterized in that, The method includes: Acquire industrial network data and generate a physical model corresponding to the industrial network data; The industrial network data is input into a time-series feature extraction network for feature prediction to obtain residual prediction values. Based on the physical model and the residual prediction values, a hybrid physical model is obtained. Based on the hybrid physics model, a hybrid predicted value for the industrial network data is obtained, and based on the difference between the hybrid predicted value and the actual observed value of the industrial network data, a residual sequence is obtained. Based on the residual sequence, the cumulative statistical value of the industrial network data is calculated, and the anomaly detection result of the industrial network data is obtained based on the cumulative statistical value.

2. The method for detecting anomalies in industrial network data according to claim 1, characterized in that, The generation of the physical model corresponding to the industrial network data includes: Obtain multiple initial physical models of the industrial network data; The fitness is calculated based on the prediction accuracy and complexity of each initial physical model. Based on the fitness of each initial physical model, a target physical model is selected from multiple initial physical models, and cross-mutation is performed on multiple initial physical models to obtain an updated physical model. The updated physical model is then used as the new initial physical model, and multiple cross-mutation updates are performed. The target physical model is updated during the cross-mutation update process. The target physical model updated by the last crossover mutation is used as the physical model.

3. The method for detecting anomalies in industrial network data according to claim 2, characterized in that, The fitness calculated based on the prediction accuracy and complexity of each initial physical model includes: Substitute the industrial network data into each of the initial physical models to obtain initial physical prediction values; The prediction accuracy of each initial physical model is obtained based on the difference between the initial physical prediction value and the actual observation value of the industrial network data. Based on the expression composition parameters of each initial physical model, the corresponding complexity is determined, wherein the expression composition parameters include at least one of the following: number of symbol nodes, tree depth, and operator type. The fitness of each initial physical model is obtained based on a weighted sum of the prediction accuracy and the complexity.

4. The method for detecting anomalies in industrial network data according to claim 1, characterized in that, The training process of the temporal feature extraction network includes: Obtain the training physical model obtained from the training industrial network data, and determine the training prediction residual of the training physical model; The training industrial network data is input into the temporal feature extraction network for feature extraction to obtain multi-scale features, and the multi-scale features are integrated to obtain the training residual prediction value. The training prediction residual is used as the learning target for the training residual prediction value, and the network parameters of the temporal feature extraction network are adjusted accordingly.

5. The method for detecting anomalies in industrial network data according to claim 1, characterized in that, The hybrid physical model, derived based on the physical model and the residual prediction values, includes: Obtain the first and second weights; The hybrid physical model is obtained by summing the product of the first weight and the physical model, the second weight and the residual prediction value.

6. The method for detecting anomalies in industrial network data according to claim 1, characterized in that, The residual sequence includes residual values ​​corresponding to multiple detection times. The calculation of the cumulative statistical values ​​of the industrial network data based on the residual sequence includes: Determine the mean and standard deviation of the residual sequence; The current residual ratio is obtained by dividing the difference between the residual value at the current detection time and the mean of the residual sequence by the standard deviation of the residual sequence. The cumulative statistical value of the industrial network data at the current detection time is obtained by subtracting the drift detection parameter from the sum of the cumulative statistical value of the previous detection time and the current residual ratio.

7. The method for detecting anomalies in industrial network data according to claim 1, characterized in that, The anomaly detection results of the industrial network data obtained based on the cumulative statistical values ​​include: Obtain the anomaly detection range; When the cumulative statistical value is not within the anomaly detection range, an anomaly detection result is generated to characterize the industrial network data as abnormal; When the cumulative statistical value is within the anomaly detection range, an anomaly detection result is generated that indicates the industrial network data is normal.

8. An anomaly detection device for industrial network data, characterized in that, The device includes: The physical model generation module is used to acquire industrial network data and generate a physical model corresponding to the industrial network data. The hybrid physical model generation module is used to input the industrial network data into a time-series feature extraction network for feature prediction, obtain residual prediction values, and obtain a hybrid physical model based on the physical model and the residual prediction values. The residual sequence generation module is used to obtain the mixed predicted value of the industrial network data based on the mixed physical model, and to obtain the residual sequence based on the difference between the mixed predicted value and the actual observed value of the industrial network data. An anomaly detection module is used to calculate the cumulative statistical value of the industrial network data based on the residual sequence, and to obtain the anomaly detection result of the industrial network data based on the cumulative statistical value.

9. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the industrial network data anomaly detection method according to any one of claims 1 to 7.

10. A storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the method for detecting anomalies in industrial network data as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Internet of Things time series data anomaly detection method and related equipment thereof

    CN111767930A

  • Industrial plant monitoring

    CN115039047A

  • Industrial fault diagnosis method and system based on intelligent causal correction

    CN120217262A

  • Self-adaptive predictive maintenance system and method for industrial equipment

    CN120634526A

Cited By

  • Dongle connection state monitoring method and system based on 5G RedCap

    CN121901059A

  • A Dongle Connection Status Monitoring Method and System Based on 5G RedCap

    CN121901059B

  • Training methods, devices, equipment, and media for predicting water and soil pollution remediation outcomes

    CN122412966A