Unsupervised industrial continuous casting machine time series data anomaly detection method

By combining the distillation learning method of Mamba network and encoder, a teacher-student model system is built, which solves the problems of low efficiency and insufficient real-time performance of unsupervised industrial anomaly detection in the timing data processing of continuous casting equipment, and realizes efficient and accurate abnormality detection and intelligent equipment monitoring, improving production efficiency and safety.

CN120337049APending Publication Date: 2025-07-18DONGHUA UNIV
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510303105.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing unsupervised industrial anomaly detection technology is inefficient, has high latency, high cost when processing time sequence data of continuous casting equipment, and it is difficult to effectively process high-dimensional features and complex patterns, resulting in poor adaptability of the model to new data, large demand for computing resources, and insufficient real-time performance.

Method used

Combining the method of distillation learning of Mamba network and encoder, the key sensor data of the continuous casting machine is monitored in real time through edge computing devices, a teacher-student model system is built, data correlation is analyzed using the fast greedy equivalent search method, unsupervised anomaly detection is performed, and lightweight student models are deployed at the edge end, and real-time monitoring and analysis is carried out in combination with the digital twin platform.

Benefits of technology

It improves the accuracy and computing efficiency of abnormal detection, meets real-time requirements, optimizes the use of hardware resources, realizes intelligent monitoring and maintenance of equipment, and improves production efficiency and equipment safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337049A_ABST
    Figure CN120337049A_ABST
Patent Text Reader

Abstract

The invention relates to an unsupervised industrial continuous casting machine time series data anomaly detection method. The method comprises the following steps: monitoring and analyzing key sensor data of continuous casting machine equipment; installing and configuring edge computing equipment, performing digital twin platform deployment, performing data preprocessing on sensor data, performing correlation analysis on the preprocessed data by using a rapid greedy equivalent search method, constructing encoder and decoder structures, introducing a Mama network to construct a teacher model, and training; based on a distillation learning principle, migrating knowledge of the teacher model to the student model; and performing anomaly detection after the student model is trained, deploying at an edge end after an expected standard is reached, and performing alarm monitoring by a digital twin platform. The method solves the problems of low efficiency, high delay and high cost when the existing unsupervised industrial anomaly detection technology is used for processing the time sequence data of the continuous casting equipment, guarantees the anomaly detection accuracy, improves the calculation efficiency of the model, and meets the requirements of real-time performance and calculation resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of industrial intelligent manufacturing applications, and particularly to a method for unsupervised anomaly detection of time-series data of a steel continuous casting machine by combining an encoder and a Mamba network for distillation learning in a digital twin system, so as to improve production efficiency and ensure equipment safety. Background Art

[0002] As an important device in the field of iron and steel metallurgy, the continuous casting machine plays an indispensable role, and its operating state directly affects production efficiency and product quality. In actual industrial production, in order to ensure the normal operation of the continuous casting machine, it is necessary to monitor and analyze the generated time-series data in real time, so as to detect and handle possible abnormal situations in time. However, due to the complex operation of the continuous casting machine, the generated time-series data has a large number of features and complex patterns, and conventional monitoring and analysis methods are often difficult to effectively process these data, thus potentially missing some potential abnormal situations.

[0003] In recent years, digital twin technology has been widely applied in the industrial field. By constructing a digital model corresponding to the real device in the virtual space, digital twin can monitor the operating state of the device in real time, predict the future behavior of the device, and identify potential anomalies and faults through simulation and analysis. For complex industrial devices such as continuous casting machines, digital twin technology can obtain multi-dimensional data of the device in real time and conduct in-depth analysis, becoming an important tool for improving production efficiency and equipment reliability. However, due to the complex operating state of the device and the huge amount of time-series data, how to extract effective anomaly features from them remains an urgent problem to be solved.

[0004] Currently, some supervised learning-based methods are often used in the industrial field for equipment anomaly detection, especially in digital twin systems. However, such methods often rely on a large amount of labeled data, and in the actual production environment, the acquisition of labeled data is very difficult and costly. In addition, supervised learning-based methods are also prone to overfitting problems, resulting in the model overfitting to the training data and having poor adaptability to new and unknown data. Therefore, how to effectively detect anomalies in the unlabeled time-series data of the continuous casting machine and then improve the monitoring accuracy of the digital twin model and equipment safety is an important issue that needs to be solved urgently.

[0005] Unsupervised anomaly detection methods do not rely on labeled data and can learn potential patterns from raw data to automatically identify abnormal behaviors. Common unsupervised methods include clustering-based algorithms and reconstruction-based algorithms. Clustering-based methods rely on the distribution characteristics of time series data, but they are difficult to handle high-dimensional features and complex non-linear relationships; reconstruction-based methods detect anomalies by learning historical data and predicting new data sequences, and then calculating the error between the prediction results and the actual data. However, these methods have a high computational complexity when dealing with long sequences. Especially when using models with self-attention mechanisms such as Transformer, the computational complexity increases exponentially with the increase in sequence length. Therefore, how to design an efficient unsupervised anomaly detection model to handle long sequence data and complex patterns has become an important challenge for digital twin systems in industrial applications.

[0006] Although unsupervised anomaly detection methods do not require data annotation, their model generalization ability is limited, and when dealing with tasks with high industrial real-time requirements, there are often problems such as large computational resource requirements and response delays. How to overcome these challenges and improve the computational efficiency and real-time performance of the model is the key to promoting the application of digital twin technology in industrial equipment monitoring. Summary of the Invention

[0007] Aiming at the problems of low efficiency, high latency, and high cost in the existing unsupervised industrial anomaly detection technology when dealing with the time series data of continuous casting equipment, an unsupervised industrial continuous casting machine time series data anomaly detection method is proposed, which combines the Mamba network and the encoder for distillation learning to detect anomalies in the unsupervised industrial time series data of continuous casting equipment.

[0008] The technical solution of the present invention is as follows:

[0009] An unsupervised industrial continuous casting machine time series data anomaly detection method, comprising the following steps:

[0010] Step 1: Monitor and analyze the key sensor data of the mold level and roll current of the continuous casting machine equipment;

[0011] Step 2: Edge device configuration: Install and configure edge computing devices at the production line site. The edge computing devices are connected to the sensors of the continuous casting equipment through a wired network; the mold level and roll current sensors on the continuous casting equipment are connected to the edge computing devices using the industrial data stream transmission protocol MQTT; the sensors convert the collected real-time data into standardized digital signals and send them to the edge computing devices through the MQTT protocol;

[0012] Step 3: Deployment of Digital Twin Platform: The platform is containerized and deployed using Docker on the server. Install Docker and Docker Compose on the server. Create a Dockerfile according to the requirements of the digital twin platform to define the platform's environment, dependencies, and configurations. Use Docker Compose to define the services required for multiple tasks. Create a docker-compose.yml file. Then use docker-compose to start all services and check if all services are running. The edge side collects and preprocesses the data of physical devices or sensors and transmits it to the digital twin platform for further data storage and real-time visualization analysis and display on the platform.

[0013] Step 4: Data Preprocessing: First, these data need to be cleaned and preprocessed, including filling missing values, removing noise, and standardization operations.

[0014] Step 5: Correlation Analysis between Data: Use the Fast Greedy Equivalence Search method FGES to perform correlation analysis on the preprocessed time series data of the industrial continuous caster. Select these dimensional data with strong correlation relationships, that is, causal relationships and dependencies, as the input of the model.

[0015] Step 6: Training of Teacher Model: Build an encoder and decoder structure, and introduce the Mamba network to build the teacher model. Train it on the processed dataset using a large amount of unlabeled data. The encoder of the teacher network is the part of the network responsible for processing the input data, and its task is to extract the key information in the input data. In the teacher model, the encoder receives the time series data from various sensors of the continuous caster as input and encodes these data to extract the deep features of these data. The decoder consists of multiple Mamba network modules. Using the output of the encoder as input, it performs calculations through multiple Mamba network modules and finally generates the prediction results.

[0016] Step 7: Training of Student Model: Based on the principle of distillation learning, transfer the knowledge of the teacher model to a lightweight student model. This process can be achieved by having the student model imitate the output of the teacher model or by having the student model and the teacher model jointly optimize an objective function. During the training process, the student model only simulates or replicates the output of the teacher model based on normal sample data. Each student model is trained independently. When the student model cannot correctly predict the data outside the normal data distribution range, these data are considered abnormal.

[0017] Step 8: Anomaly Detection: After the student model is trained, it can be used to predict new time-series data. By comparing the prediction results with the original data, abnormal situations can be detected. For minor anomalies, the teacher model, trained with a large amount of historical data, can identify complex anomaly features, while the student model is trained by mimicking the output of the teacher model, i.e., soft labels, thus learning the anomaly detection ability of the teacher model. The student model uses a loss function to minimize the difference from the teacher model and gradually masters the skills of identifying minor anomalies. In practical applications, the student model predicts new time-series data, calculates the difference by comparing with the prediction results of the teacher model, and determines whether there are anomalies. Finally, the detection results are output based on a preset threshold range to ensure accurate identification of abnormal situations. After completing the anomaly detection of the continuous casting equipment time-series data, the detection results need to be systematically evaluated to ensure that the performance of the model meets the expected goals under various abnormal conditions. If the evaluation results do not meet the requirements, the model parameters can be further adjusted or the distillation learning process can be optimized to improve the accuracy and efficiency of anomaly detection. Through continuous iteration and optimization, the optimal model configuration is finally determined to accurately identify various abnormal situations during the operation of the continuous casting equipment.

[0018] Step 9: Edge Deployment of the Student Model: When the model performance reaches the expected standard, the lightweight student model optimized through distillation learning is deployed to the edge computing device of the continuous casting equipment. The student model deployed at the edge is responsible for the preliminary processing of the real-time data collected from the device sensors, including data cleaning, filtering, and formatting. The student model will use its compressed parameters and simplified structure to quickly detect anomalies in the operating state of the device. Once an abnormal situation is detected, the edge device will respond immediately.

[0019] Step 10: Alarm Monitoring on the Digital Twin Platform: After the student model conducts a preliminary analysis of the edge device data, the edge side will upload some processing results or key data to the digital twin monitoring platform. This platform receives data from multiple edge devices and conducts in-depth analysis, trend prediction, and historical data backtracking. By integrating and analyzing these data, the digital twin platform can provide operators with a comprehensive monitoring view and display the global state of the device. The platform has a real-time monitoring function, which can show the operating health status of the device and conduct fault diagnosis, predictive maintenance, and early warning based on the analysis results. If the digital twin platform detects potential faults or trend anomalies in the edge device, the platform will automatically trigger an alarm to prompt the operator to take necessary measures. In addition, the platform also supports long-term trend analysis of the device, helping managers make data-based decisions, optimize device performance, reduce downtime, and improve production efficiency.

[0020] Furthermore, in step 3, the role of the digital twin platform is to establish a virtual model of the physical system or device by acquiring its data in real time, and then analyze and optimize it to support decision-making and operations. The digital twin platform includes the following functions: real-time data collection and integration; data storage and management; digital model and simulation; model analysis and decision support; visual display to present device status, trends, and alarm information; security and permission management to ensure the security of data and systems.

[0021] Furthermore, in step 4, interpolation method, mean filling method or machine learning-based filling method is used to fill in the missing values; filters or noise suppression techniques are used to remove high-frequency noise in the data; the data is standardized or normalized to eliminate the influence of dimensions and ensure that different feature data have similar scales; in addition, outlier detection algorithms are used to identify and remove abnormal data.

[0022] Furthermore, in step 8, during the evaluation process, the accuracy, recall rate, and F1 score of each student network in the anomaly detection task are calculated to judge the precision of each student network; the computational resource consumption of each student network on the edge device is evaluated, including memory usage and processing time; if the precision effect does not meet expectations, hyperparameter tuning can be performed and then knowledge distillation can be carried out; if the resource consumption is too large, model pruning can be performed.

[0023] Furthermore, in step 9, a secure network connection and data transmission protocol MQTT are established between the edge device and the digital twin platform to achieve two-way transmission and real-time update of model data and status information; the running status and prediction data of the virtual model are obtained in real time.

[0024] Furthermore, in step 1, the data collected by the mold part includes: the outlet flow rates of the inner arc and outer arc of the mold water wide surface, the outlet flow rates of the right and left sides of the mold water narrow surface, and the mold liquid level data; for the upper and lower guide roll currents, the upper roll current data in the bending section area and the upper and lower guide roll current data in the straightening section area are collected.

[0025] Furthermore, step 4 is specifically as follows: data processing is performed on the data of the continuous casting equipment input; first, it is necessary to ensure that the time information is correct, secondly, check whether there are missing values, and finally, standardize the data of the remaining dimensions except the time column; the output data is where X t is the input time dimension, is the other dimensions of the input; after removing the time dimension from X, it can be expressed as After standardization, the processing result is X′ s ;

[0026]

[0027] Among them, m represents the m-th data.

[0028] Furthermore, step 5 is specifically as follows: The preprocessed time series data X′ S Find the causal relationships between data through the FGES algorithm; Use an unsupervised method to analyze the correlations between various dimensions by constructing a graph structure and relying on data relationships; Specifically, find the time series X′ S The steps for finding causal relationships are as follows:

[0029] Step 5.1, Initialize the graph structure: First, initialize an empty graph with no edges.

[0030] Step 5.2, Forward search and add edges: Start forward search from the empty graph to add edges; For each possible edge (i, j), calculate the Bayesian Information Criterion BIC of the graph G after adding this edge:

[0031] ΔBIC add = BIC(G+(i,j)) - BIC(G)

[0032] Among them, G is the current graph, representing the causal relationship graph containing the already added edges; Add the edge (i, j) that can make ΔBIC add positive and the largest to the graph until no edge addition can make ΔBIC add > 0;

[0033] Step 5.3, Backward search and delete edges: Then start backward search from the current graph G to delete edges; For each existing edge (i, j), calculate the Bayesian Information Criterion BIC after deleting this edge:

[0034] ΔBIC del = BIC(G-(i,j)) - BIC(G)

[0035] Delete the edge (i, j) that can make ΔBIC del non - negative and the largest from the graph until no edge deletion can make ΔBIC del > 0;

[0036] Step 5.4, Deduce the time series X′ from the graph structure (S-C) : First, extract the significantly influential nodes from the causal relationship graph; These key variables may be root nodes or other variables that have a direct impact on the target variable in the causal relationship graph; Then, deduce other variables affected by these variables according to the dependency relationships in the causal relationship; Next, screen out the most significant variables from the deduced variables according to the screening rules of the correlation threshold or causal strength; Finally, combine the screened variables in chronological order to form the final time series X′(S-C) , which reflects the dependence structure and causal relationships among variables in the causal relationship diagram.

[0037] Further, step 6 is specifically as follows: construct a Mamba autoencoder network for distillation learning for the time series data of the continuous casting equipment; the Mamba autoencoder network includes a teacher-student model, which is divided into the following steps: First, the time series data X′ S-C is input into the teacher network. After the features are extracted by the encoder, they are passed to the decoder for data reconstruction; the goal of the teacher network is to be trained by optimizing the reconstruction error and generate prediction results; then, the student model learns from the teacher network through knowledge distillation, using the same encoder and decoder structures, but its goal is to minimize the difference from the prediction results of the teacher model as much as possible; each student model makes predictions independently and outputs the anomaly detection results; by comparing the prediction results of the student models, the digital twin platform can detect abnormal behaviors and give feedback; the specific processing process is as follows:

[0038] The input time series X′ S-C is mapped into a low-dimensional latent representation space Z through the encoder, and the result Feature of the encoder is output:

[0039] Feature = f encoder (X′ S-C )

[0040] where f encoder represents the encoder function; then Feature is reconstructed back into the original data form through the Mamba network as the decoder:

[0041]

[0042] where f decoder is the function of the Mamba network decoder, and X′ R(S-C) is the reconstruction output of the teacher network;

[0043] The Mamba network extracts features from the input time series data through convolutional layers and introduces non-linear transformations through activation functions to enhance the expression ability of the network; subsequently, the hybrid state encoder encodes the extracted features into hidden states, captures the context relationship of the time series, and models the dynamic changes in the time series through the state space model SSM; finally, the hybrid state decoder decodes these hidden states to restore the original data structure;

[0044] For the output of the encoder, it can be regarded as a triple with dimensions (Batch, V, D), where Batch is the batch size, V is the number of variables, and D is the hidden dimension; first, expand the input time series to a higher dimension ED through linear projection of the encoder result Feature:

[0045] x, z = Linear(Feature)

[0046] where x and z represent the feature representations after linear projection; Linear is the linear projection operation that maps Feature to dimension ED;

[0047] where the dimensions of x and z are (Batch, V, ED); then use the convolutional layer and activation function to process the linear projection result:

[0048] x′ = σ(Conv(x))

[0049] conv(x) is the convolution operation on x, σ is the activation function, and x′ is the new feature representation obtained after convolution and activation, with dimensions still (Batch, V, ED);

[0050] The state space model is used to describe these state representations and predict the next state based on the input; at time t, map the input sequence x′(t) to the latent state representation h(t) and derive the predicted output sequence y(t); then the future state can be predicted by the following two equations:

[0051] h′(t) = Ah(t) + Bx′(t)

[0052] y(t) = Ch(t) + Dx′(t)

[0053] where A is the state transition matrix that describes how the state changes from time t to time t + 1; B is the input matrix that represents the influence of the external input x′(t) on the current state; C is the observation matrix that describes how to calculate the output y(t) from the current state h(t); D is the direct transfer matrix that describes the relationship between how the input x′(t) directly affects the output y(t); h′(t) is the state equation that represents how the state changes according to the influence of the input on the state; the output equation y(t) describes how the state is transformed into the output through matrix C and how the input affects the output;

[0054] The current SSM is a continuous time-invariant model and needs to be discretized for processing actual data; according to the differential equation, sampling using the zero-order hold rule can obtain the discretized result:

[0055]

[0056] where is the state transition matrix after discretization, exp represents the exponential operation of the matrix, and ΔA is the time step; is the input matrix after discretization, I is the identity matrix, representing the discretization process of B; in the formula, Δ is the selectivity factor, and its shape is the same as the input;

[0057] Use the "selective information propagation mechanism" to control how information propagates in the sequence dimension; discretize the parameters Δ, A, B, and C in the SSM to transform the time-invariant model into a time-variable model; the discretized matrices and represent state transition and input mapping, and C is the observation matrix; assume D is 0, which means the input only indirectly affects the output through the state:

[0058]

[0059] After performing residual processing on the result y output output by the SSM module, and then passing through the activation function and linear projection, we obtain

[0060] Furthermore, step 7 is specifically as follows: Multiple lightweight student networks use the latent representation and reconstruction output of the teacher network for training; the encoding and decoding functions of the student network are G encoder and G decoder , and after inputting X′ S-C , the result obtained after training by the student network is

[0061]

[0062] To make the reconstruction result of the student network close to the output of the teacher network, the following loss function is used for training:

[0063]

[0064] When the loss value reaches the optimum, the distillation learning will no longer be performed, and the prediction result will be output from the trained student network model and compared with the threshold set according to industrial experience to determine whether an anomaly has occurred; what the teacher network transmits to the student network are the latent representation and reconstruction output, and multiple lightweight student networks learn based on these contents, so as to be able to perform anomaly detection efficiently on edge devices.

[0065] The beneficial effects of the present invention are as follows:

[0066] Compared with traditional methods, the present invention transfers the knowledge of a pre-trained complex teacher model to multiple lightweight student models through a distillation learning strategy, thereby improving the computational efficiency of the model while ensuring the accuracy of anomaly detection, meeting the requirements of real-time performance and computational resources. To further enhance the performance of the model, the present invention introduces the Mamba network (a deep learning architecture based on a selective state space model), and utilizes its selective mechanism, allowing the model to dynamically adjust the state according to input features, selectively propagate or ignore information. This mechanism not only improves the accuracy of anomaly detection but also optimizes the real-time monitoring ability of the device operating state. In this way, the present invention can more efficiently process high-dimensional and complex time-series data, providing more powerful data processing and analysis support for the digital twin platform. In addition, the present invention fully considers the computational and memory resource limitations of the hardware, optimizes the computational efficiency and memory usage of the model, making it more suitable for applications in edge computing devices and real-time monitoring systems. This innovative method will provide strong support for the intelligent monitoring and maintenance of continuous casting equipment, helping to improve production efficiency, ensure equipment safety, and promote the in-depth application of digital twin technology in the industrial field.

[0067] Due to the limited computational resources of edge devices, adopting lightweight student models can ensure the rapid response and efficient processing of the model at the device end. At the same time, through seamless connection with the digital twin system, the edge device feeds back the detected anomaly information and prediction results to the digital twin platform, and the digital twin model conducts further simulation and optimization analysis based on this real-time data. The digital twin platform can provide intelligent support such as state prediction, fault diagnosis, and maintenance decision-making for the device, improving the device management efficiency.

[0068] In this way, the deployment at the edge end not only improves the real-time anomaly detection ability of continuous casting equipment but also realizes the intelligent operation and predictive maintenance of the device through interaction with the digital twin. The system can detect and handle abnormal situations in a timely manner in the actual operating environment of continuous casting equipment, significantly improving the operating efficiency and safety of the equipment, and promoting the application of digital twin technology in the intelligent management of industrial equipment. Brief Description of the Drawings

[0069] Figure 1 It is the digital twin architecture diagram of the present invention for steelmaking equipment;

[0070] Figure 2 It is the network structure diagram of the present invention for anomaly detection of continuous casting equipment time-series data using the Mamba network and distillation learning;

[0071] Figure 3 It is the structure diagram of the Mamba network used in the present invention;

[0072] Figure 4The figure is a flowchart of a specific implementation of the present invention. DETAILED DESCRIPTION

[0073] The present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.

[0074] An unsupervised anomaly detection method for time series data of an industrial continuous casting machine includes the following steps:

[0075] Step 1: The present invention collects data from key parts of the continuous casting machine equipment where abnormalities are prone to occur. Specifically, key sensor data such as the crystallizer liquid level (usually made of high-temperature resistant materials, directly in contact with molten steel, and subjected to high temperature and thermal shock), the guide roller current (the current consumed by the motor of the roller used to guide the movement of the cast billet), etc. are selected for monitoring and analysis. The operating state of the crystallizer directly affects the quality and production efficiency of the billet, and abnormal conditions may cause problems with the internal structure or surface quality of the billet. The guide roller is an important component of the continuous casting machine, and its current state is crucial to the normal operation of the equipment. Abnormal current may cause the guide roller to be unstable or even shut down, seriously affecting the production schedule and equipment operation. Therefore, the present invention focuses on selecting data from the crystallizer and the current parts of each guide roller for collection and detection, and transmits these real-time collected data to the edge end.

[0076] Step 2: Edge configuration: Install and configure the NVIDIA Jetson Xavier device on the production line, which is connected to the sensors of the continuous casting equipment through a wired network. The crystallizer level and guide roller current sensors on the continuous casting equipment (such as liquid level sensors, temperature sensors, pressure sensors, current sensors, etc.) are connected to the edge computing device (NVIDIA Jetson Xavier) using the industrial data stream transmission protocol MQTT (a lightweight message transmission protocol commonly used for communication between IoT devices and suitable for real-time transmission of small-scale data). The sensor converts the collected real-time data into a standardized digital signal and sends it to the edge computing device through the MQTT protocol. Data transmission is carried out through a wired network (such as Ethernet or industrial bus protocols such as Modbus TCP, Ethernet / IP, etc.) to ensure a stable, fast, high-frequency, and low-latency data stream.

[0077] Step 3: Deployment of Digital Twin Platform: The role of the digital twin platform is to establish a virtual model of a physical system or device by acquiring its data in real time, and conduct analysis and optimization to support decision-making and operations. The platform mainly includes the following functions: real-time data collection and integration; data storage and management; digital model and simulation; model analysis and decision support; visualization display to present device status, trends, and alarm information; security and permission management to ensure the security of data and systems. The platform is containerized using Docker and deployed on a 4090 server with 24GB of video memory. Install Docker and Docker Compose on the server. Create a Dockerfile according to the requirements of the digital twin platform to define the environment, dependencies, and configurations of the platform. Use Docker Compose to define multiple services such as databases, front-ends, back-ends, etc. to simplify container management. Create a docker-compose.yml file. Then use docker-compose to start all services and check if all services are running. The edge side collects and preprocesses the data of physical devices or sensors and transmits it to the digital twin platform for further data storage and real-time visualization analysis and display.

[0078] Step 4: Data Preprocessing: The time-series data generated by industrial continuous casting machines may contain noise, missing values, or other outliers. First, these data need to be cleaned and preprocessed, including operations such as filling missing values, eliminating noise, and standardization. Specifically, interpolation methods, mean filling methods, or machine learning-based filling methods can be used to fill missing values; filters or other noise suppression techniques can be used to remove high-frequency noise in the data; the data can be standardized or normalized to eliminate the influence of dimensions and ensure that different feature data have similar scales. In addition, outlier detection algorithms (such as Z-Score, IQR, etc.) can be used to identify and remove abnormal data to ensure the quality and reliability of the data input into the analysis model.

[0079] Step 5: Correlation Analysis between Data: Use the Fast Greedy Equivalence Search method (FGES) to conduct correlation analysis on the preprocessed time-series data of industrial continuous casting machines. Through the FGES algorithm, the internal relationships between data dimensions can be automatically explored and mined, and significantly correlated variables can be identified. The FGES algorithm optimizes the Bayesian network structure by adding, deleting, or reversing edges in the network, thereby revealing the causal relationships and dependencies between variables. Finally, select the dimension data with strong correlation relationships (i.e., causal relationships and dependencies) as the input of the model to improve the accuracy and robustness of the model.

[0080] Step 6: Teacher Model Training: Construct an encoder and a decoder structure, and introduce the Mamba network to build the teacher model. Train it using a large amount of unlabeled data on the processed dataset. The encoder of the teacher network is the part of the network responsible for processing the input data, and its task is to extract the key information in the input data. In the teacher model, the encoder receives the time-series data from various sensors of the continuous casting machine as input, encodes this data, and extracts the deep features of the data. The decoder consists of multiple Mamba network modules. Taking the output of the encoder as input, it performs calculations through multiple Mamba network modules and finally generates the prediction results.

[0081] Step 7: Student Model Training: Based on the principle of distillation learning, transfer the knowledge of the teacher model to a lightweight student model. This process can be achieved by having the student model imitate the output of the teacher model or by having the student model and the teacher model jointly optimize an objective function. During training, the student model only simulates or replicates the output of the teacher model based on normal sample data. Each student model is trained independently, which can improve the robustness of the model when facing different problems. When the student model cannot correctly predict the data that goes beyond the normal data distribution range (this data is called data outside the normal manifold), these data are considered abnormal.

[0082] Step 8: Anomaly Detection: After the student model is trained, it can be used to predict new time-series data. By comparing the prediction results with the original data, abnormal situations can be detected. For minor anomalies, the teacher model, trained with a large amount of historical data, can identify complex anomaly features, while the student model is trained by imitating the output of the teacher model (i.e., "soft labels") to learn the anomaly detection ability of the teacher model. The student model uses a specific loss function (such as KL divergence) to minimize the difference from the teacher model and gradually master the skills of identifying minor anomalies. In practical applications, the student model predicts new time-series data, calculates the difference by comparing with the prediction results of the teacher model, and determines whether there are anomalies. Finally, based on a preset threshold range, the detection results are output to ensure accurate identification of abnormal situations. After completing the anomaly detection of the time-series data of the continuous casting equipment, it is necessary to systematically evaluate the detection results to ensure that the performance of the model meets the expected goals in various abnormal situations. If the evaluation results do not meet the requirements, the model parameters can be further adjusted or the distillation learning process can be optimized to improve the accuracy and efficiency of anomaly detection. During the evaluation process, indicators such as accuracy, recall rate, and F1 score are used to comprehensively analyze the stability and robustness of the model. Through continuous iteration and optimization, the optimal model configuration is finally determined to enable accurate identification of various abnormal situations during the operation of the continuous casting equipment.

[0083] Step 9: Edge Deployment of the Student Model: To reduce latency and achieve real-time response, after the model performance reaches the expected standard, the lightweight student model optimized through distillation learning is deployed to the edge computing device of the continuous casting equipment. The student model deployed at the edge is responsible for the preliminary processing of real-time data collected from device sensors, including data cleaning, filtering, and formatting. The student model will utilize its compressed parameters and simplified structure to quickly detect anomalies in the operating state of the device. For example, through real-time analysis of key data such as temperature, pressure, and speed, the model can identify whether there are faults or abnormal behaviors in the device. Once an abnormal situation is detected, the edge device will immediately respond, such as adjusting the operating parameters of the device, activating protection measures, or triggering an alarm mechanism to notify the operator for timely handling. In this way, the edge computing device can work independently, reducing the dependence on the central server and ensuring the safety and stability of the device.

[0084] Step 10: Alarm Monitoring of the Digital Twin Platform: After the student model conducts preliminary analysis on the edge device data, the edge side will upload some processing results or key data (such as anomaly detection results, device status changes, etc.) to the digital twin monitoring platform. This platform receives data from multiple edge devices and conducts in-depth analysis, trend prediction, and historical data backtracking. By integrating and analyzing this data, the digital twin platform can provide operators with a comprehensive monitoring view and display the global status of the device. The platform has a real-time monitoring function, which can display the operating health status of the device and conduct fault diagnosis, predictive maintenance, and early warning based on the analysis results. If the digital twin platform detects potential faults or abnormal trends in the edge device, the platform will automatically trigger an alarm to prompt the operator to take necessary measures. In addition, the platform also supports long-term trend analysis of the device, helping managers make data-based decisions, optimize device performance, reduce downtime, and improve production efficiency.

[0085] Specifically, the present invention mainly proposes an unsupervised anomaly detection method for industrial continuous casting machine time series data, a method combining an encoder and a Mamba network for distillation learning, for the monitoring and maintenance of industrial continuous casting equipment. As Figure 4 shown, it specifically includes the following steps:

[0086] Step 1: First, the present invention collects sensor data for the key parts in the continuous casting machine equipment that are prone to abnormalities, and uploads the collected data to the digital twin platform for real-time monitoring and analysis. Specifically, for the mold liquid level and related parameters, their operating states are collected and detected; at the same time, the upper and lower guide roll currents and their related data are monitored. The main data collected for the mold part includes: the outlet flow rates of the inner arc and outer arc of the wide face of the mold water, the outlet flow rates of the right and left sides of the narrow face of the mold water, and the mold liquid level data; for the upper and lower guide roll currents, the upper roll current data in the S1-S7 interval (the upper roll positions in the S1 to S7 regions) and the upper and lower guide roll current data in the S8-S16 interval (the upper and lower guide roll positions in the S8 to S16 regions) are mainly collected. Among them, according to the continuous casting machine segment structure, the S1-S16 monitoring regions are divided along the casting flow direction. Among them: S1-S7 is the bending section region, and the current of a single upper guide roll is monitored in each region (a total of 7 upper rolls), S8-S16 is the straightening section region, and the currents of the upper and lower guide rolls are monitored simultaneously in each region (a total of 9 regions, 18 guide rolls); the monitoring point number S1 starts from the mold outlet position, and S16 is located in front of the cutting machine.

[0087] Step 2: Install and configure NVIDIA Jetson Xavier devices at the production line site. These devices are connected to the sensors of the continuous casting equipment through a wired network. The mold liquid level and guide roll current sensors on the continuous casting equipment (such as liquid level sensors, temperature sensors, pressure sensors, current sensors, etc.) are connected to the edge computing device (NVIDIA Jetson Xavier) using the industrial data stream transmission protocol MQTT (a lightweight message transmission protocol, commonly used for communication between IoT devices and suitable for real-time transmission of small-scale data). The sensors convert the collected real-time data into standardized digital signals and send them to the edge computing device through the MQTT protocol. The data transmission is carried out through a wired network (such as Ethernet or industrial bus protocols such as Modbus TCP, Ethernet / IP, etc.) to ensure a stable, fast, high-frequency, and low-latency data stream.

[0088] Step 3: The role of the digital twin platform is to establish a virtual model of a physical system or device by obtaining its data in real time, and perform analysis and optimization to support decision-making and operations. The platform mainly includes the following functions: real-time data collection and integration; data storage and management; digital model and simulation; model analysis and decision support; visual display to present device status, trends, and alarm information; security and permission management to ensure the security of data and systems. The platform is containerized using Docker and deployed on a 4090 server with 24GB of video memory. Install Docker and Docker Compose on the server. Create a Dockerfile according to the requirements of the digital twin platform to define the environment, dependencies, and configurations of the platform. Use Docker Compose to define multiple services such as databases, front-ends, back-ends, etc. to simplify container management. Create a docker-compose.yml file. Then use docker-compose to start all services and check if all services are running. The edge side collects and preprocesses the data of physical devices or sensors and transmits it to the digital twin platform for further data storage and real-time visualization analysis and display.

[0089] Step 4: Process the data of the input continuous casting equipment. First, it is necessary to ensure that the time information is correct. Second, check for missing values. Finally, standardize the data of the remaining dimensions except the time column. The output data is where X t is the input time dimension, and is the other dimensions of the input. After removing the time dimension from X, it can be expressed as After standardization, the processed result is X′ s .

[0090]

[0091] Among them, m represents the m-th data.

[0092] Step 5: The preprocessed time series data X′ S Find the causal relationship between data through the FGES algorithm. In this process, when there is a strong correlation between current variables, they will affect each other, ultimately affecting the detection result. Use an unsupervised method to analyze the correlation between each dimension by constructing a graph structure and relying on data relationships. Specifically, the steps to find the causal relationship of the time series X′ S are as follows:

[0093] (1) Initialize the graph structure: First, initialize an empty graph with no edges. (The starting point is the initial state of the algorithm. Specifically, in this causal relationship analysis, it usually starts from an empty graph, that is, a graph with no edges. This is the starting point for searching and adding / deleting edges.)

[0094] (2) Forward search and add edges: Start forward search and add edges from the empty graph. For each possible edge (i, j), calculate the Bayesian Information Criterion (BIC) of the graph G after adding this edge:

[0095] ΔBIC add = BIC(G+(i,j)) - BIC(G)

[0096] where G is the current graph, representing the causal relationship graph that contains the edges that have been added. Add the edge (i, j) that can make ΔBIC add positive and the largest to the graph until there is no edge addition that can make ΔBIC add > 0.

[0097] (3) Backward search and delete edges: Then start backward search and delete edges from the current graph G. For each existing edge (i, j), calculate the Bayesian Information Criterion (BIC) after deleting this edge:

[0098] ΔBIC del = BIC(G-(i,j)) - BIC(G)

[0099] Delete the edge (i, j) that can make ΔBIC del non - negative and the largest from the graph until there is no edge deletion that can make ΔBIC del > 0.

[0100] (4) Derive the time series X′ from the graph structure (S-C) : The above steps mainly focus on optimizing the causal relationship graph by adding and deleting edges. First, extract the significantly influential nodes from the causal relationship graph. These key variables may be root nodes or other variables that have a direct impact on the target variable in the causal relationship graph. Then, based on the dependency relationships in the causal relationship, derive other variables affected by these variables (if a key variable affects a certain variable, determine whether this variable is part of the causal path of the target variable from the graph). Next, according to the screening rules of the correlation threshold or causal strength, screen out the most significant variables from the derived variables. Finally, combine the screened variables in chronological order to form the final time series X′ (S-C) , which reflects the dependency structure and causal connections between variables in the causal relationship graph.

[0101] Step 6: The present invention constructs a Mamba autoencoder network for distillation learning on the time series data of the continuous casting equipment, so as to perform industrial unsupervised anomaly detection tasks. The network structure diagram is as Figure 2 shown. The network structure shown in the figure is a deep learning framework based on the Mamba autoencoder, which consists of a teacher-student model and mainly includes the following steps: First, the time series data X′ S-C is input into the teacher network. After the encoder extracts features, it is passed to the decoder for data reconstruction. The goal of the teacher network is to be trained by optimizing the reconstruction error and generate prediction results. Then, the student model learns from the teacher network through knowledge distillation, adopting the same encoder and decoder structures, but its goal is to minimize the difference from the prediction results of the teacher model as much as possible. Each student model makes predictions independently and outputs the anomaly detection results. By comparing the prediction results of the student models, the digital twin platform can detect abnormal behaviors and give feedback. The specific processing process is as follows:

[0102] The input time series X′ S-C is mapped into a low-dimensional latent representation space Z through the encoder, and the result Feature of the encoder is output:

[0103] Feature = f encoder (X′ S-C )

[0104] where f encoder represents the encoder function. Then Feature is reconstructed back into the original data form through the Mamba network acting as the decoder:

[0105]

[0106] where f decoder is the function of the Mamba network decoder, and X′ R(S-C) is the reconstruction output of the teacher network.

[0107] As Figure 3As shown, the Mamba network extracts features from the input time series data through convolutional layers and introduces non-linear transformations through activation functions to enhance the network's expressive power. Subsequently, the hybrid state encoder encodes the extracted features into hidden states, captures the context relationship of the time series, and models the dynamic changes in the time series through the state space model (SSM). Finally, the hybrid state decoder decodes these hidden states to restore the original data structure. This network structure focuses on the previous window, preserves the sequential attributes of the time series, and thus improves the accuracy of anomaly detection and prediction. The Mamba network focuses on the previous window during the information extraction process, thereby preserving certain sequential attributes. The potential effectiveness of the Mamba network in time series information fusion aims to address the problem of performance degradation or stagnation as the lookback length increases.

[0108] For the output of the encoder, it can be regarded as a triple with dimensions (Batch, V, D), where Batch is the batch size, V is the number of variables, and D is the hidden dimension. First, the result Feature of the encoder is linearly projected to expand the input time series to a higher dimension ED:

[0109] x, z = Linear(Feature)

[0110] where x and z represent the feature representations after linear projection. Linear is the linear projection operation that maps Feature to dimension ED.

[0111] where the dimensions of x and z are (Batch, V, ED). Then, the convolutional layer and activation function are used to process the linear projection result:

[0112] x′ = σ(conv(x))

[0113] conv(x) is the convolution operation on x, σ is the activation function (such as ReLU), and x′ is the new feature representation obtained after convolution and activation, with the dimension still being (Batch, V, ED).

[0114] The state space model is used to describe these state representations and predict the next state based on the input. At time t, the input sequence x′(t) is mapped to the latent state representation h(t), and the predicted output sequence y(t) is derived. Then, the future state can be predicted through the following two equations:

[0115] h′(t) = Ah(t) + Bx′(t)

[0116] y(t) = Ch(t) + Dx′(t)

[0117] Where A is the state transition matrix, which describes how the state changes from time t to time t+1. B is the input matrix, which represents the influence of the external input x′(t) on the current state. C is the observation matrix, which describes how to calculate the output y(t) from the current state h(t). D is the direct transmission matrix, which describes the relationship of how the input x′(t) directly affects the output y(t). h′(t) is the state equation, which represents how the state changes according to the input's influence on the state (through the matrix A). The output equation y(t) describes how the state is transformed into the output through the matrix C and how the input affects the output (through the matrix D).

[0118] The current SSM (the next state obtained after the previous state change) is a continuous time-invariant model and needs to be discretized for processing actual data. According to the differential equation and using the zero-order hold rule for sampling, the discretized result can be obtained as follows:

[0119]

[0120] Where is the discretized state transition matrix, exp represents the exponential operation of the matrix, and ΔA is the time step (ΔA represents the product of the time increment (sampling period) Δ and the state transition matrix A during the discretization process, usually written as: ΔA = A·Δ). is the discretized input matrix, I is the identity matrix, representing the discretization process of B. In the formula, Δ is the selectivity factor, which has the same shape as the input and mainly functions to determine the degree of state update in each time step.

[0121] Time series data is sequential. Therefore, the sequence model constructed in the present invention has the ability of context awareness, can filter out unimportant information, and focus on important information. The present invention uses a "selective information propagation mechanism" to control how information propagates in the sequence dimension. This mechanism (such as a gating mechanism or an attention mechanism) determines which information is passed to the next time step at each time step. The present invention discretizes the parameters Δ, A, B, and C in the SSM to transform the time-invariant model into a time-variable model. The discretized matrices and represent state transition and input mapping, and C is the observation matrix. Assuming D is 0, this means that the input only indirectly affects the output through the state:

[0122]

[0123] After performing residual processing on the result y output output by the SSM module, and then passing through the activation function and linear projection, we obtain

[0124] Step 7: Multiple lightweight student networks are trained using the latent representations of the teacher network (containing key information of the data, such as the dynamic features of the time series) and the reconstruction output. The encoding and decoding functions of the student network are G encoder and G decoder After the input X′ S-C the result obtained through the training of the student network is

[0125]

[0126] To make the reconstruction result of the student network close to the output of the teacher network, the following loss function is used for training:

[0127]

[0128] When the loss value reaches the optimum, the distillation learning will no longer be performed, and the prediction result will be output from the trained student network model and compared with the threshold set according to industrial experience to determine whether an abnormality has occurred. What the teacher network transfers to the student network are the latent representation and the reconstruction output, and multiple lightweight student networks learn based on this content, so as to be able to perform anomaly detection efficiently on edge devices.

[0129] Step 8: Calculate the accuracy, recall rate, and F1 score of each student network in the anomaly detection task to judge the precision of each student network; evaluate the computational resource consumption of each student network on the edge device, including memory usage and processing time. If the precision effect does not meet expectations, hyperparameter tuning can be performed and then knowledge distillation can be carried out; if the resource consumption is too large, model pruning can be performed.

[0130] Step 9: To achieve efficient anomaly detection and fault diagnosis, the present invention deploys multiple lightweight student models to the edge side. As Figure 1 shown, the specific steps are as follows:

[0131] First, the optimized student model is packaged into a deployable lightweight container, and each student model is deployed as an independent service module. Then, these containers are deployed on edge computing devices to utilize the advantages of low latency and high bandwidth of edge computing to achieve real-time data processing and analysis. In addition, a secure network connection and the data transmission protocol MQTT are established between the edge device and the digital twin platform to achieve two-way transmission and real-time update of model data and status information. The operating status and predicted data of the virtual model (a model constructed by digitizing and mapping the geometric structure, production process, and key parameters of the continuous casting equipment; it can simulate the operating status and predict data during the continuous casting process) are obtained in real time, so as to help the edge device make more accurate decisions based on the virtual state of the continuous casting equipment (the real-time simulation state of the continuous casting equipment in the virtual model).

[0132] Step Ten: After the student model preliminarily analyzes the edge device data, the edge side will upload some processing results or key data (such as anomaly detection results, device status changes, etc.) to the digital twin platform. This platform receives data from multiple edge devices and conducts in-depth analysis, trend prediction, and historical data backtracking. By integrating and analyzing this data, the digital twin platform can provide operators with a comprehensive monitoring view and display the global state of the device. The platform has a real-time monitoring function, which can display the operating health status of the device and perform fault diagnosis, predictive maintenance, and early warning based on the analysis results. If the system detects potential device failures or trend anomalies, the platform will automatically trigger an alarm to prompt the operator to take necessary measures. In addition, the platform also supports long-term trend analysis of the device, helping managers make data-based decisions, optimize device performance, reduce downtime, and improve production efficiency.

[0133] Through the two-way data interaction and collaborative work between the edge device and the digital twin platform, more accurate and efficient real-time monitoring and fault diagnosis can be achieved. In addition, the digital twin platform can rely on technical means such as real-time data collection, simulation analysis, and health assessment (such as using protocols such as MQTT and WebSocket to obtain data from edge devices, and then analyzing the operating status of the device through built-in machine learning algorithms or analysis models, and sending the results back to the edge device through a secure communication mechanism), so as to timely feedback the device health status, guide the edge device to make dynamic adjustments, and further optimize the maintenance strategy. Through this distributed deployment method, both the computing and storage resources of edge computing can be fully utilized, and the virtual simulation capabilities of digital twins can be used to improve the response speed, reliability, and maintenance efficiency of the entire system.

[0134] The present invention is directed to digital twins, combining an encoder with a Mamba network for research on a method for unsupervised anomaly detection of time-series data of continuous casting equipment. This method utilizes a distillation learning strategy to transfer the knowledge of a pre-trained model to multiple lightweight models, thereby improving the anomaly detection performance of continuous casting equipment while meeting the requirements of real-time performance and computing resources. The Mamba network is introduced and combined with a selective mechanism that allows the model to dynamically adjust its state according to the input, selectively propagating or ignoring information, further improving the detection accuracy and efficiency. In addition, the present invention also optimizes the hardware mechanism, enhancing the computing efficiency and memory usage of the model. By combining with a digital twin platform, the present invention can synchronize the states of physical devices and virtual models in real time, and further enhance the accuracy and real-time response ability of anomaly detection based on the data of virtual-real fusion. This innovation not only ensures the normal operation of continuous casting equipment but also improves production efficiency, providing strong support for the intelligent application of equipment.

[0135] The above-described embodiments merely represent one implementation manner of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can be made, and these all fall within the protection scope of the present invention. Therefore, the protection scope of the present invention patent shall be subject to the appended claims.

Claims

1. An unsupervised abnormal detection method for time-series data of an industrial continuous casting machine, characterized in that, It includes the following steps: Step 1: Monitor and analyze the key sensor data of the mold level and roll current of the continuous casting machine equipment; Step 2: Edge device configuration: Install and configure edge computing devices at the production line site. The edge computing devices are connected to the sensors of the continuous casting equipment through a wired network. The mold level and roll current sensors on the continuous casting equipment are connected to the edge computing devices using the industrial data stream transmission protocol MQTT. The sensors convert the collected real-time data into standardized digital signals and send them to the edge computing devices through the MQTT protocol; Step 3: Digital twin platform deployment: The platform is containerized and deployed using Docker and deployed on the server. Install Docker and Docker Compose on the server; According to the requirements of the digital twin platform, create a Dockerfile to define the environment, dependencies, and configurations of the platform; use Docker Compose to define the services required for multiple tasks; create a docker-compose.yml file; then use docker-compose to start all services and check if all services are running. The data collected and preprocessed by the edge device is transmitted to the digital twin platform for further data storage and real-time visualization analysis and display by the platform; Step 4: Data preprocessing: First, these data need to be cleaned and preprocessed, including filling missing values, removing noise, and standardization operations; Step 5: Correlation analysis between data: Use the Fast Greedy Equivalence Search method FGES to perform correlation analysis on the preprocessed time series data of the industrial continuous casting machine; Select these dimensional data with strong correlation relationships, namely causal relationships and dependencies, as the input of the model; Step 6: Teacher model training: Construct an encoder and a decoder structure, and introduce the Mamba network to construct the teacher model; use a large amount of unlabeled data to train it on the processed data set. The encoder of the teacher network is the part of the network responsible for processing the input data, and its task is to extract the key information in the input data. In the teacher model, the encoder receives the time series data from various sensors of the continuous casting machine as input and encodes this data to extract the deep features of this data; The decoder consists of multiple Mamba network modules; using the output of the encoder as input, it is calculated through multiple Mamba network modules, and finally the prediction result is generated; Step 7: Student model training: Based on the principle of distillation learning, transfer the knowledge of the teacher model to a lightweight student model. This process can be achieved by letting the student model imitate the output of the teacher model, or by letting the student model and the teacher model jointly optimize an objective function. During the training process, the student model only simulates or replicates the output of the teacher model according to normal sample data. Each student model is trained independently. When the student model cannot correctly predict the data outside the normal data distribution range, these data are considered abnormal; Step 8: Anomaly Detection: After the student model is trained, it can be used to predict new time-series data. By comparing the prediction results with the original data, anomalies can be detected. For minor anomalies, the teacher model, trained with a large amount of historical data, can identify complex anomaly features. The student model is trained by mimicking the output of the teacher model, i.e., soft labels, thus learning the teacher model's anomaly detection ability. The student model uses a loss function to minimize the difference from the teacher model and gradually masters the skills of identifying minor anomalies. In practical applications, the student model predicts new time-series data, calculates the difference by comparing with the prediction results of the teacher model, and determines whether there are anomalies. Finally, the detection results are output based on a preset threshold range to ensure accurate identification of anomalies. After completing the anomaly detection of the continuous casting equipment time-series data, the detection results need to be systematically evaluated to ensure that the model's performance meets the expected goals in various abnormal situations. If the evaluation results do not meet the requirements, the model parameters can be further adjusted or the distillation learning process can be optimized to improve the accuracy and efficiency of anomaly detection. Through continuous iteration and optimization, the optimal model configuration is finally determined to accurately identify various anomalies in the operation of the continuous casting equipment. Step 9: Edge Deployment of the Student Model: When the model performance reaches the expected standard, the lightweight student model optimized by distillation learning is deployed to the edge computing device of the continuous casting equipment. The student model deployed at the edge is responsible for the preliminary processing of the real-time data collected from the device sensors, including data cleaning, filtering, and formatting. The student model will use its compressed parameters and simplified structure to quickly detect anomalies in the device's operating state. Once an anomaly is detected, the edge device will respond immediately. Step 10: Alarm Monitoring on the Digital Twin Platform: After the student model preliminarily analyzes the edge device data, the edge side will upload some processing results or key data to the digital twin monitoring platform. This platform receives data from multiple edge devices and conducts in-depth analysis, trend prediction, and historical data backtracking. By integrating and analyzing this data, the digital twin platform can provide operators with a comprehensive monitoring view and display the global state of the device. The platform has a real-time monitoring function, which can display the operating health status of the device and perform fault diagnosis, predictive maintenance, and early warning based on the analysis results. If the digital twin platform detects potential faults or trend anomalies in the edge device, the platform will automatically trigger an alarm to prompt the operator to take necessary measures. In addition, the platform also supports long-term trend analysis of the device, helping managers make data-based decisions, optimize device performance, reduce downtime, and improve production efficiency.

2. The unsupervised industrial continuous casting machine time series data anomaly detection method according to claim 1, wherein, In Step 3, the role of the digital twin platform is to support decision-making and operations by obtaining data from the physical system or device in real time, establishing its virtual model, and conducting analysis and optimization. The digital twin platform includes the following functions: real-time data collection and integration; Data storage and management; digital model and simulation; Model analysis and decision support; Visualization to present device status, trends, and alarm information; Security and permission management to ensure the security of data and systems.

3. The unsupervised abnormal detection method for time series data of an industrial continuous casting machine according to claim 1, wherein In Step 4, interpolation methods, mean imputation methods, or machine learning-based imputation methods are used to fill in missing values; Filters or noise suppression techniques are used to remove high-frequency noise from the data; The data is standardized or normalized to eliminate the influence of dimensions and ensure that different feature data has a similar scale; In addition, outlier detection algorithms are used to identify and remove outlier data.

4. The unsupervised industrial continuous casting machine time series data anomaly detection method according to claim 1, characterized in that In Step 8, during the evaluation process, the accuracy, recall, and F1 score of each student network in the anomaly detection task are calculated to judge the precision of each student network; The computational resource consumption of each student network on the edge device is evaluated, including memory usage and processing time; If the precision effect does not meet expectations, hyperparameter tuning can be performed and then knowledge distillation can be carried out; If the resource consumption is too large, model pruning can be performed.

5. The unsupervised industrial continuous casting machine time series data anomaly detection method according to claim 1, characterized in that In Step 9, a secure network connection and the data transmission protocol MQTT are established between the edge device and the digital twin platform to achieve two-way transmission and real-time update of model data and status information; The running status and prediction data of the virtual model are obtained in real time.

6. The unsupervised industrial continuous casting machine time series data anomaly detection method according to claim 1, wherein In Step 1, the data collected by the mold part includes: the outlet flow rates of the inner arc and outer arc of the mold water wide face, the outlet flow rates of the right and left sides of the mold water narrow face, and the mold liquid level data; For the upper and lower guide roll currents, the upper roll current data in the bending section area and the upper and lower guide roll current data in the straightening section area are collected.

7. The unsupervised industrial continuous casting machine time series data anomaly detection method according to claim 1, characterized in that Step 4 specifically involves: processing the data of the continuous casting equipment input; first, it is necessary to ensure that the time information is correct, second, check for missing values, and finally, standardize the data of the remaining dimensions except the time column; the output data is where X t is the input time dimension, is the other dimensions of the input; after removing the time dimension from X, it can be expressed as After standardization, the processing result is X′ s ; Among them, m represents the m-th data.

8. The unsupervised industrial continuous casting machine time series data anomaly detection method according to claim 7, characterized in that Step 5 is specifically as follows: the preprocessed time series data X' S Find the causal relationship between data through the FGES algorithm; use an unsupervised method to analyze the correlation between each dimension by constructing a graph structure and relying on data relationships; specifically, find the time series X' S The steps for finding the causal relationship are as follows: Step 5.1, Initialize the graph structure: First, initialize an empty graph with no edges; Step 5.2, Forward search to add edges: Start forward search to add edges from the empty graph; For each possible edge (i,j), calculate the Bayesian Information Criterion BIC of the graph G after adding this edge: ΔBIC add = BIC(G+(i,j)) - BIC(G) Among them, G is the current graph, representing the causal relationship graph that contains the added edges; add the edge (i, j) that can make ΔBIC add positive and maximum to the graph until no edge addition can make ΔBIC add > 0; Step 5.3, Backward search to delete edges: Then start backward search to delete edges from the current graph G; For each existing edge (i,j), calculate the Bayesian Information Criterion BIC after deleting this edge: ΔBIC del = BIC(G-(i,j)) - BIC(G) Delete the edge (i, j) that can make ΔBIC del non - negative and maximum from the graph until no edge deletion can make ΔBIC del > 0; Step 5.4: Derive time series X from the graph structure ( ′ S-C) : First, extract the nodes with significant influence from the causal relationship graph; these key variables may be root nodes or other variables that have a direct impact on the target variable in the causal relationship graph; then, derive other variables affected by these variables according to the dependency relationships in the causal relationship; next, screen out the most significant variables from the derived variables according to the screening rules of the correlation threshold or causal strength; finally, combine the screened variables in chronological order to form the final time series X ( ′ S-C) , which reflects the dependency structure and causal connections among the variables in the causal relationship graph.

9. The unsupervised industrial continuous casting machine time series data anomaly detection method according to claim 8, wherein, Step 6 is specifically as follows: construct a Mamba autoencoder network for distillation learning for the time series data of the continuous casting equipment; the Mamba autoencoder network includes a teacher-student model, which is divided into the following steps: First, the time series data X S ′ -C is input into the teacher network. After the features are extracted by the encoder, they are passed to the decoder for data reconstruction; the goal of the teacher network is to be trained by optimizing the reconstruction error and generate prediction results; then, the student model learns from the teacher network through knowledge distillation, adopting the same encoder and decoder structures, but its goal is to minimize the difference from the prediction results of the teacher model as much as possible; each student model makes predictions independently and outputs the anomaly detection results; by comparing the prediction results of the student models, the digital twin platform can detect abnormal behaviors and give feedback; the specific processing process is as follows: The input time series X′ S-C is mapped into a low-dimensional latent representation space Z by the encoder, and the result Feature of the encoder is obtained as the output: Feature=f encoder (X′ S-C ) where f encoder represents the encoder function; then Feature is reconstructed back into the original data form through the Mamba network acting as the decoder: where f decoder is the function of the Mamba network decoder, and X′ R(S-C) is the reconstructed output of the teacher network; The Mamba network extracts features from the input time series data through convolutional layers and introduces non-linear transformations through activation functions to enhance the network's expressive power; Subsequently, the hybrid state encoder encodes the extracted features into hidden states to capture the context relationship of the time series and models the dynamic changes in the time series through the State Space Model SSM; Finally, the hybrid state decoder decodes these hidden states to restore the original data structure; For the output of the encoder, it can be regarded as a triple with dimensions (Batch, V, D), where Batch is the batch size, V is the number of variables, and D is the hidden dimension; First, the result Feature of the encoder is linearly projected to expand the input time series to a higher dimension ED: x,z = Linear(Feature) where x and z represent the feature representations after linear projection; Linear is the linear projection operation that maps Feature to dimension ED; where the dimensions of x and z are (Batch, V, ED); then a convolutional layer and an activation function are used to process the linear projection result: x′ = σ(Conv(x)) conv(x) is the convolution operation on x, σ is the activation function, and x′ is the new feature representation obtained after convolution and activation, with the dimension still being (Batch, V, ED); The state space model is used to describe these state representations and predict the next state based on the input; at time t, the input sequence x′(t) is mapped to the latent state representation h(t), and the predicted output sequence y(t) is derived; then the future state can be predicted by the following two equations: h′(t) = Ah(t) + Bx′(t) y(t) = Ch(t) + Dx′(t) where A is the state transition matrix, which describes the way the state changes from time t to time t+1; B is the input matrix, indicating the influence of the external input x′(t) on the current state; C is the observation matrix, which describes how to calculate the output y(t) from the current state h(t); D is the direct transmission matrix, which describes the relationship between how the input x′(t) directly affects the output y(t); h′(t) is the state equation, indicating how the state changes according to the influence of the input on the state; the output equation y(t) describes how the state is transformed into the output through the matrix C and how the input affects the output; The current SSM is a continuous time-invariant model and needs to be discretized for processing actual data; according to the differential equation, sampling using the zero-order hold rule can obtain the discretized result: Among them is the state transition matrix after discretization, exp represents the exponential operation of the matrix, and ΔA is the time step; is the input matrix after discretization, I is the identity matrix, representing the discretization process of B; in the formula, Δ is the selectivity factor, and its shape is the same as the input; Use the "selective information propagation mechanism" to control how information propagates in the sequence dimension; discretize the parameters Δ, A, B, and C in the SSM to transform the time-invariant model into a time-variable model; the discretized matrices and represent state transition and input mapping, and C is the observation matrix; assume D to be 0, which means that the input only indirectly affects the output through the state: After the result y output by the SSM module output is subjected to residual processing and then passes through an activation function and a linear projection, we obtain 10. The unsupervised industrial continuous casting machine time series data anomaly detection method according to claim 9, characterized in that Step 7 specifically is: Multiple lightweight student networks are trained using the latent representation and reconstruction output of the teacher network; the encoding and decoding functions of the student network are G and G encoder and G decoder After the input X S ′ -C the result obtained through the training of the student network is To make the reconstruction result of the student network close to the output of the teacher network, the following loss function is used for training: When the loss value reaches the optimum, the distillation learning will no longer be performed, and the prediction result will be output from the trained student network model and compared with the threshold set based on industrial experience to determine whether an anomaly has occurred; What the teacher network transmits to the student network are the latent representation and the reconstruction output, and multiple lightweight student networks learn based on this content, so as to be able to perform anomaly detection on edge devices.

Citation Information

Cited By

  • Acoustic test exception processing method, system and device and storage medium

    CN120611333A

  • Hydrological flow monitoring system based on edge computing gateway equipment and software product

    CN121089688A

  • Classroom event analysis method and analysis model optimization method and system

    CN121415298A

  • Photovoltaic power generation power prediction method and device for large model distillation, and medium

    CN121479171A