Fault prediction and detection method and device based on intelligent model, equipment and medium

Through the fault prediction method based on intelligent model, the sensor data is pre-trained and fine-tuned using the self-attention mechanism and position-aware attention mask, which solves the timing accuracy and real-time response problems in industrial equipment failure prediction, and achieves efficient fault monitoring and alarm.

CN120493070APending Publication Date: 2025-08-15SUN YAT SEN UNIV

Patent Information

Application Number
CN202510572669.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

In the fault prediction of industrial equipment, the prior art problems of insufficient timing prediction accuracy, weak real-time response capabilities and poor model adaptability, making it difficult to effectively deal with the signs of failure in complex and dynamic operating environments.

Method used

The fault prediction method based on intelligent model is adopted, and the sensor historical data is obtained for preprocessing, time sequence data is generated, and pre-trained using the self-attention mechanism, combined with the position-perceived attention mask for fine-tuning, generate a fine-tuned model, obtain the current data for prediction, analyze the prediction deviation and trigger an alarm signal.

Benefits of technology

It improves the accuracy and real-timeness of fault prediction, enhances the adaptability of the model under different operating conditions, realizes intelligent monitoring of the operating status of the equipment, and improves the reliability and real-timeness of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120493070A_ABST
    Figure CN120493070A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence and the technical field of equipment fault prediction and detection, and discloses a fault prediction and detection method based on an intelligent model, and the method comprises the steps: obtaining and preprocessing the historical data of a sensor, and generating time series data; pre-training the time sequence data to generate a pre-training model; performing fine tuning on the pre-training model based on the attention mask of position perception to generate a fine-tuned model; acquiring current data of the sensor, and generating prediction data; and analyzing the deviation between the predicted data and the actual data, and triggering an alarm signal when the deviation exceeds a preset threshold value. According to the method, the time series data features are learned through pre-training, multi-segment prediction is performed in combination with an attention mechanism, and the accuracy of fault prediction is improved; the fine-tuned model adapts to different working conditions, and the generalization ability of the model is enhanced; alarm is triggered through deviation analysis, intelligent monitoring of the equipment state is achieved, and real-time performance and reliability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of artificial intelligence technology and equipment fault prediction and detection technology, and in particular to a fault prediction and detection method, device, equipment and storage medium based on an intelligent model. Background Art

[0002] With the rapid development of industrial automation and intelligent manufacturing, the operating environment of industrial equipment is becoming increasingly complex. This is especially true in the field of high-end equipment manufacturing, such as CNC machine tools and stamping equipment, where critical links require long-term operation under high-precision and high-intensity working conditions. Therefore, real-time monitoring of equipment operating status and fault detection have become important technical directions for ensuring equipment safety and improving production efficiency. However, existing technologies still have many shortcomings in data fusion, fault pattern recognition, and prediction, which are mainly reflected in the following aspects:

[0003] In modern industrial environments, equipment is often equipped with multiple types of sensors, such as vibration, temperature, pressure, and current sensors, to comprehensively monitor its operating status. However, existing technologies typically use single-channel or simplified feature extraction methods, relying on a single data source (such as vibration or temperature) for analysis, making it difficult to fully capture the complexity of equipment operating status.

[0004] In addition, different types of sensors have different data sampling frequencies and formats. Traditional methods face technical challenges in data synchronization, alignment, and fusion, resulting in the inability to fully explore the correlation characteristics between various types of sensors, affecting the accuracy of fault prediction.

[0005] Traditional fault detection methods mostly use judgment rules based on threshold settings or simple machine learning models (such as support vector machines and decision trees). Although these methods can identify certain predefined fault modes, they find it difficult to maintain stable detection performance when faced with variable working conditions (such as equipment load changes, temperature fluctuations, component wear, etc.).

[0006] In addition, although some fault detection methods based on deep learning have certain learning capabilities, they are prone to model overfitting due to the limited coverage of training data, resulting in insufficient generalization ability and decreased detection accuracy under unknown fault modes.

[0007] During the operation of industrial equipment, faults often exhibit time series characteristics, requiring prediction based on historical data. Common time series prediction methods include autoregressive models (AR) and long short-term memory networks (LSTM). While these methods can capture the temporal dependencies of data, they suffer from the following issues:

[0008] Accurate short-term predictions, but large errors in long-term predictions: Traditional methods have good prediction effects within a short time frame, but cannot accurately predict failure trends at more distant time points, resulting in insufficient early warning capabilities.

[0009] Error accumulation: Recursive time series forecasting methods will cause errors to gradually accumulate, affecting the stability of forecasts over long time windows.

[0010] High computational overhead: Although some deep learning methods can improve prediction accuracy, they have high computational complexity, making them difficult to deploy in industrial sites and making it difficult to achieve real-time prediction and detection.

[0011] Industrial equipment operates in complex environments, and the collected data is easily interfered with by external noise. For example, mechanical vibration and impact interference cause large fluctuations in vibration data; changes in ambient temperature and humidity affect the measurement accuracy of temperature sensors; and fluctuations in equipment current cause abnormal deviations in current signals.

[0012] Most existing technologies use simple filtering methods (such as mean filtering and low-pass filtering) to reduce noise interference, but these methods are often difficult to be effective when faced with non-stationary noise and sudden change signals, which may cause important fault characteristics to be ignored or normal operating conditions to be mistakenly detected as faults. Summary of the Invention

[0013] The main purpose of the present invention is to provide a fault prediction and detection method, device, equipment and storage medium based on an intelligent model, aiming to solve the technical problem that the existing technology has obvious deficiencies in timing prediction accuracy, real-time response capability and model adaptability, and is difficult to effectively respond to the constantly changing fault signs of industrial equipment in a complex and dynamic operating environment.

[0014] To achieve the above objectives, the present invention provides a fault prediction and detection method based on an intelligent model, comprising:

[0015] Acquire sensor historical data, preprocess the sensor historical data, and generate time series data;

[0016] Dividing the time series data into historical data segments, and performing pre-training based on the historical data segments through a self-attention mechanism to generate a pre-training model;

[0017] splicing the historical data segments and the predicted blank segments into an input sequence, and fine-tuning the pre-trained model in combination with the position-aware attention mask to generate a fine-tuned model;

[0018] Acquire current sensor data, input the current sensor data into the fine-tuned model, and generate prediction data within a prediction time window;

[0019] Acquiring actual sensor data within the prediction time window, and analyzing a prediction deviation between the predicted data and the actual sensor data within the prediction time window;

[0020] When the prediction deviation exceeds a preset threshold, an alarm signal is triggered.

[0021] Furthermore, to achieve the above objectives, the present invention provides a fault prediction and detection device based on an intelligent model, comprising:

[0022] A data preprocessing module is used to obtain sensor historical data, preprocess the sensor historical data, and generate time series data;

[0023] A pre-training model construction module is used to divide the time series data into historical data segments, and perform pre-training based on the historical data segments through a self-attention mechanism to generate a pre-training model;

[0024] A model fine-tuning module is used to concatenate the historical data segments and the predicted blank segments into an input sequence, and fine-tune the pre-trained model in combination with the position-aware attention mask to generate a fine-tuned model;

[0025] A real-time prediction module is used to obtain current sensor data, input the current sensor data into the fine-tuned model, and generate prediction data within a prediction time window;

[0026] A prediction deviation analysis module, configured to obtain actual sensor data within the prediction time window and analyze a prediction deviation between the prediction data and the actual sensor data within the prediction time window;

[0027] The abnormality alarm module is used to trigger an alarm signal when the prediction deviation exceeds a preset threshold.

[0028] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a computer device, which includes a memory, a processor, and an intelligent model-based fault prediction and detection program stored in the memory and runnable on the processor. When the intelligent model-based fault prediction and detection program is executed by the processor, the steps of the intelligent model-based fault prediction and detection method as described above are implemented.

[0029] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, on which an intelligent model-based fault prediction and detection program is stored. When the intelligent model-based fault prediction and detection program is executed by a processor, the steps of the intelligent model-based fault prediction and detection method as described above are implemented.

[0030] Beneficial effects: The present invention relates to the fields of artificial intelligence technology and equipment fault prediction and detection technology, and discloses a fault prediction and detection method based on an intelligent model, comprising: obtaining sensor historical data and preprocessing it to generate time series data; dividing the time series data into historical data segments, pre-training the historical data segments based on a self-attention mechanism to generate a pre-trained model; splicing the historical data segments and prediction blank segments into an input sequence, fine-tuning the pre-trained model in combination with a position-aware attention mask to generate a fine-tuned model; obtaining sensor current data, and inputting the sensor current data into the fine-tuned model to generate predicted data within a prediction time window; obtaining sensor actual data within the prediction time window, and analyzing the prediction deviation between the predicted data and the sensor actual data; and triggering an alarm signal when the prediction deviation exceeds a preset threshold. The present invention improves the accuracy of fault prediction by learning the long-term dependency of time series data through a pre-trained model and performing multi-segment parallel prediction in combination with a position-aware attention mask; performing real-time prediction through the fine-tuned model to enhance the adaptability of the model under different working conditions; triggering an alarm by analyzing the deviation between the predicted data and the actual data, thereby realizing intelligent monitoring of the equipment operating status and improving the real-time performance and reliability of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] The present invention will be further described below with reference to the accompanying drawings and embodiments, in which:

[0032] Figure 1 Schematic diagram of an application environment of a fault prediction and detection method based on an intelligent model in an embodiment of the present invention;

[0033] Figure 2 This is a flow chart of an embodiment of a fault prediction and detection method based on an intelligent model of the present invention;

[0034] Figure 3 Schematic diagram of functional modules of a preferred embodiment of a fault prediction and detection device based on an intelligent model of the present invention;

[0035] Figure 4 A schematic diagram of the structure of a computer device according to an embodiment of the present invention;

[0036] Figure 5 FIG. 2 is another structural diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0037] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0038] The fault prediction and detection method based on intelligent model provided by the embodiment of the present invention can be applied in Figure 1in an application environment, wherein the user end communicates with the server end through a network. The server end can obtain the historical data of the sensor through the user end and perform preprocessing to generate time series data; divide the time series data into historical data segments, pre-train the historical data segments based on the self-attention mechanism to generate a pre-trained model; splice the historical data segments and the prediction blank segments into an input sequence, fine-tune the pre-trained model in combination with the position-aware attention mask, and generate a fine-tuned model; obtain the current data of the sensor, and input the current data of the sensor into the fine-tuned model to generate the predicted data within the prediction time window; obtain the actual data of the sensor within the prediction time window, and analyze the prediction deviation between the predicted data and the actual data of the sensor; when the prediction deviation exceeds a preset threshold, trigger an alarm signal. The present invention learns the long-term dependency of time series data through a pre-trained model, and performs multi-segment parallel prediction in combination with the position-aware attention mask, thereby improving the accuracy of fault prediction; performs real-time prediction through the fine-tuned model, thereby enhancing the adaptability of the model under different working conditions; triggers an alarm through the deviation analysis between the predicted data and the actual data, thereby realizing intelligent monitoring of the equipment operation status and improving the real-time performance and reliability of the system. The user end may be, but is not limited to, various personal computers, laptops, smartphones, tablet computers, and portable wearable devices. The server end may be implemented as an independent server or a server cluster consisting of multiple servers. The present invention is described in detail below using specific embodiments.

[0039] See also Figure 2 , Figure 2 This is a flow chart of an embodiment of the intelligent model-based fault prediction and detection method provided by the present invention. It should be noted that although the flow chart shows a logical order, in some cases, the steps shown or described may be performed in a different order than here.

[0040] like Figure 2 As shown, the fault prediction and detection method based on the intelligent model proposed in the present invention includes the following steps:

[0041] S10, acquiring sensor historical data, preprocessing the sensor historical data, and generating time series data;

[0042] In this embodiment, acquiring historical data is fundamental to the entire method for predicting and detecting industrial equipment faults. Sensors are installed on key equipment components, such as bearings, crankshafts, motors, and hydraulic systems. The operating status of these components directly affects the overall performance of the equipment. Collected data is typically transmitted to a storage device or the cloud via wired or wireless communication, forming a historical data record.

[0043] Different data collection methods are suitable for different industrial environments. In high-precision manufacturing equipment (such as CNC machine tools), sensors typically use Industrial Ethernet for high-speed data transmission, ensuring low latency and high stability. For remote equipment monitoring, wireless sensor networks (WSNs) or low-power wide area network (LPWAN) technologies such as LoRa can enable remote data collection. During the data collection process, the Persistent Time Protocol (PTP) can be used to align the time of multi-sensor data to ensure data timing consistency.

[0044] Historical sensor data may have measurement errors, noise interference, data loss and other problems. In order to ensure the stability and availability of the data, preprocessing steps such as denoising, outlier detection, interpolation and completion, and data normalization are required.

[0045] Industrial sensor data is often subject to environmental interference, such as mechanical vibration and electromagnetic interference, leading to abnormal data fluctuations. Kalman filtering or wavelet transforms can be used to smooth this problem. Kalman filtering is suitable for real-time data correction in dynamic environments, while wavelet transforms can decompose signals and extract different frequency components to reduce the impact of high-frequency noise. In the data processing of pressure sensors in stamping equipment, using wavelet transforms to filter out periodic interference signals can improve the ability to detect abnormal impact events. In punch press vibration monitoring, Kalman filtering can be used to optimize vibration signals, reduce misjudgments of environmental vibrations, and improve the accuracy of fault detection.

[0046] Sensor data may contain outliers due to equipment malfunction, sensor failure, or network packet loss. To address these issues, outlier detection can be performed using the 3σ principle based on statistical distribution, isolation forests, or density-based clustering (DBSCAN). In current monitoring of stamping equipment, the isolation forest algorithm can identify abnormal current changes caused by motor aging or overload. In mold wear monitoring, the DBSCAN algorithm can detect abnormal patterns in continuous workpiece pressure data, identifying mold damage risks in advance.

[0047] Data loss can occur due to unstable network communications, device power outages, and other factors, necessitating the use of interpolation algorithms to fill in missing data. Linear and spline interpolation methods are suitable for short-term data loss, while LSTM-based interpolation methods are suitable for filling in missing data over longer periods. In temperature sensor data processing on stamping presses, spline interpolation can be used to fill in gaps caused by short-term data loss, improving the stability of temperature anomaly trend analysis. In hydraulic system monitoring, LSTM interpolation can predict pressure changes during missing time periods, ensuring the integrity of time series data.

[0048] Different sensor data have different dimensions (such as temperature in degrees Celsius and pressure in Pascals), so standardization is required to ensure that different types of data can be effectively integrated in the same model. Common methods include Z-score normalization (mean normalization) and Min-Max normalization. Z-score normalization is applicable to normally distributed data, while Min-Max normalization is applicable to data in a fixed range. In real-time fault prediction of stamping equipment, the use of Z-score normalization can reduce the impact of different sensor data on model training and improve data consistency. In punch vibration monitoring, Min-Max normalization can convert vibration acceleration signals into a standard range to improve the adaptability of the prediction model.

[0049] After preprocessing, data needs to be organized into a standard time series format to support predictive model learning. Generating time series data involves key steps such as data format conversion, time window division, and feature extraction.

[0050] Different sensor data may be stored in different formats, such as CSV, JSON, and binary HDF5, and needs to be converted to a unified data format to improve reading efficiency and compatibility. Industrial control systems often use the Parquet format to store time series data to support efficient queries. In the stamping machine production process, using Parquet to store pressure, temperature, and vibration signals can improve data analysis efficiency and reduce storage space usage. In remote equipment monitoring, storing data based on time series databases (such as InfluxDB) can increase the speed of parallel queries on multiple devices.

[0051] Time series prediction requires dividing data into time windows of fixed length to ensure that the model can learn short-term and long-term time dependencies. Window division methods include fixed windows, sliding windows, and adaptive windows. Fixed windows are suitable for stable working conditions, while adaptive windows are suitable for dynamically changing operating environments. In the stroke analysis of stamping equipment, the use of fixed windows (such as 1 second) can extract the periodic change characteristics of the equipment and improve the model's ability to identify repetitive faults. On automated production lines, the use of adaptive windows can dynamically adjust the time slice length according to changes in pressure and vibration signals to improve prediction accuracy.

[0052] The generated time series data can be further used to extract features, such as time series statistical features (mean, variance), frequency domain features (FFT transform), and embedded features (autoencoder), to improve the learning ability of the prediction model. In vibration monitoring of stamping machines, the use of FFT transform to extract frequency features can effectively identify whether the equipment is in an abnormal operating state. In hydraulic system fault prediction, the use of autoencoders to extract low-dimensional representations of the data helps detect potential abnormal patterns.

[0053] Example: During the operation of stamping equipment, die wear is one of the key factors leading to reduced production quality and equipment failure. To predict the lifespan of a die, it is necessary to collect real-time sensor data such as punching force, vibration, temperature, and stroke position. This data is then used to construct time series data for analyzing die status changes.

[0054] Kalman filtering is used to denoise punch pressure sensor data, removing transient outliers caused by impact force fluctuations and improving signal stability. An isolation forest algorithm is used to identify abnormal patterns in vibration sensor data and detect potential mold damage trends. A sliding window is used to segment time series data, analyze mold wear over different time periods, and predict remaining service life. FFT feature extraction is used to analyze spectral changes during the stamping cycle to determine whether the mold has experienced wear failure.

[0055] Multiple data preprocessing methods improve the quality of stamping equipment sensor data, reduce noise interference, optimize time synchronization, and enhance the reliability of time series data. This reduces data errors, improves the accuracy of the prediction model, and adapts to different operating conditions, improving the accuracy and real-time performance of fault prediction.

[0056] S20, dividing the time series data into historical data segments, and performing pre-training based on the historical data segments through a self-attention mechanism to generate a pre-training model;

[0057] In this example, time series data contains a large amount of historical operational information. To facilitate the model's learning of time series patterns, the data needs to be segmented into fixed-length input data units. Each data unit contains a continuous period of time steps, forming historical data segments. This segmentation allows the model to learn data features within different time windows and extract temporal dependencies.

[0058] Fixed-length window partitioning: The entire time series data is divided according to the preset time step, for example, every N time steps is a data segment, ensuring that the data format of each segment is consistent.

[0059] Sliding window partitioning: Based on a fixed time step, an overlapping window method is used for sliding partitioning, with a fixed step size each time to increase data coverage, ensure partial overlap between adjacent time segments, and improve the learning ability of the model.

[0060] Adaptive window partitioning: Dynamically adjust the window size based on the changing trend of the data. For example, a longer time window is used when the device is running smoothly, and a shorter window is used when the device status suddenly changes, thereby increasing the model's sensitivity to sudden failures.

[0061] In stamping equipment fault prediction, fixed window methods are suitable for detecting periodic faults, such as analyzing pressure variation patterns between strokes. Sliding window methods are suitable for capturing short-term trends in equipment state changes, such as detecting local anomalies in vibration signals. Adaptive window methods are suitable for complex industrial scenarios, such as monitoring pressure variations in hydraulic systems. They can dynamically adjust the window length to adapt to varying operating conditions.

[0062] Self-attention is a method that learns the relationships between different time steps in a sequence. By calculating the correlation between different time steps, the model can focus on key historical data and extract temporal dependencies. During the pre-training phase, the model learns the temporal patterns of a large amount of historical data, providing a foundation for subsequent fine-tuning and prediction.

[0063] Input feature mapping: Perform high-dimensional embedding mapping on historical data segments to convert the original time series data into feature representations suitable for model input.

[0064] Autoregressive encoding: An autoregressive model (such as Transformer) is used to encode input features and calculate the dependencies between time steps through a multi-layer self-attention mechanism.

[0065] Masking mechanism: During the self-attention calculation process, data of future time steps are masked so that the model can only focus on data of the current time step and before, achieving causal learning.

[0066] Feature Normalization: Use layer normalization to standardize the calculation results, improve model stability, and reduce the gradient disappearance problem.

[0067] In vibration analysis of stamping equipment, the self-attention mechanism can be used to capture abnormal patterns in vibration waveforms, improving the accuracy of fault detection. In pressure signal prediction during the stamping process, a masking mechanism is used to block data from future time steps, allowing the model to make predictions based solely on past pressure trends, thus preventing information leakage. In high-precision equipment fault prediction, the use of multi-layer self-attention structures can learn deeper temporal dependencies, improving the generalization ability of the prediction model.

[0068] Pretrained models are trained on large amounts of unlabeled data, enabling them to learn common timing patterns and provide good initial weights for subsequent tasks. Pretraining allows the model to grasp the characteristics of devices under different operating conditions, improving convergence speed and predictive capabilities during subsequent fine-tuning.

[0069] Self-supervised learning: Use self-supervised learning methods (such as mask prediction and contrastive learning) to train the model on large-scale historical data so that it can learn the normal operating patterns of the device.

[0070] Causal modeling: Through causal reasoning methods, the model can predict future time steps based on past historical data, avoiding the error accumulation problem in traditional time series methods.

[0071] Multi-task training: During the pre-training phase, multiple tasks are trained simultaneously, such as fault mode classification and trend prediction, to improve the generalization ability of the model.

[0072] In stamping equipment fault detection, a pre-trained model is trained using a mask prediction method, enabling it to automatically fill in missing data and predict future trends in stamping force. In complex industrial environments, a pre-trained model is trained using a contrastive learning method, enabling it to distinguish between normal and abnormal states, improving the model's ability to identify unknown faults. By combining industrial knowledge with a data-driven model and using multi-task learning, the pre-trained model can both predict equipment trends and identify equipment anomalies, enhancing its practical application value.

[0073] By scientifically partitioning time series data, the model can efficiently learn the temporal dependencies of historical data. Pre-training with a self-attention mechanism enables the model to capture key features, improving the accuracy and reliability of fault prediction. Causal modeling and self-supervised learning enhance the model's ability to identify unknown faults, improving fault prediction for industrial equipment. This allows for more accurate capture of changes in equipment operating status, reducing false positives and missed alerts, and improving the safety and stability of industrial production.

[0074] S30, concatenating the historical data segments and the predicted blank segments into an input sequence, and fine-tuning the pre-trained model in combination with a position-aware attention mask to generate a fine-tuned model;

[0075] In this embodiment, to improve the accuracy of fault prediction, the model needs to consider both past device status information (historical data segments) and the future time window to be predicted (prediction blank segments). The historical data segments contain the actual operation records of the equipment, while the prediction blank segments are the unknown parts used for model reasoning. By combining these two into a complete input sequence, the model can infer future device status based on historical data.

[0076] Fixed-length splicing: Set the length of the prediction window so that historical data segments and prediction blank segments are spliced according to preset rules. For example, each input sequence consists of N historical time steps + M prediction time steps, ensuring that the model can predict future states based on existing data.

[0077] Sliding window splicing: Using a sliding window strategy, each time moving a fixed step size, updating historical data segments, and adding new prediction blank segments to enhance the model's time series learning ability.

[0078] Feature dimension alignment: Since the predicted blank segments initially have no specific data values, they need to be initialized using zero padding, mean padding, or padding based on historical statistical values to keep the feature dimensions of the input sequence consistent.

[0079] In stamping equipment fault prediction, fixed-length splicing can be used for pressure signal prediction, ensuring that the model can learn trend patterns from the complete historical time window. In hydraulic system fault monitoring, sliding window splicing can dynamically update hydraulic pressure status, improving responsiveness to sudden faults. In motor load prediction, feature dimension alignment ensures consistency of data from different sensors and avoids model calculation errors caused by mismatched input formats.

[0080] In the self-attention mechanism, each time step can pay attention to information from other time steps in the sequence. To ensure that the model does not access data from future time steps during prediction, a position-aware attention mask is used. This masking mechanism allows time steps in historical data segments to pay attention to each other while shielding prediction gap segments from accessing future time steps, thereby preserving causal relationships in predictive reasoning.

[0081] Autoregressive masking: During attention calculation, data from future time steps is masked to ensure that the model's reasoning relies only on past information.

[0082] Positional encoding enhancement: Time step information is introduced through positional encoding, enabling the model to understand the temporal structure of the data and avoid the loss of time information.

[0083] Multi-layer masking mechanism: During the model's multi-layer self-attention calculation process, position-aware masks are applied layer by layer to ensure that deep networks maintain causality.

[0084] In press pressure monitoring, autoregressive masking prevents the model from accessing future pressure values during prediction, ensuring prediction accuracy. In press vibration signal modeling, position encoding enhancement enables the model to distinguish vibration patterns at different time steps, improving time series modeling capabilities. In hydraulic system trend prediction, a multi-layer masking mechanism ensures the temporal causality of hydraulic pressure, avoiding non-physical prediction errors.

[0085] Fine-tuning is the process of further optimizing the initial parameters of a pre-trained model on a target task dataset to make the model more suitable for specific equipment fault prediction tasks. The fine-tuning process can be divided into two methods: global fine-tuning and local fine-tuning.

[0086] Global fine-tuning: Adjust the parameters of the entire pre-trained model so that the model can be trained end-to-end on new device data. This is suitable for situations where the amount of training data is large.

[0087] Local fine-tuning: Adjusting only some layers of the model, such as the attention weights of the last few layers or the weights of the decoding layer, adapts the model to the new task while retaining the common features learned during pre-training. This is suitable for situations where the amount of data is small but fast adaptation is required.

[0088] Task adaptation fine-tuning: By introducing task-specific loss functions, such as prediction error minimization loss and fault classification loss, the model can focus on the target task and improve prediction performance.

[0089] In vibration fault prediction for stamping equipment, local fine-tuning allows the model to adapt to different types of equipment vibration patterns without having to start training from scratch. In stamping die life prediction, task-adaptive fine-tuning optimizes the model's ability to predict die wear and improves the accuracy of maintenance plans. In hydraulic system anomaly detection, global fine-tuning allows the model to adapt to different hydraulic operating environments, improving generalization capabilities.

[0090] After fine-tuning, the performance of the model on specific tasks is optimized, forming a fine-tuned model that can be directly used for real-time prediction and alarm detection.

[0091] Dynamic parameter update: After the model is deployed, incremental learning can be performed based on newly collected data to continuously optimize model parameters and improve long-term prediction stability.

[0092] Edge computing deployment: The fine-tuned model can be deployed on edge computing nodes to achieve real-time prediction and local alarm without relying on cloud computing, reducing prediction latency.

[0093] Model compression optimization: After model training is completed, methods such as knowledge distillation or quantization can be used to optimize computing efficiency and reduce computing resource consumption.

[0094] On stamping equipment production lines, dynamic parameter updates can continuously optimize fault predictions based on real-time data, improving the model's long-term adaptability. In remote monitoring systems, edge computing deployment enables low-latency fault prediction, preventing network transmission delays from impacting production decisions. In resource-constrained embedded systems, model compression optimization can reduce computing costs, minimizing hardware requirements while maintaining prediction accuracy.

[0095] By splicing historical data fragments with prediction blank fragments, the model ensures that predictions are based on complete time series data, avoiding information loss that affects inference performance. Using location-aware attention masks ensures that prediction tasks adhere to causal relationships, improving prediction credibility. Fine-tuning strategies optimize model parameters to enhance the model's adaptability to specific tasks. Combining dynamic updates, edge computing, and model optimization further improves the model's long-term stability and industrial application value. This enables more accurate predictions of equipment status changes, reduces false alarms and missed faults, and improves the safety and efficiency of stamping production lines.

[0096] S40, obtaining current sensor data, inputting the current sensor data into the fine-tuned model, and generating prediction data within a prediction time window;

[0097] In this embodiment, to achieve real-time fault prediction, the system needs to continuously collect sensor data on the device's operating status. Current data refers to the sensor observations of the device at the current time point or within the most recent sampling period, typically including information such as pressure, vibration, temperature, and current.

[0098] Data acquisition protocols: Use industrial communication protocols such as Modbus, OPC UA, MQTT, Industrial Ethernet (EtherCAT, PROFINET), etc. to ensure efficient transmission of different sensor data.

[0099] Real-time data update: Set a fixed sampling interval, such as milliseconds (high-precision equipment) or seconds (standard industrial equipment), to ensure the timeliness of data.

[0100] Data storage and buffering: Use edge computing cache or time series databases (such as InfluxDB and TimescaleDB) to store current data for subsequent modeling and prediction.

[0101] In pressure monitoring of stamping equipment, industrial Ethernet is used to acquire stamping force signals in real time, ensuring the timeliness of prediction model inputs. In hydraulic system monitoring, OPC UA is used to collect hydraulic pressure data and store it in a local cache to reduce network latency. In motor load monitoring, MQTT is used to collect motor current signals and pre-process them through edge computing nodes to improve data stability.

[0102] The current data needs to undergo certain preprocessing and feature conversion to ensure that it can be effectively used by the model. The model needs to map the current sensor data into the same input format as the training data and input it into the fine-tuned model for prediction.

[0103] Data standardization: Perform Z-score standardization, minimum-maximum normalization and other methods on the current data to make the data distribution consistent with the training data.

[0104] Feature mapping: Through high-dimensional embedding mapping, the current data is converted into a feature representation acceptable to the model, such as vectorized encoding.

[0105] Batch optimization: If the model supports batch prediction, data from multiple time steps can be packaged into mini-batch inputs to improve computational efficiency.

[0106] In die monitoring for stamping equipment, Z-score standardization can eliminate numerical scale differences between different devices, making model predictions more stable. In hydraulic system predictions, high-dimensional embedding mapping ensures consistent data formats across different sensors, improving multimodal data fusion capabilities. In automated production line monitoring, batch processing optimization can increase the efficiency of parallel predictions across multiple devices and reduce computation time.

[0107] The prediction time window refers to the time range predicted by the model, such as the device state in the next 1, 5, or 10 seconds. The model infers the future device state based on current data and generates predicted data.

[0108] Multi-step prediction: Use methods such as recursive prediction and parallel prediction to ensure that the model can predict future data for multiple time steps at a time.

[0109] Uncertainty estimation: Use Bayesian neural networks and confidence interval methods to calculate the confidence level of predicted data and improve prediction reliability.

[0110] Dynamically adjust the forecast window: Dynamically adjust the forecast time range based on historical data to ensure high accuracy of short-term forecasts while maintaining the availability of long-term forecasts.

[0111] In vibration prediction for stamping equipment, parallel prediction is used to generate data for multiple future time steps, improving the real-time nature of fault warnings. In motor load prediction, uncertainty estimation is incorporated to provide confidence intervals for prediction results, reducing false alarm rates. In hydraulic system anomaly monitoring, the prediction window is dynamically adjusted to ensure the accuracy of short-term trend predictions while optimizing the stability of long-term trend predictions.

[0112] By collecting current sensor data in real time, the prediction model input data is kept up-to-date, improving prediction accuracy. Data standardization and high-dimensional mapping ensure that the input data distribution is consistent with the training data, enhancing model stability. Combining multi-step prediction, uncertainty estimation, and dynamic window adjustment improves the accuracy and reliability of prediction results. This system is widely applicable to predicting key components of stamping equipment, such as die life, motor load, and hydraulic system status. Through real-time data collection, input standardization, and prediction optimization, it improves fault prediction accuracy, reduces unplanned equipment downtime, and provides data support for intelligent manufacturing.

[0113] S50, acquiring actual sensor data within the prediction time window, and analyzing a prediction deviation between the predicted data and the actual sensor data within the prediction time window;

[0114] In this embodiment, during device operation, the prediction time window refers to the time range of the model's predictions, such as the next 5 seconds, the next 1 minute, or the next N time steps. Actual sensor data refers to the actual observation data collected by the sensors within this time range. This data is used to calculate prediction deviations to assess the accuracy of the model's predictions and further optimize model performance.

[0115] Time window determination: Set a fixed time window based on the equipment operation cycle or industrial application requirements, such as updating the forecast every 5 seconds.

[0116] Data synchronization and alignment: Use time synchronization protocols (PTP, NTP) to ensure that data timestamps from different sensors are consistent, avoiding data misalignment that affects calculations.

[0117] Data storage and query: Use time series databases (InfluxDB, TimescaleDB) to store sensor data, and quickly query the actual data of the corresponding time period through indexes.

[0118] Abnormal data filtering: During the data acquisition process, methods such as threshold filtering, statistical anomaly detection, and Kalman filtering are used to remove obviously abnormal sensor data and improve calculation accuracy.

[0119] In press pressure monitoring, time windows can be determined based on stroke intervals, such as calculating prediction deviations every 10 strokes. In hydraulic system fault prediction, data synchronization ensures alignment between pressure and flow sensor data, guaranteeing timely analysis. In motor load prediction, data storage and querying enable rapid extraction of current fluctuation data within a specific time window for error analysis.

[0120] Prediction bias refers to the numerical error between the predicted value generated by the model and the actual observed value, and is used to measure the model's predictive accuracy. The calculation method for prediction bias should select an appropriate error metric based on the application scenario and combine statistical analysis to optimize the model's error correction mechanism.

[0121] Error calculation methods can include: Mean Squared Error (MSE): used to quantify overall prediction error, suitable for continuous variable predictions such as pressure and temperature. Mean Absolute Error (MAE): suitable for scenarios sensitive to error magnitude, such as vibration signal prediction. Mean Percent Error (MAPE): suitable for multi-sensor fusion predictions, such as hydraulic system state prediction. Dynamic Error Weighting: uses different weights for different sensor data to improve the rationality of error calculation.

[0122] Error trend analysis can include using a sliding window to analyze the temporal trend of the calculation error to assess model stability. Alternatively, error distribution visualization (such as histograms and error curves) can be used to provide statistical characteristics of the prediction error and optimize the prediction model.

[0123] Error compensation and optimization can include:

[0124] Error regression correction: This method primarily targets short-term error compensation. It adjusts the current forecast results by analyzing historical error patterns (e.g., the error from the previous forecast). This is a form of local optimization. It is suitable for real-time error correction, such as adjusting the forecast error of a stamping machine within a short time window.

[0125] Bayesian optimization dynamically adjusts model parameters: This method primarily optimizes global model parameters. Through Bayesian optimization, hyperparameters are automatically adjusted across multiple prediction tasks to optimize overall prediction capabilities. This is a form of global optimization. This method is suitable for optimizing a model's adaptability to new operating conditions after long-term equipment operation, such as optimizing the accuracy of long-term pressure predictions in hydraulic systems.

[0126] Incremental learning focuses on continuous learning and adaptability of the model. By continuously learning new data, the model gradually adapts to different equipment operating conditions and environmental changes, thereby reducing long-term error accumulation. This is suitable for long-term equipment maintenance, such as learning the wear trends of stamping equipment, so that the model can adapt to changes in molds from batch to batch.

[0127] By acquiring actual sensor data corresponding to the prediction time window, the comparability of prediction results with actual operating conditions is ensured, improving the prediction verification capability. Error calculation, trend analysis, and compensation optimization are used to optimize model performance, reduce prediction errors, and improve prediction accuracy and stability. This allows for more precise identification of prediction errors, enhances the model's long-term adaptability, and provides reliable data support for intelligent maintenance of industrial equipment.

[0128] S60: When the prediction deviation exceeds a preset threshold, an alarm signal is triggered.

[0129] In this embodiment, prediction deviation refers to the difference between the model's predicted value and the actual value collected by the sensor. This value is used to determine the model's prediction accuracy. If the prediction deviation exceeds a set threshold, it means that the model's prediction result may no longer be reliable, thus triggering further response mechanisms.

[0130] Benchmark error calculation: Based on indicators such as mean square error (MSE), mean absolute error (MAE), and root mean square error (RMSE), the prediction error of multiple time steps is calculated and compared with the set threshold.

[0131] Dynamic threshold adjustment: Based on the operating status of the device, an adaptive error threshold adjustment mechanism is adopted to ensure that false alarms are not triggered when the device is operating normally, and timely response is provided when fault signs appear.

[0132] Historical error comparison: Stores forecast error data over a period of time and analyzes error trends. If the error continues to increase, it may indicate an abnormal device status and an alarm needs to be triggered in advance.

[0133] During the operation of a stamping machine, each stroke generates pressure fluctuation data. The model calculates the error between the predicted and actual pressure values for each stroke and determines whether an anomaly exists based on a set threshold. In hydraulic system monitoring, the predicted error in hydraulic pressure is analyzed in real time. If it exceeds a threshold, it may indicate a leak or blockage risk in the hydraulic system, requiring further protective measures. In motor load prediction, a sliding window is used to calculate the error trend. If the error increases over multiple consecutive time steps, an early warning alert is triggered.

[0134] Alarm signal triggering means that when the prediction deviation exceeds the threshold, the system will activate the corresponding alarm mechanism, notify the operator or automatically implement protective measures. This triggering mechanism should be real-time and adaptable to different application scenarios.

[0135] A graded alarm mechanism: Different levels of alarm signals are used based on the severity of the prediction deviation. For example, a low-level alarm may be a warning message, while a high-level alarm may require an emergency shutdown or automatic protection measures.

[0136] Multi-channel alarm notification: Alarm signals can be delivered through multiple means such as sound and light alarms, SMS, email, industrial control systems (such as SCADA), mobile APP notifications, etc. to ensure that operators can receive information in a timely manner.

[0137] Alarm hysteresis protection: A fault confirmation time window is used to avoid false alarms caused by short-term error fluctuations. For example, a formal alarm is triggered only when the value exceeds the threshold three times in a row.

[0138] In die wear monitoring for stamping equipment, when the predicted pressure deviation exceeds a threshold, a low-level alarm is first triggered, alerting the operator to inspect the equipment. If the deviation persists, a high-level alarm is triggered, automatically adjusting stamping parameters or shutting down the machine for inspection. In hydraulic system pressure monitoring, if the predicted pressure value is abnormal, the operator is notified via the SCADA system and a log is recorded to facilitate subsequent analysis of whether hydraulic component replacement is necessary. In motor load prediction, if the power prediction error is consistently large, a push notification can be sent via the mobile app to alert maintenance personnel to conduct an inspection.

[0139] In certain high-precision industrial scenarios, such as die monitoring for stamping equipment, a fixed error threshold is required. For example, if the error in the predicted stamping force exceeds 10%, the system triggers an alarm. This approach is suitable for production environments with relatively stable errors. Its advantages include simple calculations and fast alarm response, making it suitable for high-precision manufacturing processes. Using a fixed threshold, for example, an alarm is triggered if the MSE exceeds a certain value. This is combined with a sliding window error calculation to avoid false alarms caused by single error fluctuations.

[0140] In complex environments like hydraulic systems and CNC machine tools, errors can be affected by factors like temperature and equipment aging. Fixed thresholds can easily lead to false alarms or missed alarms. Therefore, adaptive threshold algorithms can be used to dynamically adjust error thresholds based on equipment status, improving alarm accuracy. Historical error statistical analysis can be used to dynamically adjust thresholds based on the current state of the equipment. Incorporating anomaly detection models, such as error drift detection based on time series analysis, ensures that alarms are triggered at the appropriate time.

[0141] For applications such as motor load monitoring and stamping equipment stroke anomaly detection, a multi-level alarm mechanism ensures that different response strategies are adopted for anomalies of varying severity. If the error deviation is less than 10%, only a log is recorded, and no alarm is triggered. If the error deviation is between 10% and 20%, a low-level alarm is triggered, prompting operations and maintenance personnel to conduct an inspection. If the error deviation is greater than 20%, a high-level alarm is triggered, automatically adjusting equipment parameters or suspending operations.

[0142] By combining error calculation, dynamic threshold adjustment, and alarm signal triggering mechanisms, the reliability and real-time nature of prediction results can be effectively improved. This can reduce false positives and missed positives, improve alarm accuracy, and enhance the automated operation capabilities of equipment in intelligent manufacturing environments.

[0143] The present invention relates to the fields of artificial intelligence technology and equipment fault prediction and detection technology, and discloses a fault prediction and detection method based on an intelligent model, comprising: obtaining and preprocessing historical sensor data to generate time series data; dividing the time series data into historical data segments, pre-training the historical data segments based on a self-attention mechanism to generate a pre-trained model; concatenating the historical data segments with prediction blank segments into an input sequence, fine-tuning the pre-trained model in combination with a position-aware attention mask to generate a fine-tuned model; obtaining current sensor data, inputting the current sensor data into the fine-tuned model to generate predicted data within a prediction time window; obtaining actual sensor data within the prediction time window, analyzing the prediction deviation between the predicted data and the actual sensor data; and triggering an alarm signal when the prediction deviation exceeds a preset threshold. The present invention improves the accuracy of fault prediction by learning the long-term dependency of time series data through a pre-trained model and performing multi-segment parallel prediction in combination with a position-aware attention mask; performing real-time prediction through the fine-tuned model, thereby enhancing the adaptability of the model under different working conditions; triggering an alarm by analyzing the deviation between the predicted data and the actual data, thereby realizing intelligent monitoring of the equipment operating status and improving the real-time performance and reliability of the system.

[0144] In one embodiment, the above S10 includes:

[0145] S101, obtaining multimodal sensor historical data;

[0146] S102, performing noise smoothing processing on the multimodal sensor historical data using a Kalman filter technique;

[0147] S103, performing interpolation processing on the multimodal sensor historical data after the noise smoothing processing;

[0148] S104 , performing standardization processing on the multimodal sensor historical data after the interpolation processing to generate the time series data.

[0149] In this embodiment, during the operation of industrial equipment, sensors collect data from a variety of sources, including pressure, vibration, temperature, current, strain, displacement, and other types. This data typically comes from different sensors installed in key locations such as the main stamping mechanism, hydraulic system, drive motor, and mold assembly, and is used to monitor equipment status. The acquisition of historical multimodal sensor data is the foundation for subsequent preprocessing.

[0150] Data acquisition system: Use industrial Ethernet, wireless sensor network (WSN), field bus (PROFIBUS, MODBUS), etc. for data transmission to ensure synchronous data collection.

[0151] Data storage and indexing: Use time series databases (InfluxDB, TimescaleDB) to store historical data and perform time indexing to facilitate subsequent query and processing.

[0152] Data format conversion: Different sensors may have different sampling frequencies and data formats, so format conversion is required to ensure data availability.

[0153] In stamping equipment, pressure sensors record real-time pressure values during each stamping cycle, vibration sensors record vibration signals caused by impact forces, and temperature sensors monitor changes in stamping die temperature. All sensor data is stored locally on edge computing nodes or cloud servers as historical data. In the hydraulic system, data from flow sensors, pressure sensors, and current sensors are synchronously stored to ensure subsequent time alignment and analysis.

[0154] Sensor data collection can be affected by factors such as environmental interference, electromagnetic interference, and measurement errors, resulting in high-frequency noise in the data and affecting prediction accuracy. Kalman filtering is an optimal recursive estimation algorithm used to smooth noisy signals and improve data quality.

[0155] State estimation model: Establish observation state equations and system state equations to estimate real data.

[0156] Recursive calculation: Kalman filtering uses a prediction-update recursive method to calculate the optimal estimate of the current time step and reduce the impact of noise on the data.

[0157] Dynamically adjust weights: By updating the covariance matrix, dynamically adjust the filter parameters to adapt to data changes in different working conditions.

[0158] In vibration monitoring of stamping equipment, sensors can be affected by factors such as mechanical resonance and background vibration, resulting in high-frequency jitter in the data. Kalman filtering removes sudden interference through smoothing, improving signal stability. In hydraulic system pressure monitoring, Kalman filtering is used to smooth pressure fluctuations, reduce errors caused by sensor response delays, and improve data consistency.

[0159] In the actual operating environment of industrial equipment, data loss or discontinuity is a common problem, such as inconsistent sensor sampling intervals, data interruptions caused by communication failures, and short-term sensor failures. Interpolation processing is used to supplement missing data points to ensure the integrity of the data sequence.

[0160] Linear interpolation: Applicable to situations where data changes smoothly within a short period of time, such as interpolation filling of temperature sensor data.

[0161] Spline Interpolation: Applicable to nonlinearly changing data, such as pressure change trend compensation in stamping equipment.

[0162] Lagrange interpolation and polynomial interpolation: suitable for prediction scenarios with high precision requirements, such as pressure curve supplementation in hydraulic systems.

[0163] In pressure prediction for stamping equipment, if sensor data is lost for a short period of time, linear interpolation is used to supplement the missing points to maintain the integrity of the time series. In hydraulic system monitoring, spline interpolation is used to ensure smooth transitions between pressure data sampling points, avoiding sudden changes that could lead to misjudgment of abnormalities.

[0164] Different sensors may have different data units and value ranges. For example, temperature is expressed in degrees Celsius (°C), current is expressed in amperes (A), and vibration is expressed in g. Therefore, standardization is required to make data from different sensors comparable on the same scale.

[0165] Min-Max Scaling: Converts data to the [0, 1] range. This is suitable for situations where the maximum and minimum values are known, such as the normalization of pressure sensor data.

[0166] Z-score normalization: uses mean and standard deviation normalization, which is suitable for situations where the data distribution is unknown or has large fluctuations, such as the normalization of vibration signals.

[0167] Logarithmic transformation: used to deal with situations where data distribution is severely skewed, such as abnormal fluctuation data before equipment failure.

[0168] In stamping equipment data processing, Z-score normalization is performed on pressure, vibration, and temperature data to ensure balanced data features and improve model training results. In hydraulic system monitoring, data from different sensor types is normalized to ensure uniformity of model input data and improve prediction accuracy.

[0169] In the vibration monitoring application of stamping equipment, edge computing nodes are used to perform real-time preprocessing of sensor data to improve data quality: sensor data is acquired through a field bus (such as CAN) and synchronized; Kalman filtering is used to remove signal noise and keep data smooth; linear interpolation is combined to supplement short-term data loss and ensure data continuity; Z-score standardization is used to unify the scale of different sensor data and improve the adaptability of the model.

[0170] In the predictive maintenance scenario of hydraulic systems, historical data is stored in the cloud and regularly preprocessed in batches: a time series database (TSDB) is used to store sensor data; Bayesian filtering is used to optimize signal noise removal to improve data accuracy; spline interpolation is used to fill missing data points to ensure time series integrity; and min-max normalization is combined to process different types of data to improve subsequent prediction accuracy.

[0171] This embodiment uses multimodal sensor historical data collection, Kalman filtering for noise reduction, interpolation and completion, and standardized preprocessing to effectively improve data quality, reduce the impact of noise on model training, and ensure the stability of input data. It can complete data preprocessing in real time, reduce error accumulation, and improve prediction accuracy. It is suitable for state prediction applications in industrial equipment such as stamping equipment, hydraulic systems, and motor loads.

[0172] In one embodiment, the above S20 includes:

[0173] S201, dividing the time series data into historical data segments of preset fixed length;

[0174] S202, performing high-dimensional embedding mapping on the historical data segments to generate input features;

[0175] S203, calculating the time step correlation of the input features using a self-attention mechanism through an autoregressive encoder, and extracting the dependency between historical data segments;

[0176] S204, during the calculation process of the self-attention mechanism, a mask mechanism is applied to shield the data after the current time step, so that the model performs causal learning only based on the data of the current time step and the data before the current time step, and generates a pre-trained model based on the correlation and dependency relationship of the time steps.

[0177] In this example, time series data is a collection of chronologically ordered measurements, potentially including pressure, temperature, vibration, and current data from multimodal sensors. Because the model cannot process the entire time series at once, it needs to be divided into fixed-length historical data segments for batch training and computation.

[0178] Fixed window partitioning: divide the data according to a fixed time step (such as 10ms, 100ms) or a fixed number of data points (such as 128 points, 256 points).

[0179] Sliding window strategy: Using a sliding window (overlapping) method, there is partial overlap between each segment to enhance the temporal continuity of the model.

[0180] Data alignment and padding: For fragments that are insufficient in length, use zero-padding or repeat padding to ensure that the fragments are of consistent length.

[0181] In press equipment pressure prediction, pressure sensor data is segmented into historical data segments every 100 time steps to ensure data segmentation consistency. In hydraulic system monitoring, data is segmented using a sliding window (50% overlap) to enhance the continuity of time series information and improve prediction accuracy.

[0182] Time series data is usually a low-dimensional sequence of values. For example, a temperature sensor outputs a single value. The role of high-dimensional embedding mapping is to convert low-dimensional time series data into a high-dimensional representation, enabling the model to more effectively learn the temporal relationships and characteristic patterns of the data.

[0183] Linear Projection: Map low-dimensional data to a high-dimensional space through weight matrix mapping, for example, from a 1D vector to a 128D or 256D feature vector.

[0184] Positional Encoding: Adding temporal position information to each embedding vector ensures that the model can recognize the temporal order of data points.

[0185] Nonlinear Projection (MLP or CNN Projection): Use a multi-layer perceptron (MLP) or convolutional neural network (CNN) to learn more complex feature representations and improve understanding of temporal patterns.

[0186] In stamping equipment prediction, linear transformation embedding is used to map pressure and vibration data into 128-dimensional input features to enhance feature expression. In hydraulic system prediction, position encoding is used to enable the model to capture the time-varying patterns of fluid pressure, improving prediction capabilities.

[0187] In time series forecasting tasks, the model needs to identify the time-step correlation of the data, that is, the relationship between different time points, and extract the dependencies between data segments to understand long-term trends and short-term fluctuations.

[0188] Autoregressive Encoder: At each time step, the information of the previous time step is used to predict the current time step, ensuring the temporal dependency of the data.

[0189] Self-Attention Mechanism: Calculates the relevance weights between all time steps to ensure that the model can focus on the most relevant historical data.

[0190] Transformer Layers: Stack multiple Transformer encoding layers to improve the model's feature extraction capabilities and learn long-term dependency patterns.

[0191] In stamping equipment data processing, an autoregressive encoder is used. The model only uses past time step data to predict the current time step, improving the causal nature of time series. In hydraulic system prediction, a self-attention mechanism is used to extract the dependency between pressure fluctuations and flow rate changes, improving prediction accuracy.

[0192] In causal learning tasks, the model should predict future data points based only on past data points, without accessing information from future time steps. Therefore, a masking mechanism is needed to ensure that the model's prediction process conforms to the temporal order.

[0193] Masking mechanism: When calculating self-attention, all values after the current time step are set to invalid, that is, future data is masked.

[0194] Causal Learning: During training, the model can only see data from the current time step and historical time steps, thereby learning the true causal relationship.

[0195] Sequence training and self-supervised learning: Self-supervised learning is used to mask some historical data for training to enhance the generalization ability of the model.

[0196] In the pressure prediction task for stamping equipment, a masking mechanism ensures that the model only uses the current and previous stroke data when predicting the next stroke pressure, ensuring causal prediction. In hydraulic system fault detection, causal learning is used to ensure that the model can correctly learn the temporal dependencies of fluid pressure fluctuations, preventing future information leakage that could lead to prediction bias.

[0197] For example, in the prediction task for stamping equipment, fixed window partitioning is used to split the data into segments of equal length (e.g., 128 time steps). Linear transformation embedding is used to map the raw data into 128D features. Self-attention is used to calculate time-step correlations and extract pressure variation patterns. A masking mechanism ensures causal relationships, prevents future information leakage, and improves prediction accuracy.

[0198] In hydraulic system monitoring tasks, we use sliding window partitioning to increase time series continuity. We use CNN for feature embedding to extract short-term patterns in fluid pressure. We employ a multi-layer transformer to learn long-term dependencies. Self-supervised causal learning improves model generalization by training with partially masked historical data.

[0199] This implementation utilizes historical data segmentation, high-dimensional embedding, autoregressive encoders, self-attention mechanisms, and masked causal learning to ensure the model effectively learns long-term dependencies in time series, improving forecast accuracy. It can capture more complex temporal patterns, enhance forecast stability, and reduce errors caused by future information leakage. It is suitable for a variety of industrial scenarios, including stamping equipment, hydraulic systems, and motor load forecasting.

[0200] In one embodiment, the above S30 includes:

[0201] S301, splicing the historical data segments and the predicted blank segments according to a preset rule to form an input sequence;

[0202] S302, generating a position-aware attention mask based on the temporal relationship of the input sequence;

[0203] S303: input the input sequence into the pre-trained model, and apply the position-aware attention mask when the pre-trained model calculates the self-attention weight to block the predicted blank segment from accessing data after the current time step, and perform multi-period parallel processing on the input sequence through a shared decoding layer;

[0204] S304: Based on the output of the shared decoding layer, adjust the weight parameters of the pre-trained model to generate a fine-tuned model.

[0205] In this example, in a time series forecasting task, the historical data segments represent the time series that the model has already observed, while the prediction gap segments represent the future time steps that need to be predicted. To enable the model to handle both historical information and future time step prediction tasks, the two need to be concatenated to form a unified input sequence.

[0206] Time step arrangement: Historical data segments are arranged in chronological order, while forecast blank segments serve as placeholders for future periods to form the input sequence.

[0207] Stitching strategy: A fixed ratio of stitching (such as 80% history + 20% prediction) or a dynamic stitching ratio based on task requirements can be used.

[0208] Placeholder padding: In the prediction blank segment, the initial value can be filled with zero padding (Zero Padding) or special symbols (such as MASK) to prevent it from affecting the model calculation.

[0209] In stamping equipment failure prediction, a splicing method using 80% historical data and 20% blank prediction segments is used to ensure that the model has sufficient historical information while making reasonable predictions for future time periods. In hydraulic system pressure prediction, a variable-length splicing strategy is used to dynamically adjust the prediction time window to improve prediction accuracy.

[0210] Position-aware attention masks are used to control how the model focuses on information at different time steps, ensuring a reasonable weight distribution of historical data while shielding information leakage from future time steps.

[0211] Based on the time step index, a time series number is generated for the historical data segments and the prediction blank segments. An attention mask matrix is constructed to ensure that historical time steps can pay attention to each other, while the prediction blank segments can only access historical time steps. Positional encoding is used to provide temporal position information, ensuring that the model can perceive the temporal relationship between data points.

[0212] In the stamping equipment prediction task, position-aware attention masking ensures that only past data is accessible at the current time step, preventing information leakage at future time steps and improving prediction rationality. In the hydraulic system prediction task, variational position encoding is used to ensure that information weights are properly distributed across different time steps, improving model prediction performance.

[0213] During the training process, the pre-trained model calculates the attention weights between all time steps. In order to prevent the model from directly accessing the information of future time steps, it is necessary to use an attention mask to shield the predicted blank segments, thereby ensuring that the model learns according to causal relationships.

[0214] Masking mechanism, during the attention calculation process, blocks the information of future time steps so that the prediction is calculated only based on historical data.

[0215] Shared Decoder Layers share the same set of decoding layer weights during the prediction process of multiple time steps to reduce computational overhead and improve temporal consistency.

[0216] Multi-period parallel processing allows predictions to be made for multiple time steps simultaneously, improving computational efficiency.

[0217] In the stamping equipment prediction task, a shared decoding layer is used to calculate multiple prediction time steps, ensuring the consistency of the output results and improving prediction accuracy. In the hydraulic system prediction task, a multi-head self-attention mechanism is used to simultaneously calculate the prediction results for multiple future time steps, improving the time series modeling capabilities.

[0218] Fine-tuning refers to further training an existing pre-trained model with new data to optimize its performance on a specific task.

[0219] Gradient descent-based parameter optimization adjusts the weights of the pre-trained model to minimize the prediction error.

[0220] Backpropagation calculates the loss between the predicted value and the true value and updates the model parameters.

[0221] Learning rate adjustment (Adaptive Learning Rate), using adaptive optimization methods (such as Adam, RMSprop) to ensure stable training convergence.

[0222] In the stamping equipment prediction task, gradient clipping is used during the fine-tuning phase to prevent gradient explosion and improve training stability. In the hydraulic system prediction task, a dynamic learning rate decay strategy is used to ensure gradual model convergence and improve prediction accuracy.

[0223] For example, in the case of pressure prediction for stamping equipment, a fixed-ratio splicing approach (80% history + 20% prediction blank segments) is used to ensure input data consistency. A linear mask is used to ensure that the model does not access future data when calculating attention. A shared decoding layer is used to simultaneously calculate predictions for multiple time steps, improving computational efficiency.

[0224] Suitable for fault prediction in hydraulic systems: Dynamic window splicing is used to adaptively adjust the prediction window size based on the device status, improving prediction stability. Variational masking is used to ensure more flexible attention calculation and improve the ability to model temporal relationships. A multi-head self-attention mechanism is used to improve the model's ability to understand data at different time steps.

[0225] This example implements an efficient time series forecasting method through concatenating historical data segments with prediction blank segments, employing position-aware attention masks, sharing decoding layers, multi-period parallel processing, and fine-tuning training. This method improves the modeling capabilities of time series data, ensures the rationality of causal relationships, and enhances forecast accuracy and computational efficiency. It is applicable to a variety of industrial forecasting scenarios, including stamping equipment, hydraulic systems, and motor loads.

[0226] In one embodiment, the above S40 includes:

[0227] S401, acquiring current sensor data, and performing noise filtering, normalization processing, and time synchronization processing on the current sensor data;

[0228] S402, performing data segmentation on the processed current sensor data to generate current data segments consistent with the format of the historical data segments;

[0229] S403, converting the current data segment into input features based on high-dimensional embedding mapping, and inputting the input features into the fine-tuned model;

[0230] S404: Calculate the time step correlation of the input features through a self-attention mechanism, and generate prediction data within a prediction time window based on the time step correlation.

[0231] In this embodiment, during the operation of industrial equipment, sensors collect various types of data in real time, including pressure, vibration, temperature, and current. However, due to factors such as electromagnetic interference, mechanical vibration, and signal jitter, the raw data may contain noise. Furthermore, different sensors may have different sampling frequencies, necessitating time synchronization to ensure that all sensor data is aligned to the same time base.

[0232] Noise filtering: Kalman filtering can be used to smooth high-frequency noise, suitable for recursive updates of sensor data. Low-pass filtering can also be used to remove high-frequency interference, suitable for smoothing vibration and current signals. Wavelet denoising can also be used to separate useful signals from noise, suitable for processing non-stationary signals in stamping equipment.

[0233] Normalization: Min-Max Scaling can be used to ensure consistent numerical ranges for data from different sensors, improving model stability. Z-score normalization (conversion to a standard normal distribution) can also be used, which is suitable for scenarios with large skewed data distributions, such as when pressure peaks fluctuate significantly during stamping.

[0234] Time synchronization: Hardware clock synchronization (GPS timing / NTP synchronization) can be used, which is suitable for data alignment of distributed devices. Interpolation alignment can also be used. When the sensor data sampling intervals are inconsistent, linear interpolation or spline interpolation is used to supplement the missing points.

[0235] In stamping equipment prediction, pressure and vibration sensor data have different sampling frequencies, requiring interpolation alignment. Low-pass filtering is also used to remove noise and improve signal quality. In hydraulic system prediction, data from multiple sensors is aligned using NTP time synchronization and normalized using Z-scores to ensure consistent numerical ranges for data like pressure and current, improving model calculation stability.

[0236] In time series prediction tasks, the model needs to receive data segments in the same format, so the current sensor data needs to be divided in a way consistent with the historical data segments to ensure that the input format of the prediction model matches.

[0237] Fixed Window Segmentation: divides the current sensor data into segments according to a fixed time window (such as 100ms, 500ms) or a fixed number of samples (such as 128, 256 data points).

[0238] Sliding Window Segmentation: Use sliding windows (e.g., 50% overlap) to ensure temporal continuity between data segments and improve prediction stability.

[0239] In the prediction of stamping equipment, a fixed window approach is used to ensure that data within each stamping cycle is fully captured, improving the input consistency of the model. In the prediction of hydraulic systems, a sliding window approach is used to enhance the ability to capture short-term fluctuations and improve the accuracy of system predictions.

[0240] Industrial sensor data is typically low-dimensional numerical sequences, making it difficult to directly capture the complex relationships between data. High-dimensional embedding mapping transforms low-dimensional data into high-dimensional feature vectors, enhancing the model's ability to understand temporal patterns.

[0241] Linear Projection: Converts raw sensor data into high-dimensional features, such as from 1D to 128D, through weight matrix mapping.

[0242] Nonlinear mapping (MLP / CNN): Use multi-layer perceptron (MLP) or convolutional neural network (CNN) to learn complex feature representations.

[0243] Time series feature enhancement: Positional encoding can be used to enhance the representation of time series information. Fourier transform can also be used to extract frequency domain features to improve the model's understanding of periodic signals.

[0244] In stamping equipment prediction, linear transformation and position encoding are used to ensure that the temporal information of data segments can be effectively perceived by the model, improving prediction accuracy. In hydraulic system prediction, CNN is used for high-dimensional mapping to extract local patterns in pressure data and improve feature expression capabilities.

[0245] Time step correlation represents the relationship between data at different time steps. It is used to capture long-term and short-term dependencies and ensure that the model can correctly understand the temporal pattern.

[0246] Self-Attention Mechanism: Calculates attention weights between all time steps to ensure that the model can focus on the most relevant historical data.

[0247] Multi-Head Attention: Uses multiple attention heads to improve the model's ability to capture features at different time scales.

[0248] Transformer-based time series prediction: Causal masking is used to ensure that the model only uses information from past time steps, improving prediction rationality. Residual connections are used to improve gradient stability and prevent model degradation.

[0249] In stamping equipment prediction, self-attention is used to calculate time-step correlation, ensuring that the model focuses on key change points in the stamping process and improving prediction accuracy. In hydraulic system prediction, multi-head attention is used to ensure that the model can capture both short-term pressure changes and long-term trends, improving prediction reliability.

[0250] For example, in the pressure prediction of stamping equipment, fixed window partitioning ensures data input consistency, linear mapping + position encoding enhances time series features, and self-attention calculates time step correlation to ensure the rationality of prediction.

[0251] Suitable for fault prediction in hydraulic systems: Sliding window partitioning enhances data continuity. CNN performs high-dimensional mapping to improve the ability to extract short-term patterns. Multi-head self-attention mechanism improves prediction stability.

[0252] This embodiment uses noise filtering, normalization, time synchronization, data segmentation, high-dimensional embedding mapping, and self-attention to calculate time-step correlations to ensure the model effectively learns complex patterns in time series data and improves prediction accuracy. This improves prediction accuracy, reduces noise interference, and enhances equipment monitoring reliability, making it suitable for a variety of industrial prediction scenarios, including stamping equipment, hydraulic systems, and motor loads.

[0253] In one embodiment, the above S50 includes:

[0254] S501, acquiring actual sensor data within the prediction time window, and performing data synchronization processing on the actual sensor data to align timestamps of sensors of different types;

[0255] S502, performing format standardization processing on the actual sensor data based on the prediction time window, so that the actual sensor data and the predicted data are consistent in data format, data type and time step;

[0256] S503: Calculate the prediction deviation between the predicted data and the actual sensor data within the prediction time window.

[0257] In this example, when industrial equipment is operating, multiple sensors (such as pressure, vibration, temperature, and current) collect data separately. Due to different sampling rates, data transmission delays, or storage mechanisms, the data from different sensors may be time-shifted. Therefore, data synchronization must be performed to align the observation data from all sensors to the same time base to facilitate reasonable error analysis.

[0258] Alignment can be based on time interpolation. When timestamps from different sensors don't match, linear or spline interpolation is used to estimate the data values for the missing time steps, aligning all data to a single time step. Synchronous sampling can also be used to ensure that all sensors collect data synchronously at fixed intervals, avoiding later alignment errors. Alignment can also be based on a reference signal. For example, using a high-precision timing signal (such as a GPS timestamp or an industrial network clock) can realign all sensor data to the same time base, reducing calculation errors caused by time drift.

[0259] In the prediction task for stamping equipment, since the sampling frequencies of pressure and vibration sensors vary, time interpolation can be used to align the data to ensure consistent timestamps for all data when calculating prediction errors. In the prediction task for hydraulic systems, synchronous sampling can be used to ensure that all sensors collect data according to the same clock, reducing the accumulation of time errors.

[0260] Since the storage format, data type, and time step of different sensor data may be different, a unified format is needed to ensure that the data can be correctly compared.

[0261] Data type conversion can be used to convert all sensor data to floating-point numbers for calculations, avoiding errors caused by incompatible data formats. Unit normalization can also be performed, such as converting temperature from Celsius to standard units and pressure data to MPa, to ensure that data is calculated on the same scale. Time step alignment can also be performed to interpolate missing data points, ensuring that predicted and actual data are consistent in time, improving calculation accuracy.

[0262] In the prediction task for stamping equipment, it is necessary to normalize the data of different types of sensors so that they use the same data scale during calculations to ensure valid comparison results. In the prediction task for hydraulic systems, it is necessary to use time step alignment methods to ensure that the error calculation of the predicted data and the actual observed data is performed at the same time step, improving the accuracy of the analysis.

[0263] In the prediction system, prediction deviation is the core indicator to measure whether the prediction results are accurate. By calculating the deviation between the predicted value and the actual observed value, the reliability of the model can be judged and optimization direction can be provided.

[0264] The absolute deviation method can be used to directly calculate the difference between the predicted value and the actual observed value to measure the error size. Alternatively, the sum of squared errors method can be used to nonlinearly amplify the error, making larger prediction errors have a greater impact on the optimization process. Relative error calculation can also be used to measure the proportion of prediction error to the actual observed value, making the error calculation applicable to data of different orders of magnitude.

[0265] In the prediction task for stamping equipment, the absolute deviation calculation method is used to obtain the prediction error at each time step during a single stamping process, improving the error assessment capability at key points. In the prediction task for hydraulic systems, the sum of squared errors method is used, allowing the system to focus more on key data points when predicting large errors, improving the accuracy of the optimization direction.

[0266] It should be noted that in industrial equipment failure prediction, deviation calculation can be performed for each type of data separately, or it can be performed by integrating data from multiple sensors and evaluating the overall deviation through comprehensive error. The specific method used depends on the actual application scenario and prediction requirements.

[0267] Different sensor data such as pressure, vibration, and temperature have different physical quantities, data units, and variation ranges, so it is usually necessary to calculate the deviation separately. For example:

[0268] Pressure deviation: used to evaluate whether the mold punching pressure of the equipment is abnormal.

[0269] Temperature deviation: used to monitor whether the ambient temperature or device heating exceeds expectations.

[0270] Vibration deviation: used to determine whether there is abnormal vibration in the mechanical system, such as bearing failure.

[0271] Applicable scenarios:

[0272] Single parameter critical impact failure: If a certain sensor data (such as pressure) has a significant impact on the health status of the equipment, it is more meaningful to calculate the deviation of this parameter alone.

[0273] Independent analysis of multi-parameter changes: Some systems require separate attention to changes in pressure, temperature, and vibration rather than the overall trend. For example, mold pressure monitoring for stamping equipment only requires attention to pressure anomalies without considering temperature deviations.

[0274] For example, in stamping equipment prediction, pressure data deviation is a key concern. The system can calculate the deviation between predicted and actual pressure to assess whether the mold is worn or the equipment rigidity has decreased. In hydraulic system fault prediction, temperature data deviation can be used to identify abnormal system pressure caused by excessive hydraulic oil temperature, so calculating temperature deviation separately is particularly important.

[0275] Furthermore, in some cases, equipment failure may be the result of the combined influence of multiple parameters. For example, abnormal pressure, elevated temperature, and increased vibration may all point to a specific failure mode. Therefore, a comprehensive error calculation method can be used to integrate the prediction errors of multiple sensors to obtain an overall deviation assessment.

[0276] Applicable scenarios:

[0277] Multi-parameter correlation affects faults: If a fault occurs, it usually involves multiple data dimensions (such as vibration, temperature, and pressure), and it is necessary to combine multiple sensor data for comprehensive deviation calculation.

[0278] Fault types are complex and cannot be determined independently by a single sensor. For example, in some mechanical systems, vibration or temperature alone may not be enough to determine whether the equipment is abnormal. A combination of the two may be more valuable.

[0279] Different weights are assigned to each sensor based on its historical impact, and a weighted overall deviation is calculated. For example, if historical data indicates that pressure contributes most to fault prediction, the deviation weight of pressure data should be higher. Feature fusion methods, such as combining the deviation vectors of pressure, vibration, and temperature, are used to calculate the overall deviation using machine learning models or statistical methods.

[0280] For example, when predicting robot joints, torque, vibration, and temperature deviations must be combined, as joints can malfunction due to friction, excessive temperatures, or insufficient motor power. In stamping equipment predictions, if pressure, vibration, and temperature deviations are all simultaneously outside normal ranges, this could indicate mold cracks or abnormal hydraulic system pressure, necessitating a comprehensive assessment.

[0281] Another approach is to first calculate the independent deviation of each sensor, and then make an overall judgment through deviation aggregation analysis. The characteristics of this approach are:

[0282] It is possible to calculate the errors of each data type separately while taking into account the correlation between them in the final analysis. Avoid directly merging different physical quantities, which may cause scaling issues. Because pressure, temperature, and vibration have different numerical ranges, directly merging them may cause the effects of certain parameters to be amplified or ignored.

[0283] Example: In stamping equipment prediction, you can first calculate pressure deviation, vibration deviation, and temperature deviation separately. Then, when judging anomalies, consider whether multiple deviations simultaneously exceed the threshold to reduce false positives or omissions caused by errors in a single indicator.

[0284] If the trend of a single key parameter (such as pressure) is of interest, the pressure deviation should be calculated alone without considering other data.

[0285] If equipment failure usually involves multiple factors (such as vibration + temperature), multi-sensor fusion should be used to calculate the deviation and consider the combined impact of multiple variables.

[0286] If you want to analyze each parameter individually and evaluate the trend as a whole, you can calculate the deviations one by one and then perform a joint analysis to improve the accuracy and stability of fault prediction.

[0287] In summary, in prediction scenarios for stamping equipment, hydraulic systems, and robotics, the appropriate deviation calculation method should be selected based on the complexity of the failure mode and the correlation between the data. In practical applications, the deviation of each sensor is usually calculated first, and then the specific fault diagnosis requirements are considered to determine whether to conduct individual analysis or a comprehensive assessment.

[0288] This embodiment improves the accuracy of prediction error calculation and ensures the reliability of prediction results through steps such as data synchronization, format standardization, and prediction deviation calculation. It can also reduce the impact of time step mismatch on error calculation, improve the stability of error calculation, and enhance the adaptive optimization capabilities of the prediction system. It is suitable for a variety of industrial scenarios such as stamping equipment, hydraulic systems, and motor load prediction.

[0289] In one embodiment, after the above S60, the method further includes:

[0290] S701, based on the prediction time window corresponding to the trigger alarm signal, extracting data segments whose prediction deviation exceeds a preset threshold within the prediction time window and corresponding sensor actual data, and storing them in a data reflow module;

[0291] S702, performing local parameter adjustment on the fine-tuned model based on the latest data stored in the data reflow module;

[0292] S703, continuously monitoring changes in prediction deviations over multiple time periods, and when the cumulative prediction deviations reach a set update condition, performing periodic model optimization using the data in the data reflow module to optimize the fine-tuned model;

[0293] S704: Synchronize the optimized and fine-tuned model to multiple model deployment terminals.

[0294] In this embodiment, when the prediction deviation exceeds the set alarm threshold, the data at the time of the anomaly needs to be recorded and stored for subsequent model optimization. The prediction time window here refers to the time period corresponding to the triggering alarm signal, typically including multiple time steps before and after the alarm is triggered, so as to analyze the context of the anomaly data.

[0295] Data fragment extraction: Based on the timestamp of the alarm signal, the predicted data, actual sensor data, and relevant environmental characteristic data such as temperature and humidity within the time window are extracted from the historical data. A fixed window (such as the past 10 time steps) or a dynamic window (the window size is adjusted based on the error trend before the alarm) can be used.

[0296] Storing the extracted data fragments in the data reflow module: Using a time-series database or cloud storage, store the extracted data fragments in the data reflow module to ensure that the model can access this data for subsequent optimization. The data storage structure can use key-value pairs (KV) with "alarm timestamp + device ID" as the index to improve query efficiency.

[0297] In stamping equipment anomaly detection, if the deviation between the predicted and actual mold pressure values for a certain period exceeds a threshold, the system automatically stores all sensor data at that moment and marks the anomaly for subsequent model optimization. In hydraulic system prediction, if the system detects continuous abnormal fluctuations in hydraulic pressure, the pressure, flow, temperature, and other data for that period are stored in the data reflux module for subsequent model updates.

[0298] The stored data segments will be used for model optimization. The first step is to adjust local parameters. That is, based on the current fine-tuned model, the parameters are adjusted locally to make the model better adapt to new failure modes or changes in working conditions.

[0299] Local fine-tuning: Using incremental learning or online learning methods, fine-tuning is performed only on newly stored data without retraining the entire model. Using sliding window training, only the most recent high-error data is used for model adjustment to quickly respond to changes in failure modes.

[0300] Parameter adjustment strategy: Error regression is used to calculate the trend of model prediction error over time, and model parameters are dynamically adjusted to reduce similar errors in subsequent predictions. A gradient adjustment method is used to update only the affected network weights, ensuring that the model's global generalization ability is not affected while being locally optimized.

[0301] In stamping equipment prediction, if the pressure deviation of a batch of dies continues to increase, the model will adjust the model parameters based on this abnormal data to better adapt to the wear of the die and improve prediction accuracy. In robot joint prediction, if the torque prediction error of a joint continues to accumulate, the system will adjust the model weight to improve its adaptability to changes in the joint load.

[0302] Local parameter adjustments can quickly correct short-term errors, but to ensure long-term model stability, the overall model structure or a wider range of parameters needs to be optimized. This step monitors the forecast deviation over time and performs a more comprehensive model optimization when the cumulative error exceeds a certain threshold.

[0303] Forecast deviation monitoring: Set up an error accumulation mechanism, such as calculating the average error over the past N forecast time windows. If the accumulated error continues to exceed a set threshold, trigger periodic model optimization. Use trend analysis to calculate the error growth rate. If the error continues to grow, perform model optimization in advance.

[0304] Periodic model optimization: Batch training uses historical data stored in the data reflow module to perform global model optimization to improve the model's adaptability to long-term changes. Combined with transfer learning, data from different production environments is integrated to improve the model's generalization ability.

[0305] In the case of stamping equipment prediction, if die pressure prediction errors are consistently high over the past month, periodic optimization is triggered, using historical data for retraining to adapt to new equipment wear conditions. In the case of hydraulic system prediction, if the hydraulic pressure prediction error continues to increase over the past five cycles, periodic model updates are performed to optimize the hydraulic system's prediction capabilities.

[0306] After periodic optimization, a new fine-tuned model is generated and needs to be deployed to different production sites or multiple equipment ends to ensure that all relevant systems can use the latest optimized model to improve prediction accuracy.

[0307] Cloud synchronization: Using an edge computing architecture, after model optimization in the cloud, model parameters are pushed to each edge node through the API interface, enabling distributed model updates. Model version management ensures that new models are fully verified before deployment to avoid impacting production.

[0308] Local updates: Adopt an incremental update strategy, synchronizing only adjusted parameters rather than the entire model to reduce network transmission burden. Combined with grayscale releases, the new model is first tested on a limited number of devices to ensure that it performs better than the old version before being pushed to a larger scale.

[0309] For stamping equipment prediction, factories can deploy updated models through edge computing nodes, enabling all stamping presses to use the latest prediction algorithms and improving prediction accuracy across the entire production line. For hydraulic system prediction, local incremental updates can be used to synchronize only the latest parameters of the pressure prediction module without updating the entire model, improving update efficiency.

[0310] Example: Stamping equipment is widely used in the automotive, metalworking, and electronic component manufacturing industries. Its core components include dies, hydraulic systems, and transmission mechanisms. Over long-term operation, due to factors such as die wear, hydraulic system fluctuations, and equipment aging, the die pressure during the stamping process may fluctuate abnormally.

[0311] Failure to detect abnormal pressure changes in a timely manner can lead to the following problems: Reduced product quality, such as deformation and reduced precision of stamped parts, can affect subsequent assembly. Increased risk of equipment damage: prolonged abnormal pressure can cause mold cracking and hydraulic system failure. Increased maintenance costs: Failure to predict equipment failures in advance can lead to production halts and affect production line efficiency.

[0312] Traditional methods mainly rely on fixed threshold detection or simple trend analysis. However, due to the complex dynamic changes in pressure during the stamping process, traditional methods find it difficult to accurately predict pressure trends and detect potential anomalies in a timely manner.

[0313] This example uses an intelligent fault prediction model, combined with time series data preprocessing, deep learning modeling, self-attention prediction, error analysis and optimization, to predict and detect anomalies in the mold pressure of stamping equipment.

[0314] Multiple sensors are installed on key components of stamping equipment, such as dies, crank connecting rods, and hydraulic cylinders, including:

[0315] Pressure sensor: monitors the pressure fluctuations on the die during the stamping process in real time.

[0316] Vibration sensor: monitors mechanical vibrations that may affect pressure during equipment operation.

[0317] Temperature sensor: records the ambient temperature and considers the effect of temperature on the deformation of stamping parts.

[0318] These sensor data are transmitted in real time to the edge computing node via industrial Ethernet or fieldbus for pre-processing.

[0319] Clean and standardize the collected pressure, vibration, temperature and other sensor data:

[0320] Noise filtering: Kalman filtering is used to smooth high-frequency noise and remove interference caused by mechanical jitter.

[0321] Data synchronization: Use interpolation methods to align the time steps of data from different sensors to ensure that the data are comparable on a unified time axis.

[0322] Normalization: Convert pressure, vibration, and temperature data to the same scale to improve model stability.

[0323] The continuous time series is divided into historical data segments for model training, with future time steps serving as blank segments for prediction. A high-dimensional embedding method is used to convert multiple sensor data into feature vectors that the model can understand. A self-attention mechanism is employed to calculate dependencies between historical time steps and learn the long-term patterns of pressure fluctuations during the stamping process. Incorporating position-aware attention masks ensures that the model only uses historical data when predicting future time steps, preventing data leakage. Pressure trends in future time windows are predicted and output as a sequence of predicted pressure values.

[0324] Obtain the actual pressure data within the prediction time window and calculate the error between the predicted value and the actual value.

[0325] Calculate forecast deviation: Use the sum of squared errors method to increase sensitivity to large errors. Use a sliding window to analyze error trends and identify long-term cumulative errors.

[0326] Dynamically optimize the model: When the error exceeds a threshold, fine-tuning the model is triggered to improve its adaptability. Incremental learning methods are used to adapt the model to long-term changes in the equipment's operating status.

[0327] Set fault alarm threshold: If the prediction deviation exceeds the set threshold, the system triggers an alarm.

[0328] Intelligent Decision Support: If the predicted deviation continues to increase, it may indicate severe mold wear and maintenance is recommended. If the pressure fluctuates significantly within a short period of time, it may indicate a hydraulic system failure and the hydraulic cylinder seal should be checked. If the predicted pressure trend is low for a long time, it may indicate a decrease in equipment rigidity and mechanical adjustments are recommended.

[0329] Alarm method: Alarm information is pushed through the SCADA system or industrial cloud platform. Combined with a visual dashboard, it displays real-time forecast trends and abnormal analysis results.

[0330] This embodiment uses alarm-triggered data storage, local parameter adjustment, periodic optimization, and global model synchronization to enable the prediction model to continuously learn and adapt to changes in diverse production environments, improving prediction accuracy and robustness. It can automatically collect high-error data, optimize model weights, improve prediction accuracy, and reduce unexpected equipment downtime. It is suitable for a variety of industrial applications, including stamping equipment, hydraulic systems, and robotics prediction.

[0331] In one embodiment, a fault prediction and detection device based on an intelligent model is provided, and the fault prediction and detection device based on an intelligent model corresponds one-to-one to the fault prediction and detection method based on an intelligent model in the above embodiment. Figure 3 , Figure 3This is a functional module diagram of a preferred embodiment of the intelligent model-based fault prediction and detection device of the present invention. It includes a data preprocessing module 10, a pretrained model construction module 20, a model fine-tuning module 30, a real-time prediction module 40, a prediction deviation analysis module 50, and an abnormality alarm module 60. Each functional module is described in detail below:

[0332] The data preprocessing module 10 is used to obtain sensor historical data, preprocess the sensor historical data, and generate time series data;

[0333] A pre-training model construction module 20 is used to divide the time series data into historical data segments, and perform pre-training based on the historical data segments through a self-attention mechanism to generate a pre-training model;

[0334] A model fine-tuning module 30 is configured to concatenate the historical data segments and the predicted blank segments into an input sequence, and fine-tune the pre-trained model in combination with a position-aware attention mask to generate a fine-tuned model;

[0335] A real-time prediction module 40 is used to obtain current sensor data, input the current sensor data into the fine-tuned model, and generate prediction data within a prediction time window;

[0336] A prediction deviation analysis module 50 is configured to obtain actual sensor data within the prediction time window and analyze a prediction deviation between the predicted data and the actual sensor data within the prediction time window;

[0337] The abnormality alarm module 60 is used to trigger an alarm signal when the prediction deviation exceeds a preset threshold.

[0338] In one embodiment, the data preprocessing module 10 is specifically configured to:

[0339] Obtain multimodal sensor historical data;

[0340] Performing noise smoothing on the multimodal sensor historical data using a Kalman filter technique;

[0341] performing interpolation processing on the multimodal sensor historical data after the noise smoothing processing;

[0342] The multimodal sensor historical data after the interpolation processing is standardized to generate the time series data.

[0343] In one embodiment, the pre-training model construction module 20 is specifically configured to:

[0344] Dividing the time series data into historical data segments of preset fixed length;

[0345] Performing high-dimensional embedding mapping on the historical data segments to generate input features;

[0346] Using a self-attention mechanism through an autoregressive encoder to calculate the time-step correlation of the input features and extract the dependency between historical data segments;

[0347] During the calculation process of the self-attention mechanism, a mask mechanism is applied to shield the data after the current time step, so that the model performs causal learning only based on the data of the current time step and the data before the current time step, and generates a pre-trained model based on the correlation and dependency of the time steps.

[0348] In one embodiment, the model fine-tuning module 30 is specifically configured to:

[0349] splicing the historical data segments and the predicted blank segments according to preset rules to form an input sequence;

[0350] generating a position-aware attention mask based on a temporal relationship of the input sequence;

[0351] Inputting the input sequence into the pre-trained model, applying the position-aware attention mask when the pre-trained model calculates the self-attention weight to shield the predicted blank segment from accessing data after the current time step, and performing multi-period parallel processing on the input sequence through a shared decoding layer;

[0352] Based on the output of the shared decoding layer, the weight parameters of the pre-trained model are adjusted to generate a fine-tuned model.

[0353] In one embodiment, the real-time prediction module 40 is specifically configured to:

[0354] Acquire current sensor data, and perform noise filtering, normalization processing, and time synchronization processing on the current sensor data;

[0355] Performing data segmentation on the processed current sensor data to generate current data segments consistent with the format of the historical data segments;

[0356] Based on a high-dimensional embedding map, converting the current data segment into input features, and inputting the input features into the fine-tuned model;

[0357] The time step correlation of the input features is calculated through a self-attention mechanism, and prediction data within a prediction time window is generated based on the time step correlation.

[0358] In one embodiment, the prediction deviation analysis module 50 is specifically configured to:

[0359] Acquiring actual sensor data within the prediction time window, and performing data synchronization processing on the actual sensor data to align timestamps of sensors of different types;

[0360] Based on the prediction time window, performing format standardization processing on the actual sensor data so that the actual sensor data and the predicted data are consistent in data format, data type and time step;

[0361] The prediction deviation between the predicted data and the actual sensor data within the prediction time window is calculated.

[0362] In one embodiment, the abnormality alarm module 60 is specifically configured to:

[0363] Based on the prediction time window corresponding to the trigger alarm signal, extract the data segments whose prediction deviation exceeds the preset threshold within the prediction time window and the corresponding sensor actual data, and store them in the data reflow module;

[0364] Performing local parameter adjustment on the fine-tuned model based on the latest data stored in the data reflow module;

[0365] Continuously monitoring changes in prediction deviations over multiple time periods, and when the cumulative prediction deviation reaches a set update condition, performing periodic model optimization using the data in the data reflow module to optimize the fine-tuned model;

[0366] Synchronize the optimized and fine-tuned model to multiple model deployment terminals.

[0367] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 4 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external user terminal via a network connection. When the computer program is executed by the processor, it realizes the functions or steps of a fault prediction and detection method based on an intelligent model on the service side.

[0368] In one embodiment, a computer device is provided. The computer device may be a user terminal, and its internal structure diagram may be as follows: Figure 5As shown. The computer device includes a processor, memory, network interface, display screen and input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the user side of a fault prediction and detection method based on an intelligent model.

[0369] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:

[0370] Acquire sensor historical data, preprocess the sensor historical data, and generate time series data;

[0371] Dividing the time series data into historical data segments, and performing pre-training based on the historical data segments through a self-attention mechanism to generate a pre-training model;

[0372] splicing the historical data segments and the predicted blank segments into an input sequence, and fine-tuning the pre-trained model in combination with the position-aware attention mask to generate a fine-tuned model;

[0373] Acquire current sensor data, input the current sensor data into the fine-tuned model, and generate prediction data within a prediction time window;

[0374] Acquiring actual sensor data within the prediction time window, and analyzing a prediction deviation between the predicted data and the actual sensor data within the prediction time window;

[0375] When the prediction deviation exceeds a preset threshold, an alarm signal is triggered.

[0376] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0377] Acquire sensor historical data, preprocess the sensor historical data, and generate time series data;

[0378] Dividing the time series data into historical data segments, and performing pre-training based on the historical data segments through a self-attention mechanism to generate a pre-training model;

[0379] splicing the historical data segments and the predicted blank segments into an input sequence, and fine-tuning the pre-trained model in combination with the position-aware attention mask to generate a fine-tuned model;

[0380] Acquire current sensor data, input the current sensor data into the fine-tuned model, and generate prediction data within a prediction time window;

[0381] Acquiring actual sensor data within the prediction time window, and analyzing a prediction deviation between the predicted data and the actual sensor data within the prediction time window;

[0382] When the prediction deviation exceeds a preset threshold, an alarm signal is triggered.

[0383] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the user side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.

[0384] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0385] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0386] It should be noted that if any software tools or components other than those of the Company appear in the embodiments of this application, they are merely for illustration and do not represent actual use. The above embodiments are intended only to illustrate the technical solutions of the present invention, not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some of the technical features therein with equivalents. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the scope of protection of the present invention.

Claims

1. A fault prediction and detection method based on an intelligent model, characterized in that: The following steps are involved: Acquire sensor historical data, preprocess the sensor historical data, and generate time series data; Dividing the time series data into historical data segments, and performing pre-training based on the historical data segments through a self-attention mechanism to generate a pre-training model; splicing the historical data segments and the predicted blank segments into an input sequence, and fine-tuning the pre-trained model in combination with the position-aware attention mask to generate a fine-tuned model; Acquire current sensor data, input the current sensor data into the fine-tuned model, and generate prediction data within a prediction time window; Acquiring actual sensor data within the prediction time window, and analyzing a prediction deviation between the predicted data and the actual sensor data within the prediction time window; When the prediction deviation exceeds a preset threshold, an alarm signal is triggered.

2. The fault prediction and detection method based on intelligent model according to claim 1, characterized in that: Acquiring historical sensor data, preprocessing the historical sensor data, and generating time series data, including: Obtain multimodal sensor historical data; Performing noise smoothing on the multimodal sensor historical data using a Kalman filter technique; performing interpolation processing on the multimodal sensor historical data after the noise smoothing processing; The multimodal sensor historical data after the interpolation processing is standardized to generate the time series data.

3. The fault prediction and detection method based on intelligent model according to claim 1, characterized in that: Dividing the time series data into historical data segments, and performing pre-training based on the historical data segments through a self-attention mechanism to generate a pre-training model, including: Dividing the time series data into historical data segments of preset fixed length; Performing high-dimensional embedding mapping on the historical data segments to generate input features; Using a self-attention mechanism through an autoregressive encoder to calculate the time-step correlation of the input features and extract the dependency between historical data segments; During the calculation process of the self-attention mechanism, a mask mechanism is applied to shield the data after the current time step, so that the model performs causal learning only based on the data of the current time step and the data before the current time step, and generates a pre-trained model based on the correlation and dependency of the time steps.

4. The fault prediction and detection method based on intelligent model according to claim 1, characterized in that: The historical data segments and the predicted blank segments are concatenated into an input sequence, and the pre-trained model is fine-tuned in combination with the position-aware attention mask to generate a fine-tuned model, including: splicing the historical data segments and the predicted blank segments according to preset rules to form an input sequence; generating a position-aware attention mask based on a temporal relationship of the input sequence; Inputting the input sequence into the pre-trained model, applying the position-aware attention mask when the pre-trained model calculates the self-attention weight to shield the predicted blank segment from accessing data after the current time step, and performing multi-period parallel processing on the input sequence through a shared decoding layer; Based on the output of the shared decoding layer, the weight parameters of the pre-trained model are adjusted to generate a fine-tuned model.

5. The fault prediction and detection method based on intelligent model as claimed in claim 1, characterized in that: Acquiring current sensor data, inputting the current sensor data into the fine-tuned model, and generating predicted data within a prediction time window, including: Acquire current sensor data, and perform noise filtering, normalization processing, and time synchronization processing on the current sensor data; Performing data segmentation on the processed current sensor data to generate current data segments consistent with the format of the historical data segments; Based on a high-dimensional embedding map, converting the current data segment into input features, and inputting the input features into the fine-tuned model; The time step correlation of the input features is calculated through a self-attention mechanism, and prediction data within a prediction time window is generated based on the time step correlation.

6. The fault prediction and detection method based on intelligent model according to claim 1, characterized in that: Acquiring actual sensor data within the prediction time window, and analyzing a prediction deviation between the predicted data and the actual sensor data within the prediction time window, including: Acquiring actual sensor data within the prediction time window, and performing data synchronization processing on the actual sensor data to align timestamps of sensors of different types; Based on the prediction time window, performing format standardization processing on the actual sensor data so that the actual sensor data and the predicted data are consistent in data format, data type and time step; The prediction deviation between the predicted data and the actual sensor data within the prediction time window is calculated.

7. The fault prediction and detection method based on intelligent model according to claim 1, characterized in that: When the prediction deviation exceeds a preset threshold, after triggering an alarm signal, the method further includes: Based on the prediction time window corresponding to the trigger alarm signal, extract the data segments whose prediction deviation exceeds the preset threshold within the prediction time window and the corresponding sensor actual data, and store them in the data reflow module; Performing local parameter adjustment on the fine-tuned model based on the latest data stored in the data reflow module; Continuously monitoring changes in prediction deviations over multiple time periods, and when the cumulative prediction deviation reaches a set update condition, performing periodic model optimization using the data in the data reflow module to optimize the fine-tuned model; Synchronize the optimized and fine-tuned model to multiple model deployment terminals.

8. A fault prediction and detection device based on an intelligent model, characterized in that: The fault prediction and detection device based on the intelligent model includes: A data preprocessing module is used to obtain sensor historical data, preprocess the sensor historical data, and generate time series data; A pre-training model construction module is used to divide the time series data into historical data segments, and perform pre-training based on the historical data segments through a self-attention mechanism to generate a pre-training model; A model fine-tuning module is used to concatenate the historical data segments and the predicted blank segments into an input sequence, and fine-tune the pre-trained model in combination with the position-aware attention mask to generate a fine-tuned model; A real-time prediction module is used to obtain current sensor data, input the current sensor data into the fine-tuned model, and generate prediction data within a prediction time window; A prediction deviation analysis module, configured to obtain actual sensor data within the prediction time window and analyze a prediction deviation between the prediction data and the actual sensor data within the prediction time window; The abnormality alarm module is used to trigger an alarm signal when the prediction deviation exceeds a preset threshold.

9. A computer device, characterized in that: The computer device includes a memory, a processor, and an intelligent model-based fault prediction and detection program stored in the memory and capable of running on the processor. When the intelligent model-based fault prediction and detection program is executed by the processor, the steps of the intelligent model-based fault prediction and detection method as described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that The storage medium stores a fault prediction and detection program based on an intelligent model. When the fault prediction and detection program based on an intelligent model is executed by a processor, the steps of the fault prediction and detection method based on an intelligent model as described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Coal mill fault early warning method based on attention mechanism

    CN115730191A

  • Power grid time sequence data pre-training and prediction method and device, and pre-training model

    CN117370788A

  • Training method, prediction method, device and equipment of time sequence signal prediction model

    CN117852624A

  • Auto-regression automatic driving motion prediction method and storage medium

    CN118094481A

Cited By

  • RISC-V storage and calculation integrated chip data processing method and system

    CN121116389A

  • Fault prediction method and system for deviation rectification of laser die cutting and winding all-in-one machine

    CN121256736A

  • A fault prediction method and system for deviation correction of a laser die-cutting and winding all-in-one machine

    CN121256736B

  • Waste heat power generation time sequence data prediction and diagnosis method based on large language model

    CN121705950A