Waterproofing membrane production monitoring and early warning method and system based on multi-modal data fusion

By combining a delay compensation model and an edge-level embedded fusion model with cloud service detection algorithms, the problem of inconsistent data processing in waterproof membrane production was solved, enabling precise monitoring of the production process and rapid anomaly identification, thereby improving production quality and efficiency.

CN121259990BActive Publication Date: 2026-02-24GUIZHOU UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511758118.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-27
Publication Date
2026-02-24
Estimated Expiration
2045-11-27

AI Technical Summary

Technical Problem

In the current production process of waterproof membranes, single data analysis is insufficient to fully reflect the complexity of the production process, leading to the omission or misjudgment of potential anomalies. Furthermore, the lack of effective data processing and fusion mechanisms makes it impossible to detect and handle production anomalies in a timely manner.

Method used

Physical segment location matching is performed using a pre-established delay compensation model. Different types of features are extracted and fused. Multi-dimensional analysis is performed using an edge-level embedded fusion model. Anomaly root cause analysis is performed using cloud service detection algorithms to generate production early warning root cause tracing results.

Benefits of technology

It enables comprehensive, efficient and accurate monitoring and early warning of the waterproof membrane production process, quickly identifies potential anomalies, reduces production costs and failure risks, and improves production quality and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121259990B_ABST
    Figure CN121259990B_ABST
Patent Text Reader

Abstract

The embodiment of the application discloses a waterproof roll production monitoring and early warning method and system based on multi-modal data fusion, which comprises the following steps: matching the initial multi-modal monitoring data set containing the continuous monitoring data of each process link of the waterproof roll production line with the same point position through a pre-established delay compensation model to obtain a same point position matching data set; performing associated feature engineering processing on each group of same point position data packets in the same point position matching data set, extracting and fusing different types of features to generate a production process state feature set; inputting the feature set into an edge-level embedded fusion model for multi-dimensional analysis and outputting a preliminary warning label; and performing abnormal root cause analysis on abnormal same point position data packets based on a cloud service detection algorithm to generate a production early warning root cause tracing result containing abnormal reasons, influence degrees and position information, which is used for guiding production adjustment and fault elimination.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data analysis technology, specifically to a method and system for monitoring and early warning of waterproof membrane production based on multimodal data fusion. Background Technology

[0002] In the field of waterproof membrane manufacturing, monitoring and early warning of the production process are crucial for ensuring product quality and production efficiency. Currently, common production monitoring methods primarily rely on single-type data analysis. This approach, due to the limited data source, struggles to comprehensively reflect the complexities of the production process, easily leading to missed or misjudged potential anomalies. While some methods utilize multi-source data, they lack effective data processing and fusion mechanisms, hindering accurate correlation analysis. Furthermore, data processing delays extend response times, making it difficult to promptly detect and address anomalies in the production process. In summary, existing technologies are insufficient for timely and accurate anomaly monitoring and analysis in waterproof membrane production. Summary of the Invention

[0003] This invention provides a method and system for monitoring and early warning of waterproof membrane production based on multimodal data fusion.

[0004] In a first aspect, embodiments of the present invention provide a method for monitoring and early warning of waterproof membrane production based on multimodal data fusion, applied to a system for monitoring and early warning of waterproof membrane production based on multimodal data fusion, the method comprising:

[0005] The initial multi-mode monitoring dataset is subjected to physical segment same-point matching processing by a pre-established delay compensation model to obtain the same-point matching dataset; the initial multi-mode monitoring dataset contains continuous monitoring data of each process link of the waterproof membrane production line.

[0006] For each group of data packets at the same location in the same location matching dataset, perform association feature engineering processing, extract different types of features reflecting the production process status from each group of data packets at the same location, cross-associate and fuse the extracted different types of features, and generate a production process status feature set corresponding to each group of data packets at the same location.

[0007] The production process status feature set is input into the edge-level embedded fusion model for multi-dimensional analysis, and the primary warning label corresponding to each group of data packets at the same location is output. The primary warning label is used to indicate whether there are potential anomalies in the production process corresponding to the group of data packets at the same location.

[0008] Based on the cloud service detection algorithm, the abnormal data packets corresponding to the primary warning labels are analyzed for abnormal root causes. The resulting production warning root cause tracing results, which include the cause of the abnormality, the degree of impact, and the location information, are used to guide production adjustments and troubleshooting.

[0009] Secondly, embodiments of the present invention provide a waterproof membrane production monitoring and early warning system based on multimodal data fusion, comprising:

[0010] processor;

[0011] Storage device, on which computer programs are stored,

[0012] When the computer program is executed by the processor, the processor implements any of the aforementioned methods for monitoring and early warning of waterproof membrane production based on multimodal data fusion.

[0013] This invention provides a readable storage medium storing a program or instructions, which, when executed by a processor, implement the steps of the waterproof membrane production monitoring and early warning method based on multimodal data fusion.

[0014] This invention, through a hierarchical anomaly analysis and early warning architecture at the edge and cloud levels, achieves comprehensive, efficient, and accurate monitoring and early warning of the waterproof membrane production process. At the edge, a pre-established delay compensation model performs physical segment matching on the initial multi-mode monitoring dataset, resolving the issue of data acquisition time differences between different sensors and ensuring data accuracy and consistency. Correlation feature engineering extracts and fuses different types of features from the same-location data packets to generate a production process status feature set, which can more comprehensively and deeply reflect the production process status. The edge-level embedded fusion model performs multi-dimensional analysis on the production process status feature set and outputs primary early warning labels, enabling rapid identification of potential anomalies in the production process at the edge, reducing data transmission volume and response time, and improving real-time performance.

[0015] In the cloud, based on cloud service detection algorithms, anomaly root cause analysis is performed on the abnormal data packets corresponding to the primary warning tags, generating production warning root cause tracing results that include the cause, impact level, and location information of the anomaly. The computing power and rich knowledge base of the cloud enable in-depth root cause mining and impact range assessment, providing detailed and accurate guidance for production adjustments and troubleshooting. The collaborative working mode between the edge and cloud in this embodiment of the invention enables rapid anomaly detection and accurate tracing, improving the quality and efficiency of waterproof membrane production, and reducing production costs and failure risks. Attached Figure Description

[0016] Figure 1 The flowchart illustrates a method for monitoring and early warning of waterproof membrane production based on multimodal data fusion, as provided in an embodiment of the present invention.

[0017] Figure 2This is a schematic diagram of the basic structure of a waterproof membrane production monitoring and early warning system based on multimodal data fusion, provided in an embodiment of the present invention. Detailed Implementation

[0018] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0019] See Figure 1 As shown in the figure, this is a flowchart of a method for monitoring and early warning of waterproof membrane production based on multimodal data fusion provided by an embodiment of the present invention. This method can be applied to a monitoring and early warning system for waterproof membrane production based on multimodal data fusion. Figure 1 As shown, the method includes steps 110-140.

[0020] Step 110: Perform physical segment same-point matching processing on the initial multi-mode monitoring dataset using a pre-established delay compensation model to obtain a same-point matching dataset; the initial multi-mode monitoring dataset contains continuous monitoring data of each process link of the waterproof membrane production line.

[0021] In the monitoring and early warning scenario of waterproof membrane production, the waterproof membrane production line includes multiple process steps, such as raw material mixing, heating, molding, and cooling. In order to comprehensively monitor the production process, various types of sensors are deployed at various key locations on the production line. These sensors continuously collect data to form an initial multi-mode monitoring dataset. This dataset covers different types of monitoring data such as temperature, pressure, speed, and humidity, and the data for each process step is continuously recorded.

[0022] However, due to differences in data acquisition mechanisms and transmission methods among different sensors, data arrival times may vary, necessitating a pre-established delay compensation model. This delay compensation model is built upon long-term research and analysis of sensor characteristics and data transmission patterns. For example, temperature and pressure sensors installed in different locations may have different data acquisition cycles and transmission delays. The delay compensation model adjusts the data from each sensor in terms of time based on the sensor's location, type, and the temporal characteristics of historical data.

[0023] Specifically, the delay compensation model first labels the data in the initial multi-mode monitoring dataset, identifying the sensor and physical location corresponding to each data point. Then, according to pre-set compensation rules, it performs time calibration on the data from different sensors. For example, for temperature and pressure data at a certain location, if the acquisition and transmission of temperature data is delayed relative to pressure data, the delay compensation model will adjust the timestamp of the temperature data backward to align it with the pressure data in time. Through the above physical segment same-location matching process, a same-location matching dataset is finally obtained. In this dataset, different types of monitoring data at the same physical location are synchronized in time.

[0024] Step 120: Perform association feature engineering processing on each group of data packets in the same location matching dataset, extract different types of features reflecting the production process status from each group of data packets, and perform cross-association and fusion of the extracted different types of features to generate a production process status feature set corresponding to each group of data packets.

[0025] Step 121: Perform data parsing on each group of data packets at the same location, separate the monitoring data sequences corresponding to different monitoring types, and extract time domain features for the data sequences of each monitoring type.

[0026] After obtaining the same-location matching dataset, it contains multiple sets of same-location data packets. Taking one set of same-location data packets as an example, this data packet may contain various monitoring types of data, such as temperature, pressure, and speed, collected at a designated location on the production line. The first step is to parse this set of same-location data packets. The data parsing process involves separating the data of different monitoring types from the data packets. For example, by identifying the data labels or formats, temperature data, pressure data, and speed data are extracted separately, forming their own independent monitoring data sequences. Each monitoring data sequence is continuous data arranged in chronological order.

[0027] For each type of monitoring data sequence, time-domain features need to be extracted. Taking temperature data sequences as an example, time-domain features can be extracted from several aspects. First is the mean, which reflects the average temperature level over a period of time. By summing and averaging all data points in the temperature data sequence, the temperature mean for that time period can be obtained. This mean can help determine the overall temperature situation in the production process. Second is the variance, which measures the degree of dispersion of temperature data relative to the mean. A large variance indicates large temperature fluctuations, potentially indicating unstable production processes. When calculating variance, the difference between each data point and the mean is first calculated, then these differences are squared, summed, and averaged. Maximum and minimum values ​​are also important time-domain features. The maximum value represents the highest temperature reached during that time period, and the minimum value represents the lowest value. By analyzing the maximum and minimum values, the range of temperature fluctuations can be understood, and it can be determined whether there are abnormal conditions of excessively high or low temperatures.

[0028] In addition to the features mentioned above, rising and falling edge features can also be extracted from temperature data sequences. A rising edge indicates a rapid increase in temperature from a lower to a higher value, while a falling edge indicates the opposite. These features reflect temperature trends and are important for predicting anomalies in the production process. Similarly, similar methods are used to extract time-domain features from pressure and velocity data sequences. By extracting these time-domain features, a more comprehensive understanding of the changes in each type of monitoring data over time can be achieved.

[0029] Step 122: Extract spectral features from the data sequences of each monitoring type to obtain the energy distribution characteristics and characteristic frequency points of the data sequences in different frequency bands.

[0030] After extracting the time-domain features, it is also necessary to extract spectral features from the data sequences of each monitoring type. Spectral features can reveal the energy distribution of the data in different frequency bands and characteristic frequency points. This information is crucial for detecting periodic changes or abnormal fluctuations in the production process.

[0031] Taking temperature data sequences as an example, the process of spectral feature extraction is as follows. First, the temperature data sequences are preprocessed, including noise and trend removal. Noise may be caused by sensor accuracy issues or external interference; noise removal improves data quality. Trends refer to the long-term trends in the data sequence; removing trends makes the data more suitable for spectral analysis. Then, spectral analysis methods are used to process the preprocessed temperature data sequences. Spectral analysis methods can convert time-domain data into frequency-domain data, thereby obtaining the energy distribution of the data in different frequency bands. For example, lower frequency bands may reflect slow changes in the production process, such as the preheating or cooling process of equipment; while higher frequency bands may reflect rapid fluctuations in the production process, such as equipment vibration or pulse signals.

[0032] Spectral analysis allows us to obtain the energy distribution characteristics of temperature data sequences across different frequency bands. These characteristics can be represented by energy values; higher energy values ​​indicate stronger signals within that frequency band. Furthermore, we can identify characteristic frequency points—those with high energy values. These characteristic frequency points may be related to certain established phenomena in the production process, such as equipment resonant frequencies or periodic faults. The same method is used to extract spectral features from pressure and velocity data sequences. By extracting spectral features, we can gain a deeper understanding of data from various monitoring types along the frequency dimension, providing crucial clues for identifying potential problems in the production process.

[0033] Step 123: Extract the time-frequency domain features of the data sequence using the time-frequency domain joint analysis method to capture the joint variation characteristics of the data sequence in the time and frequency dimensions.

[0034] In addition to time-domain and spectral features, it is also necessary to extract time-frequency domain features of the data sequence through joint time-frequency analysis. This method can simultaneously consider the changes in the data in the time and frequency dimensions and capture the joint variation characteristics of the data sequence.

[0035] Taking temperature data sequences as an example, the specific implementation process of the time-frequency joint analysis method is as follows. First, a suitable time-frequency analysis method is selected, such as wavelet transform or short-time Fourier transform. These methods can decompose the data sequence in both time and frequency dimensions to obtain a time-frequency graph.

[0036] A time-frequency graph is a two-dimensional image where the horizontal axis represents time and the vertical axis represents frequency. The grayscale or color of the image represents the energy value at that time and frequency point. By observing a time-frequency graph, one can intuitively see how temperature data changes at different times and frequencies. For example, a sudden increase in energy value at a certain time and frequency range may indicate a brief high-frequency fluctuation at that moment, which could be related to an abnormal event in the production process.

[0037] After obtaining the time-frequency plot, it is necessary to extract time-frequency domain features. These features can include time-frequency energy distribution, time-frequency peak values, and time-frequency bandwidth. The time-frequency energy distribution reflects the energy distribution of the data in time and frequency. By statistically analyzing the energy value of each point in the time-frequency plot, the characteristics of the time-frequency energy distribution can be obtained. The time-frequency peak value represents the point with the largest energy value in the time-frequency plot, reflecting the main frequency components of the data and their occurrence time. The time-frequency bandwidth represents the frequency range where energy is concentrated in the time-frequency plot, reflecting the frequency stability of the data.

[0038] For pressure and velocity data sequences, a joint time-frequency domain analysis method will also be used to extract time-frequency domain features. By extracting time-frequency domain features, a more comprehensive understanding of the joint variation characteristics of the data sequence in the time and frequency dimensions can be obtained.

[0039] Step 124: Perform feature standardization processing on the extracted time-domain features, spectral features, and time-frequency features to obtain a standardized feature set.

[0040] After extracting time-domain features, spectral features, and time-frequency domain features, since these features may have different dimensions and numerical ranges, they need to be standardized to facilitate subsequent analysis and processing.

[0041] The purpose of feature standardization is to unify the numerical ranges of different features into a defined interval, making these features comparable. Taking time-domain, spectral, and time-frequency domain features of temperature as examples, these features may have different magnitudes and ranges. For instance, the mean of temperature may have a large numerical range, while the variance may be relatively small. First, each feature is processed separately. For each feature, its mean and standard deviation are calculated. The mean represents the average level of the feature, and the standard deviation represents the dispersion of the feature relative to the mean. Then, the features are standardized based on the mean and standard deviation. Specifically, for each feature value, the mean is subtracted first, and then the standard deviation is divided. Through this process, the feature values ​​are converted into standard values ​​with a mean of 0 and a standard deviation of 1. The same standardization method is used for all time-domain, spectral, and time-frequency domain features. After standardization, a standardized feature set is obtained. In this standardized feature set, different types of features are represented at the same scale.

[0042] Step 125: Construct a feature correlation matrix, calculate the correlation coefficients between different types of features in the standardized feature set, select feature combinations that meet the correlation conditions based on the correlation coefficients, perform feature fusion operation on the selected feature combinations, and generate a production process status feature set that reflects the production process status.

[0043] After obtaining the standardized feature set, a feature correlation matrix is ​​constructed. The feature correlation matrix is ​​used to describe the correlation between different types of features in the standardized feature set.

[0044] First, identify all feature types in the standardized feature set, such as time-domain features of temperature and spectral features of pressure. Then, for each pair of features of different types, calculate the correlation coefficient between them. The correlation coefficient measures the degree of linear association between two features, ranging from -1 to 1. A correlation coefficient of 1 indicates a perfect positive correlation, meaning an increase in one feature leads to an increase in the other; a correlation coefficient of -1 indicates a perfect negative correlation, meaning an increase in one feature leads to a decrease in the other; and a correlation coefficient of 0 indicates no linear association between the two features. Statistical analysis methods, such as the Pearson correlation coefficient method, can be used to calculate the correlation coefficient. By calculating the correlation coefficient for each pair of features in the standardized feature set, a matrix is ​​obtained, where each element represents the correlation coefficient between the corresponding feature pair. Based on the correlation coefficient, feature combinations that meet the correlation criteria are selected. The correlation criteria can be set according to the actual situation; for example, feature combinations with a correlation coefficient greater than a certain threshold are considered to have a strong correlation. For example, setting the threshold to 0.8 means that when the correlation coefficient between two features is greater than 0.8, they are considered to have a strong correlation. Finally, a feature fusion operation is performed on the selected feature combinations. The purpose of feature fusion is to combine highly correlated features to form a more representative feature. There are various methods for feature fusion, such as weighted averaging. For each feature in the selected feature combination, a weight is assigned according to its importance, and then these features are weighted and averaged to obtain a new feature.

[0045] By performing feature fusion on all selected feature combinations, a production process status feature set that reflects the production process status is finally generated, which can more comprehensively reflect the production process status.

[0046] Step 130: Input the production process status feature set into the edge-level embedded fusion model for multi-dimensional analysis, and output the primary warning label corresponding to each group of data packets at the same location. The primary warning label is used to indicate whether there are potential anomalies in the production process corresponding to the group of data packets at the same location.

[0047] Step 131: Input the production process status feature set into the feature input layer of the edge-level embedded fusion model, and perform dimensionality verification and format conversion on the feature set.

[0048] The edge-level embedded fusion model is specifically designed for monitoring and early warning scenarios in the production of waterproof membrane rolls. It can run on edge devices to achieve real-time data analysis and early warning. The production process status feature set contains various features extracted and fused from matching datasets at the same location. These features will be used as input data into the feature input layer of the edge-level embedded fusion model.

[0049] At the feature input layer, the first step is to perform dimensionality validation on the production process state feature set. Dimensionality validation ensures that the dimensions of the input feature set match the model's requirements. The model is designed with explicit specifications for the dimensions of the input data; for example, it may require the input data to have a set number of features and feature dimensions. If the dimensions of the input feature set do not meet the requirements, the model may not function correctly.

[0050] For example, edge-level embedded fusion models require the input feature set to have N features, each with an M-dimensional vector representation. During dimensionality verification, it checks whether the production process status feature set contains N features and whether each feature is an M-dimensional vector. If a dimensionality mismatch is found, appropriate processing is performed, such as adding or removing features, or adjusting the feature dimensions.

[0051] In addition to dimensionality verification, the feature set also undergoes format conversion. Different sensors and data acquisition systems may use different data formats, while edge-level embedded fusion models require a unified data format for processing. Therefore, at the feature input layer, the production process state feature set is converted into a format that the model can recognize and process. For example, the data type of the feature set is converted to a specified data type, or the arrangement order of the feature set is adjusted to meet the model's input requirements.

[0052] By performing dimensional verification and format conversion, we ensure that the production process status feature set can be successfully entered into the edge-level embedded fusion model for subsequent processing.

[0053] Step 132: The transformed production process state feature set is enhanced by the feature enhancement layer of the edge-level embedded fusion model, and the feature set is mapped to a high-dimensional feature space by a feature mapping algorithm.

[0054] After processing by the feature input layer, the transformed production process state feature set enters the feature enhancement layer of the edge-level embedded fusion model. The main function of the feature enhancement layer is to strengthen the feature set to improve its expressive power.

[0055] Feature mapping algorithms are the core algorithms of feature enhancement layers, mapping feature sets to a high-dimensional feature space. In low-dimensional space, the features of data may not be able to fully express the complex relationships between data, while in high-dimensional feature space, the relationships between features are more complex and richer, and can better reflect the diversity of production process states.

[0056] Let's illustrate the feature mapping process with an example. A feature in the production process state feature set is one-dimensional, representing only a single numerical value. A feature mapping algorithm maps this one-dimensional feature to a higher-dimensional space, such as two-dimensional or three-dimensional space. In this higher-dimensional space, the feature can be represented by a vector, where each dimension represents an aspect of the feature. This allows for a more comprehensive description of the feature and improves its expressive power. Feature mapping algorithms can be implemented using various methods, such as kernel function mapping. Kernel function mapping maps features from a low-dimensional space to a higher-dimensional space by selecting an appropriate kernel function. Kernel functions can be linear, polynomial, or Gaussian, etc., and different kernel functions are suitable for different data distributions and problem types. In the feature enhancement layer, feature mapping is performed on each feature in the production process state feature set. By mapping the feature set to a higher-dimensional feature space, the expressive power of the feature set is enhanced.

[0057] Step 133: Utilize the multi-scale feature extraction layer of the edge-level embedded fusion model to perform multi-scale analysis on the feature set in the high-dimensional feature space. Extract feature information at different levels through feature extraction windows of different sizes to generate a multi-scale feature set.

[0058] Step 1331: In the multi-scale feature extraction layer, set up multiple feature extraction windows with different sizes. The size of each window increases according to a preset ratio to form a sequence of extraction windows covering different feature scales.

[0059] The purpose of a multi-scale feature extraction layer is to analyze the feature set in a high-dimensional feature space at different scales to obtain feature information at different levels. To achieve this, multiple feature extraction windows of different sizes are set in the multi-scale feature extraction layer, and the sizes of these feature extraction windows increase according to a preset ratio. For example, the first feature extraction window has the smallest size, and the size of subsequent windows increases sequentially according to a certain ratio, thus forming a sequence of extraction windows covering different feature scales. Feature extraction windows of different sizes can capture feature information at different scales. Smaller feature extraction windows are suitable for extracting local, subtle features, such as short-term fluctuations or local anomalies in a production process. Larger feature extraction windows can extract more macroscopic, global features, such as long-term trends or the overall state of a production process. By setting feature extraction windows of different sizes, the feature set in the high-dimensional feature space can be analyzed from multiple perspectives to obtain more comprehensive feature information.

[0060] Step 1332: Input the feature set in the high-dimensional feature space into each feature extraction window in sequence. Each window performs local feature extraction on the input feature set to generate a local feature map corresponding to the window size.

[0061] After setting up the feature extraction window sequence, the feature set in the high-dimensional feature space is sequentially input into each feature extraction window. Each feature extraction window performs local feature extraction on the input feature set. Taking a feature extraction window of a set size as an example, when the feature set is input into the window, the window slides across the feature set, covering a local region of the feature set with each slide. For each covered local region, the window extracts the feature information of that region, generating a local feature vector. As the window slides across the feature set, multiple local feature vectors are generated. Arranging these local feature vectors according to the sliding order of the window forms a local feature map corresponding to the window size. Feature extraction windows of different sizes generate different local feature maps. Smaller windows generate local feature maps that focus more on local details, while larger windows generate local feature maps that better reflect global information. In this way, each feature extraction window can extract feature information from different scales and generate corresponding local feature maps.

[0062] Step 1333: Perform feature dimensionality reduction on each local feature map, and standardize the features of each local feature map after dimensionality reduction so that the local feature maps extracted from different windows have a uniform numerical range.

[0063] After generating local feature maps, since the dimensions of the local feature maps may be high, feature dimensionality reduction processing needs to be performed on each local feature map in order to reduce computational complexity and data redundancy.

[0064] There are various methods for feature reduction, such as Principal Component Analysis (PCA). PCA projects high-dimensional data into a low-dimensional space by identifying the principal components of the data, while preserving the main information. When performing PCA on a local feature map, the covariance matrix of the local feature map is calculated, and then the eigenvectors and eigenvalues ​​of the covariance matrix are found. Based on the magnitude of the eigenvalues, the eigenvectors corresponding to the k largest eigenvalues ​​are selected as principal components. The local feature map is then projected onto these k principal components to obtain the dimensionality-reduced local feature map.

[0065] The dimensionality of the reduced local feature maps is lower, but the main information of the original local feature maps is still retained. However, the numerical range of local feature maps extracted from different windows may be different. In order to facilitate subsequent feature stitching and analysis, feature standardization is required for each local feature map after dimensionality reduction.

[0066] The purpose of feature standardization is to ensure that local feature maps extracted from different windows have a uniform numerical range. Specifically, for each dimensionality-reduced local feature map, its mean and standard deviation are calculated. Then, the mean is subtracted from each feature value in the local feature map, and the result is divided by the standard deviation. In this way, the numerical range of the local feature maps is unified to a set interval, such as [0, 1] or [-1, 1]. Through feature dimensionality reduction and feature standardization, local feature maps extracted from different windows exhibit consistency in both dimension and numerical range.

[0067] Step 1334: The standardized local feature maps are stitched together in ascending order of window size to form a multi-scale feature set containing feature information at different levels. The feature information at each level corresponds to the local features extracted by the window of the set size.

[0068] After feature reduction and standardization, the standardized local feature maps are concatenated in ascending order of window size. Feature concatenation involves sequentially connecting local feature maps extracted from windows of different sizes to form a larger feature set. Since each local feature map represents feature information at different scales, feature concatenation integrates these different scales to form a multi-scale feature set containing feature information at different levels. For example, the local feature map extracted by the smallest window contains the most local and subtle feature information and is placed at the beginning of the multi-scale feature set. As the window size increases, the subsequent local feature maps contain increasingly macroscopic and global feature information. These local feature maps are then concatenated sequentially to form a hierarchical multi-scale feature set. Each level of feature information corresponds to the local features extracted by a specific window size. In this way, the multi-scale feature set can comprehensively reflect the feature information of the production process at different scales.

[0069] The multi-scale feature set formed by feature splicing provides rich and comprehensive feature information for subsequent anomaly pattern recognition, which can more accurately detect anomalies in the production process.

[0070] Step 134: Input the multi-scale feature set into the anomaly pattern recognition layer of the edge-level embedded fusion model. The anomaly pattern recognition layer contains multiple parallel anomaly detection sub-modules. Each anomaly detection sub-module uses a different anomaly detection algorithm to process the multi-scale feature set and outputs its own anomaly detection result.

[0071] The generated multi-scale feature set is input into the anomaly pattern recognition layer of the edge-level embedded fusion model. The main task of the anomaly pattern recognition layer is to identify whether anomalies exist in the production process. To improve the accuracy and reliability of the recognition, this layer contains multiple parallel anomaly detection sub-modules. Each anomaly detection sub-module uses a different anomaly detection algorithm to process the multi-scale feature set. For example, one anomaly detection sub-module might use a statistical analysis-based anomaly detection algorithm, another might use a machine learning-based anomaly detection algorithm, and yet another might use a rule-based anomaly detection algorithm.

[0072] Taking the anomaly detection submodule based on statistical analysis as an example, this submodule performs statistical analysis on each feature in the multi-scale feature set. First, it calculates the mean and standard deviation of each feature, and then determines the range of normal data based on the mean and standard deviation. If the value of a feature exceeds the range of normal data, it is considered that there may be an anomaly in the production process corresponding to that feature.

[0073] The machine learning-based anomaly detection submodule processes multi-scale feature sets using a pre-trained machine learning model. For example, using a neural network model, the multi-scale feature set is taken as input, and the model outputs an anomaly probability value based on the input data. If the anomaly probability value exceeds a certain threshold, an anomaly is considered to exist.

[0074] The rule-based anomaly detection submodule judges the multi-scale feature set according to pre-defined rules. For example, if the rule is set that an anomaly exists when both temperature and pressure features exceed a certain threshold, the submodule will check whether the temperature and pressure features in the multi-scale feature set meet the rule. If they do, the anomaly detection result will be output.

[0075] Each anomaly detection submodule independently processes the multi-scale feature set and outputs its own anomaly detection results, which include anomaly probability values ​​and anomaly type identifiers. By using multiple parallel anomaly detection submodules, the multi-scale feature set can be analyzed from different perspectives, improving the accuracy of anomaly identification.

[0076] Step 135: Perform fusion decision on the anomaly detection results output by each anomaly detection submodule, comprehensively judge whether there is an anomaly pattern according to the preset decision rules, and generate and output the primary warning tag corresponding to each group of data packets at the same location.

[0077] Step 1351: Collect the anomaly detection results output by each anomaly detection submodule. Each anomaly detection result includes an anomaly probability value and an anomaly type identifier.

[0078] In the anomaly pattern recognition layer, after each anomaly detection submodule outputs its anomaly detection results, these results need to be collected. Each anomaly detection result includes an anomaly probability value and an anomaly type identifier. The anomaly probability value indicates the likelihood that the submodule believes there is an anomaly in the current production process. Different anomaly detection submodules may use different methods to calculate the anomaly probability value. For example, a statistical analysis-based submodule may calculate the anomaly probability value based on the degree to which the data deviates from the normal range, while a machine learning-based submodule may directly obtain the anomaly probability value from the model's output. The anomaly type identifier indicates the possible anomaly type. For example, anomaly type identifiers may include temperature anomalies, pressure anomalies, speed anomalies, etc. Different anomaly detection submodules may have different definitions and identifiers for anomaly types, but they will all provide a clear anomaly type identifier.

[0079] Step 1352: Normalize each abnormal probability value and map each abnormal probability value to the same probability interval.

[0080] Since the anomaly probability values ​​output by different anomaly detection submodules may have different ranges, these anomaly probability values ​​need to be normalized to facilitate fusion decision-making. The purpose of normalization is to map each anomaly probability value to the same probability interval, such as [0, 1]. This makes the anomaly probability values ​​from different submodules comparable and allows for calculation and analysis on the same scale. A linear transformation can be used for normalization. For each anomaly probability value, first find the minimum and maximum values ​​among all anomaly probability values. Then, subtract the minimum value from the normalized anomaly probability value and divide by the difference between the maximum and minimum values ​​to obtain the normalized anomaly probability value. Through normalization, each anomaly probability value is represented within the same probability interval.

[0081] Step 1353: Based on the preset anomaly detection submodule weight allocation scheme, assign a corresponding decision weight to each anomaly detection submodule. The decision weight is adjusted according to the accuracy and recall of the anomaly detection submodule in historical detection tasks.

[0082] When making fusion decisions, each anomaly detection submodule needs to be assigned a corresponding decision weight. The preset weight allocation scheme for anomaly detection submodules is adjusted based on their accuracy and recall in historical detection tasks. Accuracy refers to the proportion of anomalies correctly detected by an anomaly detection submodule, while recall refers to the proportion of anomalies detected by an anomaly detection submodule relative to actual anomalies. An anomaly detection submodule with high accuracy and recall in historical detection tasks indicates strong detection capabilities and should be assigned a higher decision weight. For example, if a machine learning-based anomaly detection submodule achieved an accuracy and recall of over 80% in historical detection tasks, while another statistical analysis-based submodule had relatively lower accuracy and recall, then the machine learning-based submodule would be assigned a higher decision weight during weight allocation. By adjusting the decision weights based on the accuracy and recall of historical detection tasks, the performance differences among the anomaly detection submodules can be fully considered, improving the accuracy of the fusion decision.

[0083] Step 1354: Weight the normalized outlier probability values ​​according to the decision weights to obtain the comprehensive outlier probability value.

[0084] After normalizing the anomaly probability values ​​and assigning decision weights, the normalized anomaly probability values ​​are weighted and summed according to the decision weights. The weighted summation process involves multiplying each normalized anomaly probability value by its corresponding decision weight, and then adding all the products to obtain the comprehensive anomaly probability value. For example, if there are three anomaly detection submodules with normalized anomaly probability values ​​P1, P2, and P3, and corresponding decision weights W1, W2, and W3, then the comprehensive anomaly probability value is P = P1*W1 + P2*W2 + P3*W3. The comprehensive anomaly probability value more comprehensively reflects the likelihood of anomalies in the current production process, taking into account the detection results and performance differences of multiple anomaly detection submodules.

[0085] Step 1355: Compare the comprehensive anomaly probability value with the preset anomaly probability threshold. If the comprehensive anomaly probability value is greater than or equal to the anomaly probability threshold, it is determined that an anomaly mode exists. Based on the anomaly type identifier output by each submodule, select the anomaly type with the highest frequency as the current anomaly type and generate a primary warning label containing the anomaly type and the comprehensive anomaly probability value. If the comprehensive anomaly probability value is less than the anomaly probability threshold, it is determined that no anomaly mode exists and a primary warning label for normal status is generated.

[0086] After obtaining the comprehensive anomaly probability value, it is compared with a preset anomaly probability threshold. The preset threshold is set based on actual conditions and experience to determine the existence of anomaly patterns. If the comprehensive anomaly probability value is greater than or equal to the threshold, an anomaly pattern is determined to exist. At this point, based on the anomaly type identifiers output by each submodule, the anomaly type with the highest frequency is selected as the current anomaly type. For example, if "temperature anomaly" appears most frequently among the anomaly type identifiers output by multiple anomaly detection submodules, then "temperature anomaly" is selected as the current anomaly type. A primary warning label containing the anomaly type and the comprehensive anomaly probability value is generated. This label clearly indicates the type of anomaly and the probability of the anomaly in the current production process, providing clear information for subsequent processing. If the comprehensive anomaly probability value is less than the threshold, no anomaly pattern is determined to exist, and a primary warning label for normal status is generated. The primary warning label for normal status indicates that the current production process is operating normally, and no obvious anomalies have been detected. In this way, it is possible to accurately determine whether there are potential anomalies in the production process.

[0087] Step 140: Based on the cloud service detection algorithm, perform anomaly root cause analysis on the abnormal data packets corresponding to the primary warning tags to generate production warning root cause tracing results containing the cause of the anomaly, the degree of impact, and the location information. The production warning root cause tracing results are used to guide production adjustments and troubleshooting.

[0088] Step 141: Extract abnormal location data packets corresponding to the primary warning labels from the same location matching dataset, perform abnormal scene feature modeling on the abnormal location data packets, and construct an abnormal event context description model that includes the time of occurrence of the abnormality, the duration of the abnormality, and the associated process links.

[0089] Step 1411: Analyze the timestamp sequence of abnormal data packets at the same location to determine the starting time of the first occurrence of abnormal data and the duration of the abnormal state, and calculate the deviation ratio between the duration of the abnormality and the standard process duration.

[0090] Once the primary warning label is obtained, abnormal location data packets corresponding to the primary warning label are extracted from the same location matching dataset. These abnormal location data packets contain relevant monitoring data when anomalies occur during the production process.

[0091] First, the timestamp sequence of the abnormal data packets at the same location is analyzed. The timestamp sequence records the acquisition time of each data point. By analyzing the timestamp sequence, the starting time of the first occurrence of abnormal data can be determined. For example, in temperature monitoring data, when the temperature value suddenly exceeds the normal range, the corresponding timestamp is the start time of the anomaly.

[0092] Next, determine the duration of the abnormal state. Starting from the initial moment of the abnormality, continuously monitor data changes until the data returns to the normal range or meets the set termination conditions; this period is the duration of the abnormal state. Calculate the deviation ratio between the abnormal duration and the standard process duration. The standard process duration refers to the time required for this process step under normal production conditions. Calculate the deviation ratio by comparing the abnormal duration with the standard process duration. For example, if the abnormal duration is T1 and the standard process duration is T0, then the deviation ratio is (T1-T0) / T0. This deviation ratio reflects the degree of deviation of the abnormal duration from the normal process standard.

[0093] Step 1412: Extract the trend features of data changes in each monitoring dimension during the abnormal period and construct the abnormal time series evolution feature vector.

[0094] After determining the timeframe of the anomaly, the changing trend characteristics of the data from each monitoring dimension within the anomaly period are extracted. For example, during a period of abnormally rising temperature, changes in data such as pressure and velocity are simultaneously monitored. For each monitoring dimension, its changing trend within the anomaly period is analyzed. The slope of the data can be calculated to determine whether the data is rising, falling, or remaining stable. For example, for temperature data, a positive slope during the anomaly period indicates that the temperature is rising; a negative slope indicates that the temperature is falling. In addition to the slope, the rate of change and acceleration of the data can also be analyzed. The rate of change represents the amount of change in data per unit time, while acceleration represents the change in the rate of change. By analyzing these characteristics, a more detailed understanding of the data's changing trend can be obtained. The changing trend characteristics of the data from each monitoring dimension are combined to construct an anomaly time-series evolution feature vector. This feature vector reflects the dynamic changes of the data from each monitoring dimension within the anomaly period, providing important time-series information for the characteristic modeling of anomaly scenarios.

[0095] Step 1413: Based on the matching relationship of the same location, trace the physical roll material location and the production process link corresponding to the abnormal data, associate the upstream and downstream process information of the production process link, and generate a process link association map.

[0096] Based on the matching relationship between points, the physical location of the roll material and the corresponding production process step can be traced back to the source of the abnormal data. Since each sensor corresponds to a set physical location and production process step during the data acquisition process, the location of the roll material and the process step corresponding to the abnormal data can be accurately determined through the matching relationship between points.

[0097] For example, a temperature sensor installed in the forming stage of a production line can pinpoint an anomaly in the temperature data it collects, indicating that the anomaly occurs on the roll material at a designated location within the forming stage. This is then linked to upstream and downstream processes within the production process. Understanding the position and role of this process within the overall production flow, as well as its relationships with upstream and downstream processes, is crucial. For instance, the upstream process of the forming stage might be raw material mixing, while the downstream process might be cooling. By integrating the physical roll material location corresponding to the anomaly data, the associated production stage, and the upstream and downstream process information, a process link mapping diagram is generated. This diagram clearly shows the relationship between the process stage containing the anomaly data and other process stages.

[0098] Step 1414: Combine the abnormal time-series evolution feature vector, the process link correlation map, and the state switching records in the equipment operation log to construct a multi-dimensional abnormal scenario description matrix.

[0099] By combining the anomaly time-series evolution feature vector, the process link correlation graph, and the state transition records in the equipment operation log, a multi-dimensional anomaly scenario description matrix is ​​constructed. The state transition records in the equipment operation log provide information about the equipment's operating status before and after an anomaly occurs, such as equipment startup, shutdown, and fault status changes. For example, during a period of abnormal temperature rise, the equipment operation log may record startup or fault information for a heating device. By using the anomaly time-series evolution feature vector, the process link correlation graph, and the state transition records in the equipment operation log as different dimensions of the matrix, a multi-dimensional anomaly scenario description matrix is ​​constructed. This matrix can comprehensively describe the characteristics of anomaly scenarios from multiple dimensions, including time, space, and equipment status.

[0100] Step 1415: Compress the information of the multi-dimensional abnormal scene description matrix using the scene feature dimensionality reduction algorithm to generate a structured abnormal event context description model. The abnormal event context description model includes time dimension features, spatial process correlation features, and multi-source data collaborative abnormal features.

[0101] To reduce data redundancy and complexity, a scene feature dimensionality reduction algorithm is used to compress information from the multi-dimensional abnormal scene description matrix. Scene feature dimensionality reduction algorithms can employ methods such as Principal Component Analysis (PCA) or Linear Discriminant Analysis (LDA).

[0102] Taking principal component analysis (PCA) as an example, PCA calculates the covariance matrix of the multi-dimensional anomaly scene description matrix, and then finds the eigenvectors and eigenvalues ​​of the covariance matrix. Based on the magnitude of the eigenvalues, the eigenvectors corresponding to the k largest eigenvalues ​​are selected as principal components. The multi-dimensional anomaly scene description matrix is ​​then projected onto these k principal components to obtain a dimensionality-reduced matrix. The dimensionality-reduced matrix retains the main information of the multi-dimensional anomaly scene description matrix while reducing the dimensionality of the data. The dimensionality-reduced matrix is ​​then structured to generate a structured anomaly event context description model.

[0103] The anomaly event context description model includes temporal features, spatial process correlation features, and multi-source data collaborative anomaly features. Temporal features reflect the timing of the anomaly, such as its onset time and duration. Spatial process correlation features demonstrate the physical location and technological stage of the anomaly data, as well as its relationship with upstream and downstream processes. Multi-source data collaborative anomaly features reflect the coordinated changes in data from various monitoring dimensions during the anomaly's occurrence. By generating this anomaly event context description model, the contextual environment of anomaly events can be described more concisely and accurately.

[0104] Step 142: Based on the abnormal event context description model, call the root cause knowledge graph of the cloud service detection algorithm. The root cause knowledge graph contains the causal relationship network, propagation path rules and influence range characteristics of various known fault modes in the waterproof membrane production process.

[0105] After constructing the anomaly event context description model, the root cause knowledge graph of the cloud service detection algorithm is invoked based on this model. The root cause knowledge graph is a knowledge base containing causal relationship networks, propagation path rules, and impact range characteristics of various known failure modes in the waterproof membrane production process. The causal relationship network describes the causal connections between different failure modes. For example, excessively high temperatures may lead to increased pressure, thereby affecting the forming quality of the membrane. Propagation path rules specify the way and path of failure propagation in the production system. For example, a failure in one piece of equipment may affect downstream processes through material transfer or signal transmission. Impact range characteristics describe the potential scope and degree of impact after a failure occurs. For example, one failure may only affect a local process, while another failure may affect the entire production line. By invoking the root cause knowledge graph, comprehensive knowledge support can be provided for anomaly root cause analysis. Matching the anomaly event context description model with the root cause knowledge graph enables rapid identification of possible failure causes and propagation paths.

[0106] Step 143: Through a multi-dimensional association reasoning mechanism, the abnormal scene features are deeply matched with the root cause knowledge graph to identify the initial mapping node set of the abnormal event in the root cause knowledge graph, and the cause search is carried out along the causal relationship network according to the preset propagation rules.

[0107] Step 1431: The multi-dimensional association reasoning mechanism includes a semantic matching layer, a structural matching layer, and a constraint matching layer. The semantic matching layer is used to calculate the semantic similarity between abnormal scene features and knowledge graph nodes. The structural matching layer is used to analyze the consistency between the relationship between features and the topological structure of the graph. The constraint matching layer is used to verify the conformity between the range of process parameters and the rule triggering conditions. The initial mapping node set is determined by the weighted fusion of the corresponding matching results.

[0108] The multi-dimensional associative reasoning mechanism consists of a semantic matching layer, a structural matching layer, and a constraint matching layer. These three layers match the abnormal scene features with the root cause knowledge graph from different perspectives to improve the accuracy of the matching.

[0109] The semantic matching layer calculates the semantic similarity between anomalous scene features and knowledge graph nodes. For example, the semantic similarity between "abnormal temperature rise" described in the anomalous scene and the "overheating fault" node in the root cause knowledge graph. Semantic similarity can be calculated using methods such as word vector models, converting the textual descriptions of the anomalous scene features and knowledge graph nodes into vector representations, and then calculating the similarity between the vectors.

[0110] The structure matching layer analyzes the consistency between the relationships between features in anomaly scenarios and the knowledge graph topology. For example, in anomaly scenarios, there is a correlation between increased temperature and increased pressure; a corresponding causal relationship needs to be found in the knowledge graph. By comparing the structure of the anomaly scenario with the topology of the knowledge graph, it determines whether they are consistent.

[0111] The constraint matching layer verifies the compliance of process parameter ranges with rule triggering conditions. For example, it checks whether process parameters such as temperature and pressure at the time of an anomaly meet the triggering conditions of relevant rules in the root cause knowledge graph. If the process parameters exceed the allowed range of the rule, the rule is considered inapplicable.

[0112] By weighted fusion of the matching results from these three layers, the initial set of mapping nodes for the abnormal event in the root cause knowledge graph is determined. The weighted fusion process assigns corresponding weights based on the importance of the matching results from each layer, and then comprehensively considers the matching results from each layer to obtain the final set of initial mapping nodes.

[0113] Step 1432: Using a bidirectional breadth-first search strategy, starting from the initial mapping node, perform a forward search along the reverse edges of the causal relationship network, and perform a backward search starting from the preset root cause candidate set. When the two search directions meet at the intermediate node, a complete causal path is formed.

[0114] After determining the initial set of mapped nodes, a bidirectional breadth-first search strategy is employed for causal search. This strategy combines forward and backward search, improving search efficiency.

[0115] Starting from the initial mapping node, a forward search is performed along the reverse edges of the causal relationship network. Reverse edges represent the opposite direction of the causal relationship; through forward search, potential causes of abnormal events can be identified. For example, starting from the "overheating fault" node, the search along the reverse edges can identify possible causes of overheating, such as heating equipment malfunction or raw material problems. Simultaneously, a backward search is performed from a pre-defined root cause candidate set. This set contains possible root causes, such as common equipment malfunctions or incorrect process parameter settings. From these root cause candidate sets, the search is performed along the forward edges of the causal relationship network to identify potential abnormal events. When the two search directions meet at an intermediate node, a complete causal path is formed. This path clearly demonstrates the causal chain from potential root causes to abnormal events, providing crucial clues for identifying the root cause of the anomaly.

[0116] Step 1433: During the search process, the pruning rules in the propagation rule base are dynamically applied to prune propagation paths that do not meet the current process conditions based on the process constraint parameters. When the process parameters of the target node in the path exceed the allowed range of the rules, the path search is terminated.

[0117] During the cause-finding search process, pruning rules from the propagation rule base are dynamically applied. This base contains a series of pruning rules that filter propagation paths based on current process conditions. For example, if the process parameters of a target node in a propagation path exceed the allowed range (e.g., excessively high temperature or excessively low pressure), the search for that path is terminated. This pruning operation reduces unnecessary searches and improves search efficiency. Specifically, when each node is searched, its corresponding process parameters are checked to ensure they conform to the rules in the propagation rule base. If they do not conform, the path is considered invalid, and the search for that path is discontinued. By dynamically applying pruning rules, impossible propagation paths can be quickly eliminated, narrowing the search scope and improving the accuracy and efficiency of the cause-finding search.

[0118] Step 1434: Perform a comprehensive evaluation of the path length, number of nodes, and rule compliance of all complete causal paths generated during the search process. Retain a set number of paths before the evaluation value as potential root cause propagation paths. The starting nodes in the potential root cause propagation paths constitute the root cause hypothesis set.

[0119] All complete causal paths generated during the search process are comprehensively evaluated. Evaluation metrics include path length, number of nodes, and rule compliance. Path length refers to the number of edges in the causal path; a shorter path indicates a more direct and likely causal relationship from the root cause to the anomalous event. Number of nodes refers to the number of nodes in the causal path; fewer nodes indicate a simpler path and a higher probability of it being the true root cause path. Rule compliance refers to whether each node and edge in the causal path conforms to the rules in the propagation rule base. A high rule compliance score indicates that all nodes and edges in the path strictly conform to the rules; a low score indicates that some nodes and edges do not conform to the rules. Considering path length, number of nodes, and rule compliance, an evaluation value is calculated for each complete causal path. The evaluation value can be calculated using a weighted summation method, assigning appropriate weights to each metric based on its importance, and then summing the values ​​of each metric. A predetermined number of paths with the highest evaluation values ​​are retained as potential root cause propagation paths. For example, if the goal is to retain the top 5 paths, then the 5 paths with the highest evaluation values ​​are selected as potential root cause propagation paths. The starting nodes in the potential root cause propagation path constitute the root cause hypothesis set. The root cause candidates in this set are more targeted and reliable, providing a basis for subsequent simulation and deduction of the scope of influence.

[0120] Step 144: During the cause search process, the process constraints and equipment operating status of the current production environment are introduced as reasoning pruning factors to gradually narrow down the candidate range of potential root causes and form a set of root cause hypotheses.

[0121] In the cause-finding search process, in addition to applying pruning rules from the propagation rule base, the process constraints and equipment operating status of the current production environment are also introduced as inference pruning factors. The process constraints of the current production environment include the allowable ranges of process parameters such as temperature, pressure, and speed. For example, in a certain process step, the temperature must be within a certain range for normal production. If a root cause found would cause the temperature to exceed this range, then the root cause is considered to not meet the process constraints and is eliminated. Equipment operating status is also an important inference pruning factor. For example, if a piece of equipment in the current production environment is under maintenance, then the cause of failure related to that equipment can likely be eliminated because the equipment will not affect the production process during maintenance. By introducing these inference pruning factors, the candidate range of potential root causes is gradually narrowed. During the search process, each potential root cause is continuously checked to see if it meets the process constraints and equipment operating status; if not, the root cause is eliminated. Finally, a root cause hypothesis set is formed, and the root cause candidates in this set are more targeted and reliable.

[0122] Step 145: Perform an impact range simulation for each candidate root cause in the root cause hypothesis set. Based on the propagation path rules in the root cause knowledge graph and the current production system topology, predict the downstream anomaly chain and affected process units caused by the candidate root cause, and generate production early warning root cause tracing results that include root cause probability ranking, impact path visualization, and key affected area marking.

[0123] Step 1451: Construct a directed graph model of the production system topology. In the directed graph model, nodes represent process units and key equipment, and directed edges represent material transmission directions and signal control relationships. Edge attributes include transmission delay time and capacity limit parameters.

[0124] To simulate and extrapolate the impact range of each candidate root cause in the root cause hypothesis set, a directed graph model of the production system topology is first constructed. In this directed graph model, nodes represent process units and key equipment. For example, mixing equipment, heating equipment, and molding equipment in a production line can all be represented by nodes. Each node has its defined attributes, such as equipment type and performance parameters. Directed edges represent material transport directions and signal control relationships. The material transport direction describes the flow direction of materials between different process units; for example, raw materials are transported from mixing equipment to heating equipment, and then from heating equipment to molding equipment. Signal control relationships represent the transmission of control signals between equipment; for example, the start or stop signal of a certain equipment can be transmitted to other related equipment through signal control relationships. Edge attributes include transmission delay time and capacity limit parameters. Transmission delay time represents the time required for material or signal to be transmitted between two nodes, and capacity limit parameters represent the transmission capacity limit between two nodes. For example, the maximum amount of material that can be transmitted in a certain transmission pipeline is a capacity limit parameter. By constructing a directed graph model of the production system topology, the structure and operation mechanism of the production system can be accurately described.

[0125] Step 1452: For each candidate root cause, extract the standard propagation path template corresponding to the candidate root cause from the root cause knowledge graph, instantiate the path by combining it with the current production system topology graph, replace the abstract nodes in the template with actual production units, and calculate the actual propagation distance and time.

[0126] For each candidate root cause in the root cause hypothesis set, a standard propagation path template corresponding to that candidate root cause is extracted from the root cause knowledge graph. The standard propagation path template describes the possible anomaly propagation path triggered by the candidate root cause; it is an abstract path model containing a series of nodes and edges. The path is instantiated using the current production system topology. The abstract nodes in the standard propagation path template are replaced with actual production units. For example, "heating equipment" in the template is replaced with specific heating equipment in the production line. During the node replacement process, the connection relationships and attributes between nodes need to be determined based on the production system topology. For example, the material transfer direction and signal control relationship between the actual heating equipment and other equipment, as well as the transmission delay time and capacity limitation parameters, need to be determined. The actual propagation distance and time are calculated. The actual propagation distance refers to the actual physical distance the anomaly travels in the production system, while the actual propagation time refers to the time required for the anomaly to propagate from the root cause node to other nodes. When calculating the actual propagation distance and time, the transmission delay time and capacity limitation parameters in the edge attributes need to be considered. For example, if the material's transmission speed in a certain transmission pipeline is slow, the anomaly propagation time will increase accordingly. By instantiating the path and calculating the actual propagation distance and time, the specific propagation path and propagation time of each candidate root cause in the current production system can be obtained.

[0127] Step 1453: Use a discrete event simulation algorithm to perform influence range deduction, inject candidate root causes as initial events into the simulation system, and trigger the abnormal state transition of downstream units in sequence according to the propagation path rules and topology graph relationships. Record the abnormal start time, duration and abnormality level of each unit.

[0128] Discrete event simulation algorithms are used to extrapolate the scope of influence. Discrete event simulation algorithms are event-driven simulation methods that predict system behavior by simulating the occurrence and propagation of events within the system. Candidate root causes are injected into the simulation system as initial events. For example, if a candidate root cause is a malfunction of a heating device, then that device malfunction event is injected into the simulation system as the initial event. According to propagation path rules and topology relationships, abnormal state transitions of downstream units are triggered sequentially. For example, when a heating device malfunctions, according to propagation path rules and topology relationships, abnormal states of connected downstream devices are triggered. For instance, a heating device malfunction may cause abnormal material temperatures, thereby affecting the normal operation of the molding equipment and triggering abnormal state transitions in the molding equipment.

[0129] During the simulation, the start time, duration, and severity level of the anomaly for each unit are recorded. The start time refers to the time when the unit begins to exhibit an abnormal state, the duration refers to the duration of the abnormal state, and the severity level indicates the degree of severity of the anomaly. For example, for temperature anomalies, the severity level can be determined based on the degree to which the temperature exceeds the normal range. By recording this information, the propagation process and impact of the anomaly within the production system can be described in detail.

[0130] Step 1454: Introduce random disturbance factors during the simulation process to simulate uncertainties in actual production, dynamically adjust the propagation probability and impact intensity, perform Monte Carlo simulations for each candidate root cause a certain number of predictions, and statistically analyze the frequency distribution and average anomaly level of each process unit affected.

[0131] During the simulation, a random disturbance factor is introduced to simulate uncertainties in actual production. In actual production, many uncertainties exist, such as sudden equipment failures and changes in the external environment. These factors can affect the propagation and severity of anomalies.

[0132] The random disturbance factor can be a random variable that dynamically adjusts the propagation probability and the intensity of the impact. For example, in a certain propagation path, the original propagation probability might be 80%, but after introducing the random disturbance factor, the propagation probability might fluctuate within a certain range, such as 70%-90%. Each candidate root cause undergoes a prediction number of Monte Carlo simulations. Monte Carlo simulation is a method for estimating system behavior through random sampling. Through multiple simulations, the propagation and impact of anomalies under different conditions can be obtained. The frequency distribution and average anomaly level of each process unit affected are statistically analyzed. The frequency distribution represents the proportion of times each process unit is affected in multiple simulations out of the total number of simulations, while the average anomaly level is the average of the anomaly levels of each process unit across multiple simulations. By statistically analyzing this information, a more comprehensive understanding of the probability and extent to which each process unit is affected by anomalies can be obtained.

[0133] Step 1455: Draw a heat map of the impact of the anomaly based on the simulation results, and overlay the probability of each process unit being affected and the time axis of the anomaly propagation to form a spatiotemporal integrated visualization result of the impact range.

[0134] An anomaly impact heatmap is generated based on the simulation results. A heatmap is a visualization method that uses color to represent data magnitude; in an anomaly impact heatmap, different colors represent different degrees of anomaly impact. For example, the darker the color, the more severe the anomaly impact. The impact probability of each process unit is overlaid with the anomaly propagation timeline. The impact probability indicates the likelihood of each process unit being affected by the anomaly in the simulation, while the anomaly propagation timeline indicates the time required for the anomaly to propagate from its root cause to each process unit. By overlaying the impact probability and the anomaly propagation timeline, a spatiotemporally integrated visualization of the impact range can be formed. This visualization clearly shows the propagation of the anomaly in time and space, providing an intuitive way to present the root cause tracing results for production early warnings. Finally, a production early warning root cause tracing result is generated, including root cause probability ranking, impact path visualization, and key affected area marking. Root cause probability ranking helps production personnel quickly locate possible root causes, impact path visualization shows the propagation path of the anomaly, and key affected area marking clarifies the areas most severely affected by the anomaly.

[0135] As a non-limiting embodiment, it also includes the step of introducing interpretable AI technology to carry out iterative optimization for the anomaly root cause analysis process in the cloud service detection algorithm:

[0136] Step 211: Collect the root cause tracing results of production early warnings generated in the historical production process and the actual production failure investigation results to construct a root cause analysis interpretive assessment dataset.

[0137] To iteratively optimize the anomaly root cause analysis process in the cloud service detection algorithm, it is necessary to collect the root cause tracing results of production warnings and the results of actual production fault investigations generated during historical production processes. In past production processes, each time an anomaly occurred, a root cause tracing result for a production warning was generated. These results included possible causes of the anomaly, the degree of impact, and location information. Simultaneously, production personnel would investigate the actual production faults to determine the true cause and scope of impact. These root cause tracing results for production warnings and the results of actual production fault investigations are then organized and summarized to construct a root cause analysis interpretive evaluation dataset. This dataset contains historical anomaly root cause analysis results and actual fault situations, providing a data foundation for subsequent evaluation and optimization. For example, within a certain time period, 10 anomaly events occurred, each with corresponding root cause tracing results for production warnings and actual production fault investigations. These results are then organized into the dataset according to a specific format, with each event corresponding to one record containing detailed information about the root cause tracing results for production warnings and the results of actual production fault investigations.

[0138] Step 212: Modularly reconstruct the root cause reasoning engine of the cloud service detection algorithm, embedding an interpretive analysis module during the reasoning process. The interpretive analysis module is used to record the contribution changes of each feature, rule matching paths, and intermediate reasoning results during the root cause reasoning process.

[0139] The root cause reasoning engine of the cloud service detection algorithm is modularly restructured. The engine is decomposed into multiple independent modules, each responsible for different functions, such as feature extraction, rule matching, and inference computation. The purpose of modular restructuring is to improve the maintainability and scalability of the root cause reasoning engine. An interpretive analysis module is embedded in the inference process. The main function of the interpretive analysis module is to record the changes in the contribution of each feature, the rule matching path, and intermediate inference results during root cause reasoning. Different features may contribute differently to the inference result during root cause reasoning. The interpretive analysis module records the changes in the contribution of each feature during the inference process. For example, in a certain inference step, the temperature feature has a larger contribution, while the pressure feature has a smaller contribution; the interpretive analysis module will record this information. The rule matching path refers to the sequence of rules matched during root cause reasoning. The interpretive analysis module records the order of rule matching and the specific rule content. For example, when reasoning about a certain abnormal root cause, if rule A is matched first, and then rule B is matched, the interpretive analysis module will record this rule matching path. Intermediate inference results refer to the intermediate conclusions obtained during root cause reasoning. The explanatory analysis module records detailed information for each intermediate inference result, including inference steps, input data, and output results. By embedding the explanatory analysis module, the root cause reasoning process can be understood more clearly.

[0140] Step 213: Evaluate the output of the root cause reasoning engine using the model interpretability evaluation index. Based on the evaluation results, adjust the weight parameters and rule matching thresholds of each decision node in the root cause reasoning engine through the backpropagation algorithm, and optimize the feature contribution calculation method.

[0141] The output of the root cause inference engine is evaluated using model interpretability evaluation metrics. These metrics include precision, recall, and F1 score, which measure the interpretability and accuracy of the root cause inference results. Precision represents the proportion of root causes correctly predicted by the root cause inference engine, recall represents the proportion of root causes predicted by the engine relative to the actual root causes, and the F1 score is the harmonic mean of precision and recall. Based on the evaluation results, the weight parameters and rule matching thresholds of each decision node in the root cause inference engine are adjusted using the backpropagation algorithm. The backpropagation algorithm is an algorithm used to train neural networks. It calculates the gradient of the error and propagates it back to each decision node, adjusting the node's weight parameters and rule matching thresholds to reduce the error. For example, if the evaluation results show that the accuracy of the root cause inference engine is low, it indicates that the weight parameters of some decision nodes are set unreasonably and need to be adjusted using the backpropagation algorithm. Simultaneously, the rule matching thresholds may also need to be adjusted to improve the accuracy of rule matching. In addition to adjusting the weight parameters and rule matching thresholds, the feature contribution calculation method is also optimized. Optimizing the feature contribution calculation method can improve the accuracy of the contribution of each feature to the inference result. For example, a more reasonable algorithm can be used to calculate the feature contribution, or the calculation method can be adjusted according to different inference scenarios. By continuously adjusting the parameters of the root cause inference engine and optimizing the feature contribution calculation method, the performance and interpretability of the root cause inference engine can be continuously improved.

[0142] Step 214: Periodically deploy the optimized root cause reasoning engine to the cloud service platform and verify the optimization effect through online real-time data. If the explanatory evaluation index does not reach the preset target, continue to execute the above iterative optimization process until the explanatory evaluation index of the root cause analysis results reaches the preset target.

[0143] The optimized root cause inference engine is regularly deployed to the cloud service platform. The deployment process includes uploading the optimized code and model files to the cloud service platform, performing corresponding configuration and testing to ensure the root cause inference engine can run normally on the cloud service platform. After deployment, the optimization effect is verified using real-time online data. Real-time production data generated on the cloud service platform is collected, and the optimized root cause inference engine is used to perform root cause analysis on this data to obtain new root cause inference results. These new root cause inference results are compared with actual production fault diagnosis results, and model interpretability evaluation metrics, such as precision, recall, and F1 score, are calculated. If the interpretability evaluation metrics do not meet the preset targets, it indicates that the root cause inference engine needs further optimization. At this point, the above iterative optimization process continues, including collecting new historical data to construct a root cause analysis interpretability evaluation dataset, modularly reconstructing the root cause inference engine and embedding the interpretability analysis module, evaluating using model interpretability evaluation metrics, adjusting parameters through the backpropagation algorithm, and optimizing the feature contribution calculation method. This iterative optimization process is repeated until the interpretability evaluation metrics of the root cause analysis results meet the preset targets. The preset targets are set based on actual needs and experience, such as requiring an accuracy rate of over 90% and a recall rate of over 80%. When the interpretability evaluation metrics of the root cause reasoning engine reach the preset targets, it indicates that the root cause reasoning engine can perform root cause analysis accurately and interpretably, providing reliable support for production early warning and troubleshooting.

[0144] In yet another non-limiting embodiment, the step of establishing a heterogeneous feature interaction mechanism between the edge-level embedded fusion model and the cloud service detection algorithm is also included:

[0145] Step 311: Set up a feature preprocessing relay station at the edge end, encapsulate the primary early warning labels and corresponding production process status feature sets output by the edge-level embedded fusion model, add feature description header information and data verification code according to the preset data transmission protocol, and form a standardized feature interaction data package.

[0146] A feature preprocessing relay station is set up at the edge to process the primary warning labels and corresponding production process status feature sets output by the edge-level embedded fusion model. After completing data analysis and warning, the edge-level embedded fusion model outputs primary warning labels and production process status feature sets. The primary warning labels contain information about whether there are potential anomalies in the production process, while the production process status feature sets are a set of features reflecting the production process status after a series of processing steps.

[0147] The feature preprocessing relay station first encapsulates these data into features. It combines the initial warning labels and the production process status feature set to form a complete data unit. Then, it adds feature description header information and a data checksum according to a preset data transmission protocol. The feature description header information includes the feature type, dimension, and meaning, which helps the cloud server parse and understand the data. The data checksum is used to verify whether errors occurred during data transmission. By calculating the data checksum, the cloud server recalculates the checksum after receiving the data and compares it with the transmitted checksum. If the two do not match, it indicates that an error may have occurred during data transmission.

[0148] After feature encapsulation and information addition, a standardized feature interaction data packet is formed. This data packet has a unified format and structure, which facilitates transmission and processing between the edge and cloud servers.

[0149] Step 312: Set up a feature receiving and parsing module on the cloud server to parse and verify the received feature interaction data packets, extract the primary warning tags and production process status feature sets, and convert the production process status feature sets into a feature format compatible with the cloud service detection algorithm.

[0150] A feature receiving and parsing module is set up on the cloud server to receive and process feature interaction data packets transmitted from the edge. When the feature interaction data packet arrives at the cloud server, the feature receiving and parsing module first parses it. Based on the feature description header information, different parts of the data packet are separated to extract the primary warning label and the production process status feature set.

[0151] Next, the data is validated. A data verification code is used to verify the data to ensure that no errors occurred during transmission. If the validation passes, subsequent processing continues; if the validation fails, the edge device needs to be notified to retransmit the data or perform error handling.

[0152] After extracting the initial warning labels and production process status feature set, the feature set needs to be converted into a feature format compatible with the cloud service detection algorithm because the edge-level embedded fusion model and the cloud service detection algorithm may have different requirements for feature format. For example, the feature set output by the edge-level embedded fusion model may use a certain set data structure and encoding method, while the cloud service detection algorithm requires data in a different format. The feature receiving and parsing module will convert the production process status feature set according to pre-set conversion rules so that it can be correctly processed by the cloud service detection algorithm.

[0153] Step 313: Construct an edge-cloud feature mapping relationship library to record the correspondence and conversion rules between edge features and cloud features.

[0154] An edge-cloud feature mapping library is constructed to record the correspondence and conversion rules between edge features and cloud features. Since the design and implementation of edge-level embedded fusion models and cloud service detection algorithms may differ, the features they use may also differ. The purpose of the edge-cloud feature mapping library is to establish the connections between these differing features, ensuring that the edge and cloud servers can correctly understand and process each other's data.

[0155] When building a relational database, a detailed analysis and comparison of features from both the edge and cloud sides is required. This involves determining which edge features correspond to which cloud features, and how to perform the transformation. For example, a feature on the edge might require certain calculations or transformations to match a feature on the cloud; the relational database records these transformation rules.

[0156] Relationship databases can be stored in the form of databases or files for easy management and retrieval. In practical applications, when feature interaction is required, the relationship database can be queried to quickly find the correspondence and transformation rules between edge features and cloud features, achieving accurate feature mapping and transformation.

[0157] Step 314: When the feature definition changes due to the upgrade of the edge-level embedded fusion model or the update of the cloud service detection algorithm, the feature mapping relationship library is automatically updated, and the updated information is pushed to the edge through the feature synchronization mechanism.

[0158] In practical applications, edge-level embedded fusion models and cloud service detection algorithms may be upgraded and updated. When these upgrades and updates result in changes to feature definitions, the edge-cloud feature mapping relationship library needs to be automatically updated. For example, after an edge-level embedded fusion model is upgraded, some features may be added or deleted, or the calculation methods and meanings of certain features may change. Similarly, after a cloud service detection algorithm is updated, it may have different requirements for input features. In this case, the system automatically detects the changes in feature definitions and updates the feature mapping relationship library according to the new feature definitions. The process of updating the feature mapping relationship library includes re-analyzing and comparing features at the edge and in the cloud, determining new correspondences and transformation rules, and updating this information in the relationship library. After the update is completed, the updated information is pushed to the edge through a feature synchronization mechanism. The feature synchronization mechanism can be implemented using message queues, network communication, etc. After receiving the updated information, the edge will adjust its feature processing and transmission methods accordingly to ensure that feature interaction with the cloud service can proceed normally. Through this automatic update and synchronization mechanism, the stability and reliability of the heterogeneous feature interaction mechanism between the edge-level embedded fusion model and the cloud service detection algorithm are guaranteed.

[0159] In another non-limiting embodiment, the method further includes:

[0160] Step 410: Deploy a load monitoring agent in the edge computing node cluster to collect load metrics of each edge node when running the edge-level embedded fusion model in real time.

[0161] Load monitoring agents are deployed in edge computing node clusters. Their primary task is to collect real-time load metrics from each edge node as it runs edge-level embedded fusion models. An edge computing node cluster consists of multiple edge nodes, each capable of running edge-level embedded fusion models to analyze and process production process status features. The load monitoring agents periodically collect load metrics from each edge node. These metrics include CPU utilization, memory utilization, disk I / O, and network bandwidth usage. CPU utilization reflects the activity level of the edge node's CPU; high CPU utilization indicates that the node's computing resources may be nearing saturation. Memory utilization reflects the edge node's memory usage; insufficient memory can lead to slow model execution or errors. Disk I / O and network bandwidth usage reflect the edge node's disk read / write and network data transmission activities, respectively. The load monitoring agents send the collected load metrics to a centralized management node, which then aggregates and analyzes these metrics to understand the overall load status of the edge computing node cluster.

[0162] Step 420: Construct a load assessment model, perform a weighted comprehensive assessment of the collected load indicators, and calculate the current load index of each edge node.

[0163] A load assessment model is constructed to comprehensively evaluate the collected load metrics. Different load metrics may have varying degrees of impact on the load of edge nodes, therefore, each metric needs to be assigned a corresponding weight. For example, CPU utilization may have a higher weight because CPU processing power has a significant impact on the operation of the edge-level embedded fusion model. The weights of memory utilization, disk I / O, and network bandwidth usage can be adjusted according to actual conditions. The load assessment model performs a weighted comprehensive evaluation of the load metrics collected for each edge node. Specifically, each load metric is multiplied by its corresponding weight, and then all products are summed to obtain the current load index of that edge node. The current load index comprehensively reflects the load status of the edge node; the higher the index, the heavier the node's load. By calculating the current load index, the load status of each edge node can be understood more accurately.

[0164] Step 430: Set the load threshold range. When the load index of the target edge node exceeds the upper limit threshold, trigger the load balancing adjustment mechanism; when the load index of all nodes is lower than the lower limit threshold, start the node hibernation power saving mechanism.

[0165] Set load threshold ranges, including upper and lower thresholds. The upper threshold indicates that the load on edge nodes has reached a certain level, which may affect the normal operation of the edge-level embedded fusion model, requiring load balancing adjustments. The lower threshold indicates that the load on edge nodes is low, and the resources of the entire edge computing node cluster are largely idle, allowing the node hibernation energy-saving mechanism to be activated. When the load index of a target edge node exceeds the upper threshold, it indicates that the node's load is too high, which may lead to slow model operation and prolonged response time. At this time, the load balancing adjustment mechanism is triggered. The load balancing adjustment mechanism will migrate some tasks from this node to other nodes with lower loads to balance the load of each node. When the load index of all nodes is below the lower threshold, it indicates that the load of the entire edge computing node cluster is low, and resource utilization is not high. At this time, the node hibernation energy-saving mechanism is activated. The node hibernation energy-saving mechanism will put some edge nodes into a hibernation state to reduce energy consumption. When a node is hibernating, the tasks on that node will be migrated to other nodes to ensure that the analysis and processing of the production process state feature set can proceed normally. By setting load threshold ranges and corresponding mechanisms, the load of the edge computing node cluster can be effectively managed, improving resource utilization and energy efficiency.

[0166] Step 440: Based on the current load index, historical load change trend and task processing capacity of each edge node, assign the newly received production process status feature set processing task to the edge nodes that are not overloaded.

[0167] When a new production process state feature set processing task arrives, the task is assigned to an unloaded edge node based on its current load index, historical load trend, and task processing capacity. The current load index reflects the current load status of the edge node; prioritizing nodes with lower load indices avoids overload. Historical load trends help predict future load conditions. If a node's historical load has been relatively stable and its current load is low, it is more likely to handle the new task effectively. Task processing capacity is also a crucial factor. Different edge nodes may have different hardware configurations and computing capabilities, resulting in varying speeds and efficiencies in task processing. For example, a higher-configuration edge node may complete model analysis and processing tasks faster. When assigning tasks, the system first filters out unloaded edge nodes and then evaluates and ranks them based on the factors mentioned above. The most suitable node is then assigned the new production process state feature set processing task. This task allocation method rationally utilizes the resources of the edge computing node cluster, improving task processing efficiency and reliability.

[0168] Step 450: For overloaded nodes that are processing tasks, adopt a task migration strategy to migrate some unfinished tasks to nodes that are not overloaded.

[0169] When an edge node becomes overloaded, meaning its load exceeds the upper limit threshold, a task migration strategy is needed to alleviate the pressure on that node. The specific implementation process of the task migration strategy is as follows: First, the system identifies the overloaded node processing tasks and analyzes the unfinished tasks on that node. Based on factors such as task priority, progress, and data relevance, some unfinished tasks are selected for migration. For example, tasks with lower priority and less data relevance can be migrated first. During task migration, it is necessary to ensure that the task's context information and data can be completely transferred from the overloaded node to the unloaded node. This may involve data copying, transmission, and recovery operations. After the migration is complete, the unloaded node will continue processing these migrated tasks. Simultaneously, the load on the overloaded node will be correspondingly reduced, allowing it to handle the remaining tasks more efficiently. Through the task migration strategy, the load of the edge computing node cluster can be balanced, improving the overall system performance and stability.

[0170] Step 460: Periodically evaluate the load balancing effect and adjust the weight parameters and load threshold range of the load evaluation model based on the evaluation results.

[0171] Regularly evaluate the load balancing effect. The purpose of the evaluation is to check whether the load balancing mechanism effectively balances the load of the edge computing node cluster and improves system performance and resource utilization. Evaluation metrics can include multiple aspects, such as the load balancing degree of each edge node, the average response time of task processing, and the overall system throughput. The degree of load balancing can be measured by calculating the standard deviation of the load index of each edge node. A smaller standard deviation indicates that the load of each node is relatively balanced. The average response time of task processing reflects the efficiency of the system in processing tasks. A shorter average response time indicates better system performance. The overall system throughput represents the number of tasks that the system can process per unit time. Based on the evaluation results, adjust the weight parameters and load threshold range of the load assessment model. If the evaluation results show that the weight settings of certain load metrics are unreasonable, leading to inaccurate load assessment, the weights of these metrics can be adjusted. For example, if it is found that CPU utilization has a significant impact on system performance but has a low weight in the load assessment model, its weight can be appropriately increased. Similarly, if the load threshold range is set inappropriately, causing the load balancing mechanism to not function effectively and in a timely manner, it can also be adjusted based on the evaluation results. For example, if the upper limit threshold is found to be set too high, causing nodes to frequently become overloaded before load migration is initiated, the upper limit threshold can be appropriately lowered. Through regular evaluation and adjustment, the load balancing mechanism can be continuously optimized, enabling the edge computing node cluster to operate more efficiently and stably.

[0172] This invention, through a hierarchical anomaly analysis and early warning architecture at the edge and cloud levels, achieves comprehensive, efficient, and accurate monitoring and early warning of the waterproof membrane production process. At the edge, a pre-established delay compensation model performs physical segment matching on the initial multi-mode monitoring dataset, resolving the issue of data acquisition time differences between different sensors and ensuring data accuracy and consistency. Correlation feature engineering extracts and fuses different types of features from the same-location data packets to generate a production process status feature set, which can more comprehensively and deeply reflect the production process status. The edge-level embedded fusion model performs multi-dimensional analysis on the production process status feature set and outputs primary early warning labels, enabling rapid identification of potential anomalies in the production process at the edge, reducing data transmission volume and response time, and improving real-time performance.

[0173] In the cloud, based on cloud service detection algorithms, anomaly root cause analysis is performed on the abnormal data packets corresponding to the primary warning tags, generating production warning root cause tracing results that include the cause, impact level, and location information of the anomaly. The computing power and rich knowledge base of the cloud enable in-depth root cause mining and impact range assessment, providing detailed and accurate guidance for production adjustments and troubleshooting. The collaborative working mode between the edge and cloud in this embodiment of the invention enables rapid anomaly detection and accurate tracing, improving the quality and efficiency of waterproof membrane production, and reducing production costs and failure risks.

[0174] See Figure 2 As shown in the figure, this is a schematic diagram of the basic structure of a waterproof membrane production monitoring and early warning system 200 based on multimodal data fusion provided in an embodiment of the present invention. The waterproof membrane production monitoring and early warning system 200 based on multimodal data fusion includes:

[0175] Processor 201;

[0176] Storage device 202, on which computer program 2020 is stored;

[0177] When the computer program 2020 is executed by the processor 201, the processor 201 implements any of the aforementioned methods for monitoring and early warning of waterproof membrane production based on multimodal data fusion.

[0178] Based on the above, a readable storage medium is provided, on which a program or instructions are stored, and when the program or instructions are executed by a processor, the steps of the above method are implemented.

[0179] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the systems or apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and relevant parts can be referred to the method section.

Claims

1. A method for monitoring and early warning of waterproof membrane production based on multimodal data fusion, characterized in that, The method includes: The initial multi-mode monitoring dataset is subjected to physical segment same-point matching processing by a pre-established delay compensation model to obtain the same-point matching dataset; the initial multi-mode monitoring dataset contains continuous monitoring data of each process link of the waterproof membrane production line. For each group of data packets at the same location in the same location matching dataset, perform association feature engineering processing, extract different types of features reflecting the production process status from each group of data packets at the same location, cross-associate and fuse the extracted different types of features, and generate a production process status feature set corresponding to each group of data packets at the same location. The production process status feature set is input into the edge-level embedded fusion model for multi-dimensional analysis, and the primary warning label corresponding to each group of data packets at the same location is output. The primary warning label is used to indicate whether there are potential anomalies in the production process corresponding to the group of data packets at the same location. Based on the cloud service detection algorithm, the abnormal root cause analysis is performed on the abnormal data packets corresponding to the primary warning labels to generate production warning root cause tracing results containing the cause of the abnormality, the degree of impact and the location information. The production warning root cause tracing results are used to guide production adjustments and troubleshooting. The process involves inputting the production process status feature set into an edge-level embedded fusion model for multi-dimensional analysis, and outputting a primary warning label corresponding to each group of data packets at the same location, including: The production process status feature set is input into the feature input layer of the edge-level embedded fusion model, and the feature set is subjected to dimension verification and format conversion. The feature enhancement layer of the edge-level embedded fusion model is used to enhance the features of the transformed production process state feature set, and the feature mapping algorithm is used to map the feature set to a high-dimensional feature space. A multi-scale feature extraction layer of an edge-level embedded fusion model is used to perform multi-scale analysis on feature sets in a high-dimensional feature space. Feature information at different levels is extracted through feature extraction windows of different sizes to generate a multi-scale feature set. The multi-scale feature set is input into the anomaly pattern recognition layer of the edge-level embedded fusion model. The anomaly pattern recognition layer contains multiple parallel anomaly detection sub-modules. Each anomaly detection sub-module uses a different anomaly detection algorithm to process the multi-scale feature set and outputs its own anomaly detection result. The system integrates and makes decisions based on the anomaly detection results output by each anomaly detection submodule. It comprehensively judges whether there is an anomaly pattern based on the preset decision rules, and generates and outputs a primary warning tag corresponding to each group of data packets at the same location.

2. The method according to claim 1, characterized in that, The process involves performing correlation feature engineering on each group of data packets in the same location matching dataset, extracting different types of features reflecting the production process status from each group of data packets, and cross-correlated and fused the extracted different types of features to generate a production process status feature set corresponding to each group of data packets, including: Data parsing is performed on each group of data packets at the same location to separate the monitoring data sequences corresponding to different monitoring types, and time domain features are extracted for the data sequences of each monitoring type. Spectral features were extracted from data sequences of various monitoring types to obtain the energy distribution characteristics and characteristic frequency points of the data sequences in different frequency bands. The time-frequency domain features of the data sequence are extracted by the joint time-frequency domain analysis method, capturing the joint variation characteristics of the data sequence in the time and frequency dimensions; The extracted time-domain features, spectral features, and time-frequency domain features are subjected to feature standardization processing to obtain a standardized feature set; Construct a feature correlation matrix, calculate the correlation coefficients between different types of features in the standardized feature set, select feature combinations that meet the correlation conditions based on the correlation coefficients, perform feature fusion operation on the selected feature combinations, and generate a production process status feature set that reflects the production process status.

3. The method according to claim 1, characterized in that, The multi-scale feature extraction layer of the edge-level embedded fusion model performs multi-scale analysis on the feature set in the high-dimensional feature space, extracting feature information at different levels through feature extraction windows of different sizes to generate a multi-scale feature set, including: Multiple feature extraction windows of different sizes are set in the multi-scale feature extraction layer, and the size of each window increases in a preset ratio to form a sequence of extraction windows covering different feature scales. The feature set in the high-dimensional feature space is input into each feature extraction window in sequence. Each window performs local feature extraction on the input feature set and generates a local feature map corresponding to the window size. For each local feature map, feature dimensionality reduction is performed, and the features of each dimensionality-reduced local feature map are standardized so that the local feature maps extracted from different windows have a uniform numerical range. The standardized local feature maps are stitched together in ascending order of window size to form a multi-scale feature set containing feature information at different levels. The feature information at each level corresponds to the local features extracted by the window of the set size.

4. The method according to claim 1, characterized in that, The process of fusing and deciding on the anomaly detection results output by each anomaly detection submodule, comprehensively judging whether an anomaly pattern exists based on preset decision rules, and generating and outputting a primary warning tag corresponding to each group of data packets at the same location, includes: Collect the anomaly detection results output by each anomaly detection submodule. Each anomaly detection result includes an anomaly probability value and an anomaly type identifier. Normalize each anomaly probability value and map each anomaly probability value to the same probability interval; Based on the preset weight allocation scheme for anomaly detection sub-modules, each anomaly detection sub-module is assigned a corresponding decision weight, which is adjusted according to the accuracy and recall of the anomaly detection sub-module in historical detection tasks. The normalized anomaly probability values ​​are weighted and summed according to the decision weights to obtain the comprehensive anomaly probability value. The comprehensive anomaly probability value is compared with the preset anomaly probability threshold. If the comprehensive anomaly probability value is greater than or equal to the anomaly probability threshold, an anomaly mode is determined to exist. Based on the anomaly type identifier output by each submodule, the anomaly type with the highest frequency of occurrence is selected as the current anomaly type, and a primary warning label containing the anomaly type and the comprehensive anomaly probability value is generated. If the overall anomaly probability value is less than the anomaly probability threshold, it is determined that there is no abnormal pattern, and a primary warning label for normal status is generated.

5. The method according to claim 1, characterized in that, The cloud service detection algorithm performs anomaly root cause analysis on the abnormal data packets corresponding to the primary warning tags, generating production warning root cause tracing results that include the cause of the anomaly, the degree of impact, and location information, including: Extract abnormal data packets corresponding to the primary warning labels from the same location matching dataset, perform abnormal scene feature modeling on the abnormal data packets, and construct an abnormal event context description model that includes the time of occurrence of the abnormality, the duration of the abnormality, and the associated process links. Based on the abnormal event context description model, the root cause knowledge graph of the cloud service detection algorithm is invoked. The root cause knowledge graph contains the causal relationship network, propagation path rules and impact range characteristics of various known fault modes in the waterproof membrane production process. By using a multi-dimensional associative reasoning mechanism, the features of abnormal scenarios are deeply matched with the root cause knowledge graph, the initial mapping node set of abnormal events in the root cause knowledge graph is identified, and the cause search is carried out along the causal relationship network according to the preset propagation rules. During the cause-finding process, the process constraints and equipment operating status of the current production environment are introduced as reasoning pruning factors to gradually narrow down the candidate range of potential root causes and form a set of root cause hypotheses. For each candidate root cause in the root cause hypothesis set, the impact range simulation is performed. Based on the propagation path rules in the root cause knowledge graph and the topology of the current production system, the downstream abnormal chain and affected process units caused by the candidate root cause are predicted, and the production early warning root cause tracing results including root cause probability ranking, impact path visualization and key affected area marking are generated.

6. The method according to claim 5, characterized in that, The step of performing abnormal scenario feature modeling on abnormal data packets at the same location, and constructing an abnormal event context description model that includes the time of occurrence of the abnormality, its duration, and related process steps, includes: Analyze the timestamp sequence of abnormal data packets at the same location to determine the start time of the first occurrence of abnormal data and the duration of the abnormal state, and calculate the deviation ratio between the duration of the abnormality and the standard process duration. Extract the changing trend features of data from each monitoring dimension during the abnormal period and construct an abnormal time series evolution feature vector; Based on the matching relationship of the same location, trace the physical roll material location and the production process of the abnormal data, associate the upstream and downstream process information of the production process, and generate a process link association map. By combining the abnormal time-series evolution feature vector, the process link correlation map, and the state switching records in the equipment operation log, a multi-dimensional abnormal scenario description matrix is ​​constructed. The multi-dimensional abnormal scene description matrix is ​​compressed by scene feature dimensionality reduction algorithm to generate a structured abnormal event context description model. The abnormal event context description model includes time dimension features, spatial process correlation features and multi-source data collaborative abnormal features.

7. The method according to claim 5, characterized in that, The root cause knowledge graph adopts a three-layer architecture design: the bottom layer is the basic fault event node layer containing basic fault types; the middle layer is the fault propagation carrier node layer containing propagation medium types; and the top layer contains the derivative impact node layer containing the final impact manifestation. Nodes are connected by directed edges to form a causal relationship network. The edge attributes include propagation probability, influence strength, and triggering conditions. Propagation probability represents the possibility that a change in the state of an upstream node will cause an anomaly in a downstream node, and influence strength represents the propagation coefficient of the degree of anomaly. The root cause knowledge graph has a built-in fault propagation path rule base. It uses production rules to represent propagation path constraints under different process conditions. The antecedent of the rule includes process parameter constraints, and the consequent of the rule defines the propagation direction and path length limits. The influence range characteristics are represented by the set of associated nodes and the spatial distribution function. Each basic fault event node is associated with a set of polygon coordinates of the influence area and an influence decay function, which are used to describe the diffusion law of the anomaly in the physical space. The root cause knowledge graph supports a dynamic update mechanism. It absorbs new fault case data through incremental learning algorithms, automatically updates node attributes and edge propagation probability parameters, and maintains consistency between the knowledge graph and actual production fault modes.

8. The method according to claim 5, characterized in that, The multi-dimensional association reasoning mechanism includes a semantic matching layer, a structural matching layer, and a constraint matching layer. The semantic matching layer is used to calculate the semantic similarity between abnormal scene features and knowledge graph nodes. The structural matching layer is used to analyze the consistency between the relationship between features and the topological structure of the graph. The constraint matching layer is used to verify the conformity between the range of process parameters and the rule triggering conditions. The initial mapping node set is determined by weighted fusion of corresponding matching results; the step of using a multi-dimensional association reasoning mechanism to perform deep matching between abnormal scene features and the root cause knowledge graph, identifying the initial mapping node set of abnormal events in the root cause knowledge graph, and performing causal search along the causal relationship network according to preset propagation rules includes: A bidirectional breadth-first search strategy is adopted. Starting from the initial mapping node, a forward search is performed along the reverse edge of the causal relationship network, and a backward search is performed starting from the preset root cause candidate set. When the two search directions meet at the intermediate node, a complete causal path is formed. During the search process, pruning rules in the propagation rule base are dynamically applied to prune propagation paths that do not meet the current process conditions based on process constraint parameters. When the process parameters of the target node in the path exceed the allowed range of the rules, the path search is terminated. A comprehensive evaluation of the path length, number of nodes, and rule compliance of all complete causal paths generated during the search process is performed. A set number of paths before the evaluation value are retained as potential root cause propagation paths. The starting nodes in the potential root cause propagation paths constitute the root cause hypothesis set. The simulation and deduction of the impact range for each candidate root cause in the root cause hypothesis set, based on the propagation path rules in the root cause knowledge graph and the current production system topology, predicts the downstream anomaly chain triggered by the candidate root cause and the affected process units, including: Construct a directed graph model of the production system topology. In the directed graph model, nodes represent process units and key equipment, and directed edges represent material transmission directions and signal control relationships. Edge attributes include transmission delay time and capacity limit parameters. For each candidate root cause, the standard propagation path template corresponding to the candidate root cause is extracted from the root cause knowledge graph, and the path is instantiated by combining it with the current production system topology. The abstract nodes in the template are replaced with actual production units, and the actual propagation distance and time are calculated. The discrete event simulation algorithm is used to perform the scope of influence deduction. Candidate root causes are injected into the simulation system as initial events. According to the propagation path rules and topology relationship, the abnormal state transition of downstream units is triggered in sequence. The abnormal start time, duration and abnormality level of each unit are recorded. During the simulation, a random disturbance factor is introduced to simulate the uncertainties in actual production. The propagation probability and impact intensity are dynamically adjusted. Each candidate root cause undergoes a number of Monte Carlo simulations to predict the number of times it is affected. The frequency distribution and average anomaly level of each process unit are statistically analyzed. Based on the simulation results, a heat map of the impact of the anomaly is drawn, and the probability of each process unit being affected and the time axis of the anomaly propagation are superimposed to form a spatiotemporally integrated visualization result of the impact range.

9. A monitoring and early warning system for waterproof membrane production based on multimodal data fusion, characterized in that, include: processor; A storage device storing a computer program, which, when executed by the processor, causes the processor to implement the waterproof membrane production monitoring and early warning method based on multimodal data fusion as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Gas leakage time-space correlation early warning method and system based on multi-modal data fusion

    CN120312994A

  • Display panel defect detection method and system based on multi-modal data fusion

    CN120563511A