Intelligent factory fault diagnosis method and system based on AI prediction model
By constructing state evolution features and component association features and using pre-trained models for fault prediction, the accuracy and efficiency issues of smart factory fault diagnosis in existing technologies are solved, and stable operation and efficient production of smart factory production lines are achieved.
Patent Information
- Application Number
- CN202510873159.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-06-27
AI Technical Summary
Existing smart factory fault diagnosis methods rely on manual experience and simple rules, which makes it difficult to comprehensively and accurately identify complex or new faults, and cannot adapt to the complex operating environment of the production line, resulting in low efficiency and prone to errors.
By acquiring equipment monitoring data streams, building state evolution features and component association features, and using pre-trained fault prediction models to predict faults, diagnostic result data is generated, and maintenance guidelines for fault location identification are generated to achieve timely detection and accurate diagnosis of faults.
It improves the accuracy of fault diagnosis and the operational stability of the production line, reduces the losses caused by faults, and improves production efficiency.
Smart Images

Figure CN120387002B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method and system for intelligent factory fault diagnosis based on an AI prediction model. Background Art
[0002] In smart factory operations, stable production line operation is crucial for ensuring product quality, improving production efficiency, and reducing production costs. As the level of automation in smart factories continues to increase, the number and complexity of equipment on the production line are increasing, and the frequency and impact of equipment failures are also expanding.
[0003] At present, the fault diagnosis methods commonly used in smart factories mainly rely on manual experience judgment and monitoring systems based on simple rules. Although manual experience judgment can quickly locate some common faults based on the experience of professionals, this method is not only inefficient but also easily affected by personal subjective factors, making it difficult to comprehensively and accurately identify all potential faults. Monitoring systems based on simple rules have the problem of fixed rule settings and lack of flexibility. They cannot adapt to the complex operating environment and diverse fault modes of the production line, and are often difficult to effectively identify some new or complex faults. Therefore, there is an urgent need for a more intelligent and efficient fault diagnosis method that can accurately predict and diagnose faults in smart factory production lines, take intervention measures in advance, and ensure the stable operation of the production line. Summary of the Invention
[0004] In view of the above-mentioned problems, in combination with the first aspect of the present invention, an embodiment of the present invention provides a smart factory fault diagnosis method based on an AI prediction model, the method comprising:
[0005] Acquire an equipment monitoring data stream of a target production line in a smart factory, wherein the equipment monitoring data stream includes a plurality of groups of status data segments that are continuously collected and have timestamps;
[0006] Performing diagnostic feature construction processing on the device monitoring data stream to generate state evolution features reflecting the device operating state and component association features reflecting the interaction relationship between device components;
[0007] Inputting the state evolution features and the component association features into a pre-trained fault prediction model to perform fault prediction and generate diagnostic result data including a fault risk level;
[0008] Identifying the type of potential fault currently existing in the target production line and the propagation characteristics of the fault during equipment operation based on the diagnostic result data;
[0009] Maintenance guidance data including a fault location identifier is generated based on the potential fault type and the propagation characteristic information, and the maintenance guidance data is transmitted to a plant operation and maintenance system to trigger a fault intervention operation.
[0010] On the other hand, an embodiment of the present invention also provides an intelligent factory fault diagnosis system based on an AI prediction model, including a processor and a machine-readable storage medium, wherein the machine-readable storage medium is connected to the processor, the machine-readable storage medium is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the machine-readable storage medium to implement the above method.
[0011] Based on the above aspects, the embodiment of the present invention obtains the equipment monitoring data stream of the target production line of the smart factory and constructs and processes the diagnostic features thereof to generate state evolution features reflecting the operating status of the equipment and component association features reflecting the interactive relationship between the equipment components, thereby effectively improving the accuracy of fault diagnosis. The constructed state evolution features and component association features are input into the pre-trained fault prediction model for fault prediction, which can make full use of the powerful learning and analysis capabilities of the AI model to generate diagnostic result data including the fault risk level, realize accurate prediction and risk assessment of faults, identify the potential fault types currently existing in the target production line and the propagation characteristic information of the fault during the operation of the equipment according to the diagnostic result data, which helps to gain an in-depth understanding of the nature and development trend of the fault, generate maintenance guidance data including fault location identification based on the potential fault type and propagation characteristic information, and transmit it to the factory operation and maintenance system to trigger the fault intervention operation, thereby realizing timely discovery, accurate diagnosis and effective intervention of faults, greatly improving the operating stability and production efficiency of the smart factory production line, and reducing the losses caused by faults. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 It is a schematic diagram of the execution flow of the intelligent factory fault diagnosis method based on the AI prediction model provided by an embodiment of the present invention.
[0013] Figure 2 Schematic diagram of exemplary hardware and software components of an intelligent factory fault diagnosis system based on an AI prediction model provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0014] The present invention will be described in detail below with reference to the accompanying drawings. Figure 1 This is a flow chart of an intelligent factory fault diagnosis method based on an AI prediction model provided by an embodiment of the present invention. The intelligent factory fault diagnosis method based on the AI prediction model is introduced in detail below.
[0015] Step S110: Acquire a device monitoring data stream of a target production line in a smart factory, wherein the device monitoring data stream includes a plurality of groups of status data segments that are continuously collected and identified by timestamps.
[0016] For example, in a smart factory for papermaking blanket production, the target production line includes several key pieces of equipment, such as looms, setting machines, and heat presses. To accurately monitor equipment operating status, sensors with different functions are installed on each piece of equipment. The loom is equipped with a speed sensor to monitor the main shaft's rotational speed, which directly affects the blanket's weaving efficiency. A tension sensor measures the tension in the warp and weft threads in real time during the weaving process; stable tension is crucial for the blanket's smoothness. A vibration sensor detects vibration during loom operation; abnormal vibration may indicate wear or looseness of mechanical components. The setting machine is equipped with a temperature sensor to monitor the temperature during the setting process; proper temperature is crucial for ensuring blanket setting quality. A pressure sensor measures the pressure applied during setting to ensure the blanket meets specified physical properties. In addition to temperature and pressure sensors, the heat press also includes a displacement sensor to monitor the displacement of the press plate during the hot-pressing process, which is crucial for controlling the blanket's thickness and density.
[0017] In this embodiment, the aforementioned sensors continuously collect equipment operating data at their respective set collection frequencies, and each piece of collected equipment operating data can be accompanied by a precise timestamp. For example, a speed sensor collects loom speed data at set short intervals, marking each data point with the time of collection. Through wired or wireless network transmission, the time-stamped status data from various sensors is aggregated and integrated to form an equipment monitoring data stream. This stream contains multiple sets of continuously collected, time-stamped status data segments, thereby recording the operating status of each piece of equipment on the target production line at different times.
[0018] Step S120: performing diagnostic feature construction processing on the device monitoring data stream to generate state evolution features reflecting the device operating state and component association features reflecting the interaction relationship between device components.
[0019] In order to accurately diagnose possible faults in papermaking blanket production equipment, it is necessary to deeply process the acquired equipment monitoring data stream and construct features that can reflect the equipment operating status and component interaction relationships.
[0020] Step S121: performing data alignment processing on the device monitoring data stream, uniformly mapping status data segments of different acquisition frequencies to a standard time axis, and generating a time-aligned data set.
[0021] During the papermaking fabric production process, different sensors collect data at different frequencies. For example, a tension sensor might collect warp tension data on a loom at a higher frequency, while a temperature sensor might collect temperature data on a setting machine at a lower frequency. This frequency discrepancy results in inconsistent temporal data being collected, hindering subsequent analysis.
[0022] To solve this problem, a standard time interval must be established as a benchmark. For data collected frequently, the data is filtered based on the time points on the standard time axis, retaining only the state data corresponding to the standard time points. For data collected infrequently, interpolation is used to supplement missing data points. For example, when temperature data collected by a temperature sensor is missing on the standard time axis, the temperature value of the missing point is calculated through linear interpolation based on the data points collected before and after. Taking the speed and tension data of a loom as an example, the speed data is filtered to find the values corresponding to the standard time points, and the tension data is interpolated to supplement the missing values. This unifies the data collected at different frequencies onto the standard time axis, generating a time-aligned data set. This ensures temporal consistency for all state data, facilitating subsequent feature extraction and analysis.
[0023] Step S122: performing state evolution extraction processing on the time-aligned data set, analyzing the change trend of continuous state data segments through a sliding time window, and generating state evolution features including state fluctuation patterns and periodic anomaly identifiers.
[0024] Step S1221: Setting a sliding time window of fixed length to segment the time-aligned data set to obtain a plurality of data window units with temporal continuity.
[0025] In the papermaking fabric production scenario, the fixed length of the sliding time window must be set based on the equipment's operating characteristics and data variation patterns. For example, a loom's operating state fluctuates rapidly. For example, the loom's speed may change within a short period of time due to varying weaving process requirements. Therefore, a shorter window length can be used to more accurately capture these state changes. On the other hand, the setting machine operates relatively stably, with temperature and pressure fluctuations being minimal over time. Therefore, a longer window length can be used. The set fixed-length window is sequentially slid from left to right across the time-aligned data set. Each slide captures the data within a window, forming a data window unit. These data window units are continuous in time, each covering state data for a specific time period. For example, for loom speed data, a sliding window can generate a series of data window units containing speed information for different time periods.
[0026] Step S1222: Calculate the mean, variance, and extreme point distribution of the state data segments within each data window unit to generate basic statistical features reflecting state stability.
[0027] A detailed statistical analysis is performed on the status data segments within each data window unit. To calculate the mean, the values of all status data within that data window unit are summed and then divided by the number of data points. The result is the mean. The mean reflects the average operating status of the device within that window period. For example, the mean speed of a loom within a data window unit can reflect the average operating speed of the loom during that period. To calculate the variance, the square of the difference between each data point and the mean is calculated, and then the average of these squared values is taken. The variance measures the dispersion of the data relative to the mean. A larger variance indicates more significant fluctuations in the device status during that period. For example, a large variance in the loom speed within a data window unit indicates significant speed fluctuations and unstable operating conditions. Furthermore, extreme values are identified in the data, including maximum and minimum values. The distribution of extreme values can reflect the extreme operating conditions of the device during that period. For example, the maximum and minimum speed values of a loom within a data window unit can reveal the range of speed fluctuations within that period. These mean, variance, and extreme point distributions constitute the basic statistical characteristics that reflect the stability of the equipment state. By analyzing these characteristics, we can preliminarily determine whether the equipment is operating stably within each time window.
[0028] Step S1223: Analyze the variation of basic statistical features between adjacent data window units, and extract trend parameters of the state data segment in the continuous time window, wherein the trend parameters include an upward trend parameter, a downward trend parameter, and a trend persistence index.
[0029] Compare adjacent data window units and analyze the changes in their basic statistical characteristics. For the mean, if the mean of the subsequent data window unit is greater than that of the previous data window unit, it indicates that the device status is showing an upward trend. To calculate the upward trend parameter, subtract the mean of the previous window from the mean of the subsequent window, and then divide by the time interval to represent the magnitude and speed of the increase. Conversely, if the mean of the subsequent window is less than the mean of the previous window, it indicates a downward trend. The downward trend parameter is calculated similarly. Furthermore, by observing the persistence of trends within multiple consecutive windows, a trend persistence index is determined. If the device status maintains an upward or downward trend over multiple consecutive windows, the trend is highly persistent, and the trend persistence index is high. Conversely, if the trend changes frequently, the trend persistence index is low. For example, for loom speed data, if the mean speed continuously increases over multiple consecutive data window units, it indicates that the loom speed is on an upward trend with strong persistence, and the trend persistence index is high. These trend parameters can be used to describe the dynamic changes in device status within consecutive time windows.
[0030] Step S1224: performing periodic pattern detection processing on the data window unit, identifying the recurring fluctuation period in the state data segment through autocorrelation analysis, and generating a period length parameter and a period stability score.
[0031] Autocorrelation analysis is used to detect periodic patterns in the state data within each data window unit. Autocorrelation analysis identifies periodic patterns in the data by calculating the correlation between the data and itself at different time delays. Specifically, for the state data within a data window unit, correlations are calculated with the data at different time delays, yielding a series of correlation coefficients. A high correlation coefficient at a particular time delay indicates similar fluctuations in the data after that time interval, i.e., periodic fluctuations. When periodic fluctuations are detected, their period length is determined, generating a period length parameter representing the duration of the fluctuation period. A period stability score is also generated based on the stability of the period across multiple data window units. If the period length is relatively stable across different windows with minimal fluctuations, the period stability is high, resulting in a high period stability score. Conversely, if the period length varies significantly, the period stability is low, resulting in a low period stability score. For example, for loom speed data, autocorrelation analysis reveals similar fluctuations after a set time interval. This time interval is then determined as the period length, and a period stability score is then assigned based on the stability of this period length across multiple data window units.
[0032] Step S1225: Integrate the basic statistical characteristics, trend parameters, cycle length parameters and cycle stability scores to obtain state evolution characteristics including state fluctuation patterns and periodic anomaly indicators.
[0033] The previously calculated basic statistical features (mean, variance, and extreme point distribution), trend parameters (upward trend parameter, downward trend parameter, and trend persistence index), cycle length parameter, and cycle stability score are integrated. This integration can be achieved by arranging and combining these features in a predetermined order to form a multidimensional feature vector. These different types of features reflect the changing state of the equipment from different perspectives. Combining them together forms a comprehensive state evolution signature. Analysis of this state evolution signature can identify fluctuation patterns in the equipment state, such as stable fluctuations, violent fluctuations, or cyclical fluctuations. Furthermore, if a low cycle stability score or abnormal trend persistence index is detected, these can be used as indicators of cyclical anomalies. These cyclical anomaly indicators can help identify potential issues in equipment operation. For example, a low cycle stability score in a loom's state evolution signature indicates an unstable operating cycle and a potential failure risk.
[0034] Step S123: performing component association extraction processing on the time-aligned data set, analyzing the coordinated change rules of the state data segments corresponding to different components based on the physical connection relationship of the devices, and generating component association features including interaction strength parameters and association stability indicators.
[0035] Step S1231: Determine key component pairs based on the equipment connection map of the target production line, and extract the time series of status data segments corresponding to each pair of key components.
[0036] In the target production line for papermaking blanket production, specific physical connections exist between the components of different equipment. By consulting the equipment connection map, key component pairs that are critical to the production process are identified. For example, the main shaft and weft feeder of the loom are a pair of key components, and their collaborative operation directly affects the weaving quality of the blanket. From the time-aligned data set, the time series of the corresponding status data segments of each pair of key components are extracted. For the pair of components of the loom main shaft and weft feeder, the time series of the main shaft speed and the time series of the weft feeder feeding speed are extracted respectively. These time series record the operating status of the components at different times. When extracting time series, it is important to ensure the integrity and accuracy of the data to avoid missing or erroneous data.
[0037] Step S1232: Calculate the mutual correlation coefficient of each pair of key component time series to generate an interaction strength parameter that reflects the degree of collaborative change between components.
[0038] For each pair of key component time series, the cross-correlation coefficient is calculated. The cross-correlation coefficient measures the linear correlation between two time series and ranges from -1 to 1. To calculate the cross-correlation coefficient, the two time series are first centered (meaning their respective values are subtracted). The sum of the products of the corresponding time series is then calculated and divided by the product of the standard deviations of the two time series. If the cross-correlation coefficient is close to 1, it indicates that the state data segments of the two components exhibit a strong positive correlation, meaning that their operating state trends are very similar, with a high degree of synergy. If the cross-correlation coefficient is close to -1, it indicates a strong negative correlation. If the cross-correlation coefficient is close to 0, it indicates that the correlation between the state data segments of the two components is weak, with a low degree of synergy. By calculating the cross-correlation coefficient, an interaction strength parameter is generated to reflect the degree of synergy between the components. This interaction strength parameter can be used to characterize the degree of mutual influence between different components. For example, if the cross-correlation coefficient between the time series of the loom spindle speed and the time series of the weft feeder speed is close to 1, it indicates that the operating state trends of the two components are highly consistent, with a high degree of synergy.
[0039] Step S1233: Analyze the fluctuation range of the mutual correlation coefficient within the continuous time window to generate a correlation stability index reflecting the persistence of the component association relationship.
[0040] Observe the changes in the cross-correlation coefficient within consecutive time windows and analyze its fluctuation range. The maximum and minimum values of the cross-correlation coefficient within multiple consecutive data window units can be calculated, and the difference between the two is the fluctuation range. If the cross-correlation coefficient is relatively stable and has a small fluctuation range within multiple consecutive time windows, it indicates that the association between the components is relatively stable, and the generated association stability index is high. Conversely, if the cross-correlation coefficient fluctuates greatly, it indicates that the association between the components is unstable, and the association stability index is low. For example, for the main shaft and weft feeder of a loom, if their cross-correlation coefficient remains within a small fluctuation range within multiple consecutive time windows, it indicates that their collaborative working relationship is relatively stable and the association stability index is high. If the cross-correlation coefficient changes frequently and has a large fluctuation range, it indicates that their collaborative working relationship is unstable and the association stability index is low.
[0041] Step S1234: performing causal relationship detection processing on the state data segments of the key component pairs, determining the causal influence direction between the components through Granger causality test, and generating a causal relationship identifier.
[0042] The Granger causality test method is used to test causal relationships between the status data segments of key component pairs. The basic idea of Granger causality testing is that if one time series (e.g., the status data segment of component A) can help predict another time series (e.g., the status data segment of component B), but not vice versa, then component A can be considered to be the Granger cause of changes in component B. Specifically, two regression models are established: one model contains only the lagged terms of the predicted time series, and the other model contains the lagged terms of the predicted time series and the lagged terms of the predicted time series. The prediction results of the two models are compared to determine whether a Granger causal relationship exists. Using this test method, the direction of causal influence between components is determined, and a causal relationship signature is generated. Causal relationship signatures can clarify the influence relationships between different components and help to deepen understanding of the interaction mechanisms between equipment components. For example, for the spindle speed and weft feed speed of a loom, if changes in the spindle speed can predict changes in the weft feed speed to a certain extent, but not vice versa, then the spindle speed can be considered to be the Granger cause of changes in the weft feed speed, and a corresponding causal relationship signature is generated.
[0043] Step S1235: Integrate the interaction strength parameter, the association stability index and the causal relationship identifier to obtain a component association feature including the interaction strength parameter and the association stability index.
[0044] In this embodiment, the integration method can be to combine the interaction strength parameters, association stability indicators and causal relationship identifiers into a multidimensional feature vector according to set rules. The interaction strength parameters, association stability indicators and causal relationship identifiers describe the interaction relationship between equipment components from different aspects, and they are combined together to form a comprehensive component association feature. The component association feature can comprehensively reflect the coordinated change rules, association stability and causal influence relationship between different components. For example, by integrating the interaction strength parameters, association stability indicators and causal relationship identifiers of the loom main shaft and the weft feeding device, the component association feature obtained can show the interaction situation and causal influence direction between the two components.
[0045] Step S124: inputting the state evolution features and the component association features into a feature calibration module for dimension matching processing, eliminating scale differences between different features, and generating a calibration feature set with a unified representation form.
[0046] Because state evolution features and component association features may have different dimensions and scales, directly using them for subsequent fault prediction may result in inaccurate results. Therefore, these two features are input into the feature calibration module for processing. The feature calibration module first analyzes the dimensions of the state evolution features and component association features. If the dimensions of the two features differ, dimensionality adjustment may be necessary, such as through feature selection or feature extraction, to align their dimensions. At the same time, standardization or normalization is used to eliminate scale differences between the different features. Standardization can be performed by subtracting the mean and dividing by the standard deviation to convert the feature data to data with a mean of 0 and a standard deviation of 1. Normalization maps the feature data to a set interval, such as [0, 1]. For example, some parameters in the state evolution features may have a wide range of values, while some parameters in the component association features have a narrow range. Through processing in the calibration module, their value ranges are unified to a similar scale. After this processing, a set of calibrated features with a unified representation is generated, allowing the different features to function more effectively in the subsequent fault prediction model.
[0047] Step S125: dynamically adjusting the feature weight coefficients of the state evolution feature and the component association feature in fault prediction according to the evaluation result of the degree of influence of each dimensional feature in the calibration feature set on fault diagnosis.
[0048] Evaluate the impact of each dimension of the calibration feature set. For example, machine learning methods such as random forests and gradient boosting can be used to calculate the importance of each dimension to fault diagnosis. This allows each feature to be assigned an importance score based on its performance during model training. For example, in the random forest algorithm, a feature's importance score can be determined by calculating the information gain that the feature brings during the decision tree splitting process. Based on the evaluation results, dynamically adjust the feature weights of the state evolution and component association features in fault prediction. If certain dimensions within the state evolution features have a greater impact on fault diagnosis, increase the weight of the state evolution features accordingly. If certain dimensions within the component association features are more important, increase the weight of the component association features accordingly. This dynamic adjustment allows the fault prediction model to prioritize features with a significant impact on fault diagnosis, improving fault prediction accuracy. For example, if the evaluation results indicate that the speed fluctuation characteristics of a loom have a greater impact on fault diagnosis, while the temperature stability characteristics of a setting machine have a relatively smaller impact, increase the weight of the dimensions within the state evolution features related to loom speed fluctuation and decrease the weight of the dimensions related to setting machine temperature stability.
[0049] Step S130: Inputting the state evolution characteristics and the component association characteristics into a pre-trained fault prediction model to perform fault prediction and generate diagnostic result data including a fault risk level.
[0050] Step S131: inputting the state evolution feature and the component association feature into the feature fusion layer of the fault prediction model, performing cross-dimensional information fusion processing in combination with the feature weight coefficient, and generating a fused feature vector.
[0051] The processed and weighted state evolution features and component association features are input into the feature fusion layer of the fault prediction model. In this layer, cross-dimensional information fusion of the state evolution features and component association features is performed based on the previously determined feature weight coefficients. For example, a dimension feature in the state evolution features and a dimension feature in the component association features are weighted and combined according to the corresponding weight coefficients. This method fuses feature information of different types and dimensions to generate a fused feature vector that integrates information about the device's operating status and component interactions.
[0052] Step S132: performing time-dependency modeling processing on the fused feature vector through the temporal encoder of the fault prediction model to extract temporal context features reflecting the fault development process.
[0053] Step S1321: Input the fused feature vector into the bidirectional recurrent neural network layer of the temporal encoder, perform sequence processing on the fused feature vector from the forward time sequence and the reverse time sequence respectively, and generate a bidirectional hidden state containing past state information and future state information.
[0054] In a papermaking blanket production equipment fault prediction scenario, the fused feature vector is input into a bidirectional recurrent neural network (BRNN) layer. This bidirectional RNN consists of a forward recurrent neural network (RNN) and a reverse recurrent neural network. The forward RNN processes the fused feature vectors sequentially, starting from the start of the time series and proceeding in forward chronological order. At each time step, the forward RNN calculates the hidden state for the current time step based on the input (i.e., the value of the fused feature vector at that time step) and the hidden state from the previous time step. This hidden state contains information about the changes in the equipment state from the start time to the current time, i.e., past state information. For example, for the operating status of a loom, the forward RNN can capture the cumulative changes in the loom's speed, tension, and other states from the time the loom was turned on to the current moment.
[0055] The reverse RNN, on the other hand, starts at the end of the time series and processes the fused feature vector in reverse chronological order. Similarly, at each time step, the reverse RNN calculates the hidden state for the current time step based on the input of the current time step and the hidden state of the next time step. This hidden state contains information about the possible changes in the device state from the current time to a certain point in the future—in other words, future state information. For example, the reverse RNN can infer the possible development of loom failure trends from the current moment forward.
[0056] Through the parallel processing of the forward and backward RNNs, a bidirectional hidden state containing both past and future state information is ultimately generated. This step fully utilizes the contextual information of the time series, laying the foundation for accurately capturing the subsequent fault development process.
[0057] Step S1322: performing attention weight calculation on the bidirectional hidden state, and assigning an attention weight reflecting its importance to fault diagnosis to the hidden state of each time step through the temporal attention mechanism.
[0058] The purpose of the temporal attention mechanism is to highlight time-step information that is crucial for fault diagnosis. For each time-step hidden state in the bidirectional hidden state, it is first input into a fully connected layer for linear transformation, resulting in a new feature vector. This new feature vector is then nonlinearly transformed using an activation function (such as the tanh function) to keep its value within an appropriate range. The transformed feature vector is then dot-producted with a pre-trained attention weight vector to produce a scalar value representing the importance of the hidden state at that time-step for fault diagnosis.
[0059] The importance scores of all time steps are normalized, for example using a softmax function, to convert them into probability distributions, known as attention weights. Attention weights range from 0 to 1, with the sum of the attention weights for all time steps equal to 1. This way, each hidden state at a time step is assigned an attention weight that reflects its importance to fault diagnosis. For example, in the papermaking blanket production process, if the loom speed suddenly fluctuates abnormally at a certain time step, the attention weight corresponding to the hidden state at that time step will be relatively high because of its greater importance to fault diagnosis.
[0060] Step S1323: performing weighted summation processing on the bidirectional hidden state according to the attention weight to generate a temporal aggregation feature containing key time step information.
[0061] After obtaining the attention weight for the hidden state at each time step, the hidden state of each time step in the bidirectional hidden state is multiplied by the corresponding attention weight. Then, the weighted hidden states of all time steps are summed to obtain a comprehensive feature vector, namely, the time series aggregate feature containing key time step information. This weighted summation method highlights time steps that are important for fault diagnosis and filters out relatively unimportant time steps. For example, in the operation data of a loom, the hidden states of stable operation periods that have little impact on fault diagnosis will contribute less to the final time series aggregate feature during the weighted summation process due to their low attention weights. However, the hidden states of critical time steps experiencing abnormal fluctuations will occupy a prominent position in the time series aggregate feature due to their high attention weights.
[0062] Step S1324: input the temporal aggregation features into the gated recurrent unit layer of the temporal encoder for state update processing, filter redundant historical state information and retain key fault evolution clues, and generate temporal context features reflecting the fault development process.
[0063] The Gated Recurrent Unit (GRU) layer consists of a reset gate and an update gate. When time-series aggregated features are input to the GRU layer, the reset gate is first calculated. The reset gate determines whether to ignore the previous hidden state information. It performs a linear transformation and a sigmoid activation function on the current input time-series aggregated features and the hidden state of the previous time step to obtain a value between 0 and 1. If the reset gate value is close to 1, the previous hidden state information is retained; if it is close to 0, the previous hidden state information is ignored.
[0064] Next, the update gate is calculated. The update gate determines how much new information to add to the hidden state and how much of the previous hidden state to retain. It similarly computes a value by applying a linear transformation and a sigmoid activation function to the current input and the previous time step's hidden state. The hidden state for the current time step is updated based on the values of the reset and update gates.
[0065] During this process, the GRU layer automatically filters out redundant historical state information that is unhelpful for current fault diagnosis, retaining only key clues to the fault's evolution. For example, in the operational data of papermaking blanket production equipment, the GRU layer uses a gating mechanism to filter out early, restored equipment state information. However, key information related to the onset and progression of the fault, such as abnormal trends in loom speed and persistent tension instability, is retained and updated to the hidden state. Ultimately, after state updates at the GRU layer, temporal contextual features reflecting the fault's evolution are generated.
[0066] Step S133: Utilize the association decoder of the fault prediction model to perform component interaction relationship analysis on the temporal context features to generate association context features reflecting the fault propagation path.
[0067] Step S1331: Input the temporal context features into the graph convolutional network layer of the association decoder, construct an adjacency matrix with the equipment connection relationship of the target production line as the graph structure, and extract local interaction features between components through graph convolution operations.
[0068] In the target production line for papermaking fabrics, specific physical connections exist between different equipment components. Based on these connections, a graph structure is constructed, where nodes represent equipment components and edges represent connections between components. Based on this graph structure, an adjacency matrix is constructed. Elements in the adjacency matrix indicate information such as the presence and strength of connections between components.
[0069] The temporal context features are input into the graph convolutional network (GCN) layer. During the graph convolution operation, the temporal context features of each node (component) are first linearly transformed to obtain a new feature representation. Then, the feature information of each node's neighboring nodes is aggregated based on the adjacency matrix. Specifically, for each node, the features of its neighboring nodes are multiplied by the weight of the corresponding edge in the adjacency matrix. These weighted neighboring node features are then summed, and the linearly transformed features of the node itself are added. Finally, a nonlinear transformation is performed using an activation function (such as the ReLU function) to obtain the updated features of the node. This graph convolution operation extracts local interaction features between components. For example, in papermaking fabric production equipment, the main shaft of the loom is connected to the transmission components. Graph convolution can capture the local interaction features between the main shaft and the transmission components, understanding their collaborative working, such as how changes in the main shaft speed affect the operation of the transmission components.
[0070] Step S1332: performing global pooling processing on the local interaction features to generate global correlation features that reflect the interaction relationship of the components of the entire production line.
[0071] After obtaining the local interaction features between components, these features are globally pooled. The purpose of global pooling is to integrate the local interaction features to generate a global correlation feature that reflects the interaction relationship between components across the entire production line. Common global pooling methods include global average pooling and global maximum pooling.
[0072] Taking global average pooling as an example, it averages all local interaction features across dimensions. For each dimension, the local interaction feature values of all nodes (components) in that dimension are summed and then divided by the number of nodes to obtain the average value for that dimension. Combining the average values across all dimensions yields a comprehensive feature vector that represents the global interaction relationships between components across the entire production line. For example, in the papermaking blanket production process, global pooling can be used to integrate the local interaction features between various equipment components, such as looms, setting machines, and hot presses, to form a global correlation feature that reflects the collaborative working of the entire production line equipment, thereby understanding the mutual influence and degree of collaboration between the various equipment components throughout the production line.
[0073] Step S1333: Input the local interaction features and the global correlation features into the attention fusion layer of the correlation decoder, assign weights to the local interaction features of different component pairs through the spatial attention mechanism, and highlight the key component pairs that have a significant impact on fault propagation.
[0074] The local interaction features and global correlation features are input into the attention fusion layer of the association decoder. In this attention fusion layer, a spatial attention mechanism is used to assign weights to the local interaction features of different component pairs. First, the local interaction features and global correlation features are concatenated to produce a new feature vector. This new feature vector is then input into a fully connected layer for a linear transformation, followed by a nonlinear transformation using an activation function (such as the tanh function). Next, a dot product is performed on the transformed feature vector with a pre-trained attention weight vector to produce a scalar value representing the importance of the local interaction features of that component pair to fault propagation.
[0075] The importance scores of all component pairs are normalized, for example, using a softmax function, to convert these scores into probability distributions, namely attention weights. Based on the attention weights, corresponding weights are assigned to the local interaction features of different component pairs. Higher weights are assigned to component pairs that play a key role in fault propagation, while lower weights are assigned to component pairs that have less impact on fault propagation. For example, in papermaking blanket production equipment, if a fault between the main shaft and the transmission components of the loom is likely to cause a fault in the entire loom system, then the local interaction features of these two component pairs will be assigned higher weights to highlight their importance in fault propagation; while lower weights will be assigned to some component pairs that have less impact on fault propagation, such as the interaction features between some auxiliary components on the loom.
[0076] Step S1334: The weighted local interaction features and global correlation features are concatenated to generate correlation fusion features containing local interaction details and global correlation information.
[0077] After assigning weights to the local interaction features, the weighted local interaction features and global correlation features are concatenated. This concatenation combines these two features in a set order to form a new feature vector, the correlation fusion feature, which incorporates both the local interaction details between components and the global correlation information for the entire production line. For example, in papermaking blanket production, the correlation fusion feature can simultaneously reflect the local interactions between loom components (such as the details of the coordinated operation between the main shaft and transmission components, and the feeder) as well as the global coordination relationships between the entire production line equipment (such as the mutual influence between the loom, setting machine, and heat press).
[0078] Step S1335: Input the associated fusion features into the fully connected layer of the associated decoder for dimensionality compression processing to generate associated context features reflecting the fault propagation path.
[0079] The associated fusion features are input into the fully connected layer of the associated decoder. The fully connected layer performs dimensionality compression on the associated fusion features. Since the associated fusion features may be high-dimensional and contain a large amount of information, the fully connected layer compresses them to a suitable dimension, extracting the most critical information and generating associated context features that reflect the fault propagation path. The fully connected layer applies a linear transformation to the associated fusion features, mapping them to a lower-dimensional space. In this process, the weight parameters of the fully connected layer are learned through model training, automatically filtering out important information related to the fault propagation path and filtering out redundant information. For example, in the fault analysis of papermaking fabric production equipment, the fully connected layer processes the associated fusion features, removing features that have little impact on the fault propagation path analysis and retaining only key features related to fault propagation, such as which component interactions are critical to fault propagation and the possible propagation direction of the fault. This results in an associated context feature that clearly reflects the fault propagation path between equipment components.
[0080] Step S134: inputting the temporal context features and the associated context features into the risk assessment layer of the fault prediction model for comprehensive assessment processing to generate an intermediate assessment result including the probability value of each fault type.
[0081] The time series context features and the associated context features are input into the risk assessment layer of the fault prediction model. The risk assessment layer first concatenates these two features, combining them into a higher-dimensional feature vector. This concatenated feature vector is then input into a multi-layer perceptron (MLP) for processing. The MLP consists of multiple fully connected layers and activation functions. In each fully connected layer, the input feature vector is linearly transformed, and then nonlinear factors are introduced through activation functions (such as the Reluctant Unified Unit (ReLU) function) to enhance the model's expressiveness.
[0082] In the final layer of the MLP, a softmax activation function is used to convert the output into a probability distribution representing the probability of different fault types occurring. This generates an intermediate assessment result containing the probability values for each fault type. For example, in papermaking blanket production equipment, the risk assessment layer calculates the probability of different fault types, such as loom faults, setting machine faults, and hot press faults, based on temporal and associative context features. This layer comprehensively considers the chronological order of equipment fault development, the interactions between components, and the fault propagation paths.
[0083] Step S135: performing a grading process on the intermediate evaluation results, mapping the probability value of each fault type to a corresponding risk level according to a preset probability threshold interval, and generating diagnostic result data including the fault risk level.
[0084] In this embodiment, a series of probability threshold intervals can be pre-set, with each probability threshold interval corresponding to a specific risk level. These probability threshold intervals are determined based on a large amount of historical failure data and actual production experience. For example, after analyzing numerous failure cases, it is determined that the probability threshold interval corresponding to a low risk level is a relatively low probability range, the probability threshold interval corresponding to a medium risk level is a medium probability range, and the probability threshold interval corresponding to a high risk level is a relatively high probability range.
[0085] For each fault type probability value in the intermediate assessment results, it is compared with the preset probability threshold range. If the probability value of a fault type falls within the threshold range corresponding to the low risk level, then the fault type is judged as low risk; if it falls within the threshold range corresponding to the medium risk level, then it is judged as medium risk; if it falls within the threshold range corresponding to the high risk level, then it is judged as high risk.
[0086] By integrating each fault type and its corresponding risk level information, diagnostic results data including the fault risk level is generated, which can intuitively reflect the potential fault risk of the equipment. For example, the diagnostic results data for papermaking blanket production equipment clearly shows that loom faults are at a high risk level, setting machine faults are at a medium risk level, and hot press faults are at a low risk level, making it easier for operation and maintenance personnel to take appropriate measures based on different risk levels.
[0087] Step S140: Identify the potential fault type currently existing in the target production line and the propagation characteristic information of the fault during equipment operation according to the diagnosis result data.
[0088] Step S141: parsing the fault risk levels and corresponding fault type probability values in the diagnosis result data, and selecting the first several fault types with the highest probability values as potential fault types.
[0089] The diagnostic result data is parsed to extract the fault risk levels and corresponding fault type probabilities. The top several fault types with the highest probability values are then selected from these fault types and identified as potential fault types for the target production line. For example, the diagnostic result data for papermaking blanket production equipment may include probabilities and risk levels for various fault types, such as loom failures, setting machine failures, and hot press failures. The top few fault types with the highest probability values, such as loom spindle failures and setting machine temperature control failures, are selected as potential fault types.
[0090] Step S142: extracting the temporal context features and associated context features corresponding to each potential fault type, analyzing the association relationship between the fault type and the state evolution features and component association features, and generating a fault trigger factor identifier.
[0091] For each potential fault type, its corresponding temporal context features and associated context features are extracted. By analyzing these features, the association between the fault type and the previously extracted state evolution features and component association features is studied. For example, if the temporal context features corresponding to a potential fault type show that the device state has experienced abnormal fluctuations within a certain period of time, and the associated context features show that the interaction relationship between a key component pair has changed, further analysis is conducted in combination with the state evolution features and component association features to find out the possible causes of the fault type and generate a fault trigger factor identification. The fault trigger factor identification can clearly indicate which factors triggered the potential fault. For example, a loom spindle failure may be caused by long-term high-load operation (reflected by the state evolution features) and a loose connection between the spindle and the transmission parts (reflected by the component association features).
[0092] Step S143: performing a retrospective analysis on the status data segments related to the potential fault type in the device monitoring data stream to determine the time point when the fault first occurs and the corresponding initial impact component.
[0093] Perform a retrospective analysis of the status data segments related to potential fault types in the equipment monitoring data stream. Starting from the current moment, trace the data back to look for signs of the first occurrence of the fault. By analyzing the changes in the status data segments, determine the time when the fault first occurs. At the same time, combined with the component association characteristics, find the components affected at that point in time and use them as the initial influencing components. For example, in papermaking blanket production equipment, if the main shaft fault of the loom is found to be a potential fault type, the time when the main shaft fault first occurs is determined by retrospectively analyzing the status data segments such as the speed and vibration of the loom main shaft in the equipment monitoring data stream. Based on the component association characteristics, it is determined that the transmission component connected to the main shaft may be the initially affected component.
[0094] Step S144: Track the state evolution characteristics of the initial impact component in the subsequent time window, and deduce the propagation order and propagation speed of the fault between components by combining the interaction strength parameters and causal relationship identifiers in the component association characteristics.
[0095] Step S1441: Taking the initial impact component as the starting point, determine the set of downstream components directly affected by it according to the causal relationship identifier in the component association feature.
[0096] Starting with the identified initial influencing component, the causal relationship identifiers within the component association features are used to identify the set of downstream components directly affected by the initial influencing component. Causal relationship identifiers clarify the direction of influence between components, and based on these identifiers, it is possible to determine which components are directly affected by the initial influencing component. For example, in papermaking blanket production equipment, if the main shaft of the loom is the initial influencing component, the causal relationship identifiers can be used to identify the transmission components and feeders connected to the main shaft as its directly affected downstream components, forming a downstream component set.
[0097] Step S1442: extract the state evolution characteristic variation of each downstream component in each time window after the initial influencing component fails, and calculate the delay time of the failure impact transmission in combination with the interaction strength parameter.
[0098] For each component in the set of downstream components, the change in its state evolution characteristics in each time window after the failure of the initial influencing component is extracted. The change in the state evolution characteristics reflects the degree of change in the state of the component after the failure occurs. At the same time, combined with the interaction strength parameter in the component association characteristics, the delay time for the failure impact to be transmitted from the initial influencing component to the downstream component is calculated. The interaction strength parameter reflects the degree of coordinated change between components. The stronger the interaction strength, the shorter the delay time for the failure impact to be transmitted may be. For example, in a loom, if the interaction strength between the main shaft and the transmission components is high, then the impact of the main shaft failure on the transmission components may be reflected more quickly and the delay time is shorter.
[0099] Step S1443: determining the order in which the fault propagates from the initial impact component to the downstream components according to the delay time, and generating a propagation sequence identifier including a component sequence list.
[0100] The order in which the fault propagates from the initially impacting component to downstream components is determined based on the calculated fault impact propagation delay time. The shorter the delay time, the earlier the fault propagates to that downstream component. Following this order, the downstream components are arranged into a list, and a propagation sequence identifier containing a component sequence list is generated. For example, in papermaking fabric production equipment, if the delay time for a spindle fault to affect the transmission components is shorter than the delay time for affecting the feeder, the component sequence list in the propagation sequence identifier will list the transmission components before the feeder, indicating that the fault propagates to the transmission components before the feeder.
[0101] Step S1444: Calculate the time interval for the fault to propagate between adjacent components, and generate a propagation speed index reflecting the fault diffusion rate in combination with the physical distance parameters between the components.
[0102] The time interval for a fault to propagate between adjacent components—that is, the time it takes for a fault to propagate from one component to its adjacent components—is calculated. Simultaneously, the physical distance between components is combined to generate a propagation speed metric, reflecting the fault's diffusion rate. This metric can be calculated by dividing the physical distance by the propagation time interval, providing a direct reflection of the fault's propagation speed between components. For example, in papermaking carpet production equipment, if the physical distance between the loom's main shaft and transmission components is constant, and the time interval for a fault to propagate from the main shaft to the transmission components is also constant, calculating these two factors yields a fault propagation speed metric between these two components.
[0103] Step S1445: recursively analyze and process the downstream components of the downstream component until all components affected by the fault are covered, and generate a complete fault propagation sequence list and propagation speed distribution.
[0104] Recursively analyze the downstream components of the downstream components. For example, repeat the steps of identifying downstream components, calculating delay times, and determining the propagation order and velocity until all components affected by the fault are covered. This recursive analysis generates a complete list of fault propagation sequences, clarifying the order in which the fault propagates across all affected components. A propagation velocity distribution is also generated, showing the propagation velocity of the fault across different components. For example, in papermaking fabric production equipment, starting with the initial impacting component, analyze its downstream components, then its downstream components, and so on, ultimately generating a complete list of propagation sequences and a propagation velocity distribution for all affected components.
[0105] Step S1446: Integrate the propagation order list and propagation speed distribution to obtain the propagation order and propagation speed of the fault between components.
[0106] The generated propagation order list and propagation velocity distribution are combined to obtain the fault propagation order and velocity between components, which can comprehensively reflect the fault propagation process between equipment components. For example, in papermaking fabric production equipment, the combined propagation order and velocity can clearly show how the fault propagates from the initial impacting component to other components, as well as the propagation speed between different components.
[0107] Step S145: Integrate the fault trigger factor identifier, initial impact component, propagation sequence and propagation speed to generate propagation feature information reflecting the propagation characteristics of the fault during the operation of the equipment.
[0108] The fault triggering factor identifier, initial impacting component, propagation sequence, and propagation speed are integrated to generate propagation signature information that reflects the fault's propagation characteristics during equipment operation. This propagation signature information includes key information such as the cause of the fault, the starting component, the propagation path, and the propagation speed. For example, in papermaking blanket production equipment, the propagation signature information can clearly indicate that a loom spindle failure was triggered by high load operation and loose connections. Starting from the spindle, it affects other components in a predetermined propagation sequence. The propagation speed of the fault between different components is also given, helping operation and maintenance personnel quickly locate the fault and formulate appropriate maintenance strategies.
[0109] Step S150: generating maintenance guidance data including a fault location identifier based on the potential fault type and the propagation characteristic information, and transmitting the maintenance guidance data to a plant operation and maintenance system to trigger a fault intervention operation.
[0110] Step S151: matching a preset maintenance strategy library according to the potential fault type, and extracting maintenance operation steps and required tool identifiers corresponding to each potential fault type.
[0111] Based on the previously determined potential fault type, a match is performed against a pre-set maintenance strategy library. This library stores the maintenance steps and required tool identifiers corresponding to each fault type. For each potential fault type, the corresponding maintenance steps and required tool identifiers are extracted from the maintenance strategy library. For example, in papermaking carpet production equipment, if the potential fault type is a loom spindle failure, the maintenance strategy library extracts the specific maintenance steps, such as checking the spindle connection and replacing worn parts. The library also extracts the required tool identifiers, such as wrenches and screwdrivers.
[0112] Step S152: parse the initial impact components and the propagation order list in the propagation feature information to determine a set of key components that need to be checked first.
[0113] Parse the initial impacting components and propagation sequence list in the propagation signature information. Based on the initial impacting component being the starting point of the fault and the fault propagation path reflected in the propagation sequence list, determine the set of critical components that require priority inspection. These critical components are the ones most likely to be affected by the fault, and their inspection can quickly pinpoint the fault's scope and extent. For example, in papermaking blanket production equipment, if the propagation signature information indicates that the fault originates from the loom main shaft and affects the transmission components, feeder, and other components in a predetermined order, then the main shaft, transmission components, and feeder are identified as the critical component set requiring priority inspection.
[0114] Step S153: extracting the state data segments corresponding to the key components in the device monitoring data stream, and generating abnormal state descriptions of the key components in combination with the fluctuation patterns and periodic abnormality identifiers in the state evolution characteristics.
[0115] Extract the status data segments corresponding to key components from the equipment monitoring data stream. Combined with the fluctuation patterns and periodic anomaly indicators from the previously extracted state evolution features, analyze the status data segments of key components to generate abnormal state descriptions for these components. For example, for a loom's spindle, extract the status data segments for speed and vibration. If the state evolution features indicate abnormal periodic fluctuations in spindle speed, then generate a spindle abnormal state description as "abnormal periodic fluctuations in spindle speed."
[0116] Step S154: performing an association mapping process on the physical location information of the key component and the abnormal state description to generate a fault location identifier including the component location coordinates and abnormal characteristics.
[0117] The physical location information of key components is mapped to abnormal status descriptions. Each key component on the production line has specific physical location coordinates. These coordinates are associated with the corresponding abnormal status description to generate a fault location identifier that contains the component location coordinates and the abnormal characteristics. For example, in papermaking blanket production equipment, the physical location coordinates of the loom spindle are associated with the abnormal status description of "periodic abnormal fluctuations in spindle speed." This creates a clear fault location identifier, allowing maintenance personnel to quickly find the faulty component.
[0118] Step S155: Integrate the maintenance operation steps, required tool identifiers, key component sets and fault location identifiers to generate a maintenance task list in chronological order.
[0119] The extracted maintenance steps, required tool identifiers, key component sets, and fault location identifiers are integrated. A maintenance task list is generated based on the logical and chronological order of fault handling. This maintenance task list clearly defines the content of each maintenance task, the required tools, the key components involved, and the component locations, and arranges them in a prioritized order. For example, in papermaking blanket production equipment, the maintenance task list might begin with checking the connection of the loom main shaft using a wrench (corresponding to the main shaft fault location identifier), followed by replacing worn parts of the transmission components using a screwdriver (corresponding to the transmission component fault location identifier), and so on.
[0120] Step S156: Binding the maintenance task list with the fault location identifier to generate maintenance guidance data for guiding on-site operation and maintenance personnel.
[0121] Bind the maintenance task list with the fault location identifier. Through the above binding, each maintenance task is associated with the corresponding fault location identifier to form a complete maintenance guide data. The maintenance guide data can intuitively guide on-site operation and maintenance personnel to handle faults. They can quickly and accurately reach the location of the faulty component based on the task list and fault location identifier in the maintenance guide data and take corresponding maintenance operations. For example, in papermaking mesh production equipment, on-site operation and maintenance personnel can first find the location of the loom main shaft according to the maintenance guide data, inspect and repair it according to the steps in the maintenance task list, and then find the location of the transmission components according to the instructions and take corresponding measures. Finally, the generated maintenance guide data is transmitted to the factory operation and maintenance system to trigger the fault intervention operation. The factory operation and maintenance system will arrange the corresponding personnel and resources to handle the fault according to the maintenance guide data.
[0122] Throughout the entire process, the data collected may contain privacy-sensitive data, such as specific device operating parameters. To protect this privacy-sensitive data, various privacy protection and anti-leakage technologies are employed. For example, during data storage, encryption algorithms are used to encrypt data, ensuring it cannot be stolen or tampered with during transmission. For data storage, access control technology is implemented to ensure that only authorized personnel can access this data, and detailed logging of data access is maintained for tracking and auditing purposes. Furthermore, data is anonymized to remove any potentially sensitive information, further enhancing data security. Furthermore, data storage systems are regularly tested for security and vulnerability remediation to prevent the leakage of privacy-sensitive data due to system vulnerabilities.
[0123] In terms of AI model construction and training, the fault prediction model primarily consists of modules such as the feature fusion layer, time series encoder, association decoder, and risk assessment layer. The feature fusion layer is responsible for cross-dimensional information fusion of state evolution features and component association features, combining them with previously dynamically adjusted feature weight coefficients to generate a fused feature vector. During the fusion process, it is crucial to ensure that the dimensions of different features match and their dimensionality is consistent to avoid unreasonable calculations. For example, when performing a weighted addition operation on different features, the dimensionality of these features must be consistent; otherwise, the calculation result may lose practical significance.
[0124] The temporal encoder consists of a bidirectional recurrent neural network layer, a temporal attention mechanism, and a gated recurrent unit layer. The bidirectional recurrent neural network layer processes the fused feature vector sequentially in both forward and reverse time order, generating a bidirectional hidden state containing both past and future state information. This step aims to fully capture the temporal contextual relationship between device states. The temporal attention mechanism assigns attention weights to the hidden state at each time step, highlighting time step information that is important for fault diagnosis. The gated recurrent unit layer performs state update processing on the temporal aggregate features, filtering out redundant information while retaining key fault evolution clues. Ultimately, it generates temporal contextual features that reflect the fault development process.
[0125] The association decoder consists of a graph convolutional network layer, a global pooling layer, an attention fusion layer, and a fully connected layer. The graph convolutional network layer constructs an adjacency matrix based on the equipment connectivity of the target production line. It then extracts local interaction features between components through graph convolution operations. The global pooling layer processes these local interaction features to generate global association features that reflect the component interactions across the entire production line. The attention fusion layer uses a spatial attention mechanism to assign weights to the local interaction features of different component pairs, highlighting key component pairs. The fully connected layer performs dimensionality reduction on the association fusion features to generate association context features that reflect the fault propagation path.
[0126] The risk assessment layer performs a comprehensive assessment on the temporal context features and the associated context features to generate an intermediate assessment result containing the probability value of each fault type. After level classification, it generates diagnostic result data containing the fault risk level.
[0127] Training the fault prediction model requires a large amount of historical equipment monitoring data and corresponding fault label data. The training steps are as follows:
[0128] Historical monitoring data from papermaking fabric production equipment, including status data collected by various sensors, is collected and annotated with information such as the fault type and time of occurrence to form a training dataset. Data preprocessing includes data cleaning to remove noise and outliers, and data normalization or standardization to ensure consistent dimensionality across different features, facilitating model learning.
[0129] Initialize the various modules and parameters of the fault prediction model, including the weights of the feature fusion layer, the neuron parameters of the bidirectional recurrent neural network layer, and the convolution kernel parameters of the graph convolutional network layer. These parameters can be initialized using random initialization, but the parameter values must be within a reasonable range to avoid numerical instability.
[0130] The preprocessed training data is fed into the fault prediction model and trained using a forward and backward propagation process. During forward propagation, the data passes through the feature fusion layer, time series encoder, association decoder, and risk assessment layer to generate predicted fault risk levels and fault type probabilities. During backward propagation, a loss function, such as the cross-entropy loss function, is calculated based on the predicted results and the actual fault label data. This function measures the degree of discrepancy between the model's predictions and the actual labels. Using optimization algorithms, such as stochastic gradient descent, the model's parameters are updated, gradually decreasing the loss function and improving the model's predictive performance.
[0131] Evaluate the trained model using a portion of test data that was not used in training. Calculate model evaluation metrics, such as precision, recall, and F1 score, to assess model performance. If the evaluation metrics do not meet the requirements, adjust the model parameters or add more training data and retrain.
[0132] When the model's evaluation indicators reach a satisfactory level, the trained model is deployed in the actual smart factory environment for real-time fault prediction and diagnosis.
[0133] During the entire fault diagnosis process, when the equipment monitoring data stream indicates abnormal changes in certain loom status data, data alignment is used to unify data collected at different frequencies onto a standard timeline. State evolution and component association extraction are then performed. State evolution features may indicate changes in the loom's speed fluctuation pattern or abnormalities in periodic anomaly indicators. Component association features may indicate changes in the interaction strength parameters and associated stability indicators between the loom's main shaft and transmission components. These features are then fed into a pretrained fault prediction model, which, based on the knowledge and rules acquired through training, predicts the likely fault type and risk level.
[0134] Next, the diagnostic data is used to identify potential fault types and fault propagation characteristics. If a loom main shaft failure is predicted, retrospective analysis is used to determine the time the fault first occurred and the initially affected component. The order and speed of the fault's propagation across components are tracked. For example, a fault might originate in the main shaft and subsequently affect transmission components, feeders, and other components in a predetermined sequence. The propagation speed metric can be used to predict how quickly the fault will spread.
[0135] Finally, maintenance guidance data is generated based on the potential fault type and propagation characteristic information. The preset maintenance strategy library is matched according to the potential fault type, and the maintenance operation steps and required tool identifications are extracted. The set of key components that need to be checked first is determined, such as the main shaft, transmission parts, etc., the abnormal status descriptions of these key components are extracted, and the fault location identification is generated by associating and mapping them with their physical location information. The maintenance task list is bound to the fault location identification to form maintenance guidance data and transmit it to the factory operation and maintenance system. The factory operation and maintenance system arranges operation and maintenance personnel to perform fault intervention operations based on the maintenance guidance data, such as checking the connection of the main shaft, replacing worn transmission parts, etc., so as to solve equipment failures in a timely manner and ensure the normal production of papermaking mesh.
[0136] Figure 2 A schematic diagram illustrates exemplary hardware and software components of an AI-based predictive model-based smart factory fault diagnosis system 100 that can implement the concepts of the present application, as provided in some embodiments of the present application. For example, the processor 120 can be used in the AI-based predictive model-based smart factory fault diagnosis system 100 to perform the functions described in the present application.
[0137] The AI prediction model-based smart factory fault diagnosis system 100 can be a general-purpose server or a special-purpose server, both of which can be used to implement the AI prediction model-based smart factory fault diagnosis method of the present application. Although only one server is shown in this application, for convenience, the functions described in this application can be implemented in a distributed manner on multiple similar platforms to balance the processing load.
[0138] For example, the intelligent factory fault diagnosis system 100 based on the AI prediction model may include a network port 110 connected to the network, one or more processors 120 for executing program instructions, a communication bus 130, and different forms of storage media 140, such as a disk, ROM, or RAM, or any combination thereof. Exemplarily, the intelligent factory fault diagnosis system 100 based on the AI prediction model may also include program instructions stored in ROM, RAM, or other types of non-transitory storage media, or any combination thereof. The method of the present application can be implemented according to these program instructions. The intelligent factory fault diagnosis system 100 based on the AI prediction model also includes an I / O interface 150 between the computer and other input and output devices.
[0139] For ease of explanation, only one processor is described in the smart factory fault diagnosis system 100 based on the AI prediction model. However, it should be noted that the smart factory fault diagnosis system 100 based on the AI prediction model in the present application may also include multiple processors, so the steps performed by one processor described in the present application may also be performed jointly or individually by multiple processors. For example, if the processor of the smart factory fault diagnosis system 100 based on the AI prediction model executes step A and step B, it should be understood that step A and step B may also be executed jointly by two different processors or individually in one processor. For example, the first processor executes step A, the second processor executes step B, or the first processor and the second processor execute steps A and B together.
[0140] In addition, an embodiment of the present invention also provides a readable storage medium, in which computer-executable instructions are preset. When the processor executes the computer-executable instructions, the above-mentioned smart factory fault diagnosis method based on the AI prediction model is implemented.
[0141] It should be noted that in order to simplify the description of the present invention and thus help understand one or more embodiments of the invention, in the foregoing description of the embodiments of the present invention, multiple features are sometimes combined into one embodiment, figure or description thereof.
Claims
1. A smart factory fault diagnosis method based on AI prediction model, characterized in that: The method comprises: Acquire an equipment monitoring data stream of a target production line in a smart factory, wherein the equipment monitoring data stream includes a plurality of groups of status data segments that are continuously collected and have timestamps; Performing diagnostic feature construction processing on the device monitoring data stream to generate state evolution features reflecting the device operating state and component association features reflecting the interaction relationship between device components; Inputting the state evolution features and the component association features into a pre-trained fault prediction model to perform fault prediction and generate diagnostic result data including a fault risk level; Identifying, based on the diagnostic result data, the potential fault types currently existing in the target production line and the propagation characteristic information of the potential fault types during equipment operation; generating maintenance guidance data including a fault location identifier based on the potential fault type and the propagation characteristic information, and transmitting the maintenance guidance data to a plant operation and maintenance system to trigger a fault intervention operation; The step of identifying the potential fault type currently existing on the target production line and the propagation characteristic information of the potential fault type during equipment operation according to the diagnostic result data includes: Analyze the fault risk levels and corresponding fault type probability values in the diagnostic result data, and select the top N fault types with the highest probability values as potential fault types; Extract the time series context features and correlation context features corresponding to each potential fault type, analyze the correlation between the fault type and the state evolution features and component correlation features, and generate the fault trigger factor identification; Performing a retrospective analysis on the status data segments related to the potential fault type in the device monitoring data stream to determine the time point when the potential fault type first appeared and the corresponding initial impact component; Track the state evolution characteristics of the initial impact component in the subsequent time window, and deduce the propagation order and propagation speed of potential fault types between components by combining the interaction strength parameter and causal relationship identifier in the component association characteristics. The interaction strength parameter is the mutual correlation coefficient of each pair of key component time series, which is used to reflect the degree of coordinated change between components. Integrate the fault trigger factor identifier, initial impact component, propagation sequence, and propagation speed to generate propagation characteristic information reflecting the propagation characteristics of the fault during equipment operation; The step of inputting the state evolution features and the component association features into a pre-trained fault prediction model to perform fault prediction and generate diagnostic result data including a fault risk level includes: Inputting the state evolution features and the component association features into the feature fusion layer of the fault prediction model, performing cross-dimensional information fusion processing in combination with feature weight coefficients, and generating a fused feature vector; Performing time-dependency modeling processing on the fused feature vector through the temporal encoder of the fault prediction model to extract temporal context features reflecting the fault development process; Utilizing the association decoder of the fault prediction model to perform component interaction relationship analysis on the temporal context features, and generate association context features reflecting the fault propagation path; Inputting the temporal context features and the associated context features into the risk assessment layer of the fault prediction model for comprehensive assessment processing to generate an intermediate assessment result including a probability value of each fault type; The intermediate evaluation results are graded and processed, and the probability value of each fault type is mapped to the corresponding risk level according to the preset probability threshold interval, so as to generate diagnostic result data including the fault risk level.
2. The intelligent factory fault diagnosis method based on AI prediction model according to claim 1 is characterized in that: The diagnostic feature construction process is performed on the device monitoring data stream to generate a state evolution feature reflecting the device operating state and a component association feature reflecting the interaction relationship between device components, including: Performing data alignment processing on the device monitoring data stream, uniformly mapping status data segments with different acquisition frequencies to a standard time axis, and generating a time-aligned data set; Performing state evolution extraction processing on the time-aligned data set, analyzing the changing trend of continuous state data segments through a sliding time window, and generating state evolution features including state fluctuation patterns and periodic anomaly indicators; Performing component association extraction processing on the time-aligned data set, analyzing the coordinated change patterns of state data segments corresponding to different components based on the physical connection relationship of the devices, and generating component association features including interaction strength parameters and association stability indicators; Inputting the state evolution features and the component association features into a feature calibration module for dimension matching processing to eliminate scale differences between different features and generate a calibration feature set with a unified representation form; According to the evaluation result of the influence degree of each dimensional feature in the calibration feature set on fault diagnosis, the feature weight coefficients of the state evolution feature and the component association feature in fault prediction are dynamically adjusted.
3. The intelligent factory fault diagnosis method based on AI prediction model according to claim 2 is characterized in that: The state evolution extraction process is performed on the time-aligned data set, and the change trend of the continuous state data segment is analyzed through a sliding time window to generate a state evolution feature containing a state fluctuation pattern and a periodic anomaly indicator, including: Setting a sliding time window of fixed length to segment the time-aligned data set to obtain a plurality of data window units with temporal continuity; Calculate the mean, variance and extreme point distribution of the state data segment in each data window unit to generate basic statistical characteristics reflecting state stability; Analyze the changes in basic statistical characteristics between adjacent data window units and extract trend parameters of the state data segment in the continuous time window, wherein the trend parameters include an upward trend parameter, a downward trend parameter and a trend persistence indicator; Performing periodic pattern detection processing on the data window unit, identifying the recurring fluctuation period in the state data segment through autocorrelation analysis, and generating a period length parameter and a period stability score; The basic statistical characteristics, trend parameters, cycle length parameters and cycle stability scores are integrated to obtain state evolution characteristics including state fluctuation patterns and periodic anomaly markers.
4. The intelligent factory fault diagnosis method based on AI prediction model according to claim 2 is characterized in that: The component association extraction process is performed on the time-aligned data set, and based on the physical connection relationship of the devices, the coordinated change rules of the corresponding status data segments of different components are analyzed to generate component association features including interaction strength parameters and association stability indicators, including: Determine key component pairs based on the equipment connection map of the target production line and extract the time series of the corresponding status data segments of each key component pair; The mutual correlation coefficient of each pair of key component time series is calculated to generate the interaction strength parameter reflecting the degree of synergistic change between components; Analyzing the fluctuation range of the mutual correlation coefficient within a continuous time window to generate a correlation stability index reflecting the persistence of the component correlation relationship; Performing causal relationship detection processing on the state data segments of the key component pairs, determining the causal influence direction between the components through Granger causality test, and generating a causal relationship identifier; The interaction strength parameter, the association stability index, and the causal relationship identifier are integrated to obtain a component association feature including the interaction strength parameter and the association stability index.
5. The intelligent factory fault diagnosis method based on AI prediction model according to claim 1 is characterized in that: The performing time-dependency modeling processing on the fused feature vector by the time series encoder of the fault prediction model to extract time series context features reflecting the fault development process includes: Inputting the fused feature vector into the bidirectional recurrent neural network layer of the temporal encoder, performing sequence processing on the fused feature vector from a forward time sequence and a reverse time sequence respectively, to generate a bidirectional hidden state containing past state information and future state information; Performing attention weight calculation on the bidirectional hidden state, and assigning an attention weight reflecting its importance to fault diagnosis to the hidden state of each time step through a temporal attention mechanism; Performing weighted summation processing on the bidirectional hidden state according to the attention weight to generate a temporal aggregation feature containing key time step information; The temporal aggregation features are input into the gated recurrent unit layer of the temporal encoder for state update processing, redundant historical state information is filtered and key fault evolution clues are retained, thereby generating temporal context features reflecting the fault development process.
6. The intelligent factory fault diagnosis method based on AI prediction model according to claim 1 is characterized in that: The utilizing the association decoder of the fault prediction model to perform component interaction relationship parsing processing on the temporal context features to generate association context features reflecting the fault propagation path includes: Inputting the temporal context features into the graph convolutional network layer of the association decoder, constructing an adjacency matrix based on the equipment connection relationship of the target production line as a graph structure, and extracting local interaction features between components through graph convolution operations; Performing global pooling processing on the local interaction features to generate global correlation features that reflect the interaction relationship between components of the entire production line; Inputting the local interaction features and the global correlation features into the attention fusion layer of the correlation decoder, assigning weights to the local interaction features of different component pairs through a spatial attention mechanism, and highlighting the key component pairs that have a significant impact on fault propagation; The weighted local interaction features and global correlation features are concatenated to generate correlation fusion features that contain local interaction details and global correlation information; The associated fusion features are input into the fully connected layer of the associated decoder for dimensionality compression processing to generate associated context features reflecting the fault propagation path.
7. The intelligent factory fault diagnosis method based on AI prediction model according to claim 1 is characterized in that: Tracking the state evolution characteristics of the initially impacting component in a subsequent time window, combining the interaction strength parameter and causal relationship identifier in the component association characteristics, and deducing the propagation order and propagation speed of the fault between components includes: Taking the initial impact component as a starting point, determining a set of downstream components directly affected by the initial impact component according to the causal relationship identifier in the component association feature; Extract the state evolution characteristics of each downstream component in each time window after the initial impact component fails, and calculate the delay time of the failure impact transmission based on the interaction strength parameter; determining the order in which the fault propagates from the initial impact component to the downstream components based on the delay time, and generating a propagation sequence identifier including a component sequence list; Calculate the time interval for fault propagation between adjacent components and generate a propagation speed index reflecting the fault diffusion rate by combining the physical distance parameters between components; Recursively analyzing the downstream components of the downstream component until all components affected by the fault are covered, generating a complete fault propagation order list and propagation speed distribution; The propagation order list and the propagation speed distribution are integrated to obtain the propagation order and propagation speed of the fault among the components.
8. An intelligent factory fault diagnosis system based on AI prediction model, characterized in that: It includes a processor and a memory, the memory is connected to the processor, the memory is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the memory to implement the smart factory fault diagnosis method based on the AI prediction model as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Electricity consumption data intelligent monitoring method and system for environmental protection equipment
CN118551206A
Chip fault prediction method and device, equipment and storage medium
CN119598222A