UV curing machine fault detection method based on data analysis
By collecting and analyzing the voltage data points of the UV curing machine, calculating the noise probability and optimizing the clustering algorithm, the noise interference problem of the voltage data points is solved, and the accuracy and reliability of the UV curing machine voltage fault detection are improved.
Patent Information
- Application Number
- CN202511234589.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-01
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-09-01
AI Technical Summary
In the prior art, the voltage data points of the UV curing machine generate noise data due to electromagnetic interference, which leads to inaccurate clustering results of the agglomerative hierarchical clustering algorithm, affecting the accuracy and effectiveness of voltage fault detection.
By continuously collecting voltage data points of UV curing machines, the suspected noise factor and fluctuation possibility of each data point are calculated, the noise probability is determined, and the cluster score value is optimized in the agglomerative hierarchical clustering algorithm to eliminate noise interference and improve the accuracy of the clustering results.
It achieves accurate identification of UV curing machine voltage data points, improves the accuracy and reliability of voltage fault detection, reduces misjudgment, and ensures stable operation of the equipment.
Smart Images

Figure CN120742006A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electrical parameter measurement and industrial equipment fault diagnosis, and in particular to a UV curing machine fault detection method based on data analysis. Background Art
[0002] In modern manufacturing, UV curing machines are commonly used in processes such as surface coating and curing light-curing resins. As essential equipment in these production lines, the voltage stability of UV curing machines directly determines the lifespan of electrical components and the curing effect. Voltage anomalies (such as transient fluctuations and sustained offsets) are the primary cause of equipment failure, resulting in incomplete curing and product failure at the very least, or even serious failures such as lamp burnout and power module damage.
[0003] Traditional manual detection and regular inspections are unable to capture subtle voltage anomalies in real time, and often require passive repairs only after the equipment shuts down due to voltage failures, significantly increasing maintenance costs and downtime losses.
[0004] In the prior art, the agglomerative hierarchical clustering algorithm is a commonly used data anomaly detection algorithm. It uses this algorithm to identify abnormal patterns in voltage data and determine whether a voltage fault exists, allowing timely intervention to prevent damage to electrical components. This clustering method is a recursive aggregation process that uses a single-link aggregation method. Specifically, the distance between the two closest data points in two clusters is selected as the distance between the two clusters. Clusters are then merged based on this distance to generate a hierarchical clustering result tree. In other words, the clustering operation is based on the distance between the two closest data points in the two clusters. This group of data points is the key to single-link aggregation and determines the final clustering results.
[0005] However, when collecting voltage data points from UV curing machines, electromagnetic interference from surrounding electrical equipment like motors and switching power supplies can cause noise in the voltage data. This noise interferes with the sensor's normal readings, distorting the voltage data points. When merging two clusters using single-link aggregation, if noise exists in both clusters, it reduces the accuracy of the distance calculation between the two clusters, affecting the accuracy of the clustering results. This can ultimately lead to misjudgments of the UV curing machine's voltage during operation, compromising the accuracy and effectiveness of voltage fault detection. Summary of the Invention
[0006] In order to solve the problem that when detecting abnormal conditions of voltage data points of UV curing machines, the clustering results are inaccurate due to noise data interference, which in turn causes misjudgment of the voltage abnormality of the UV curing machine and affects the accuracy and effectiveness of UV curing machine voltage fault detection, the present invention proposes a UV curing machine fault detection method based on data analysis.
[0007] In one aspect, the present invention provides a UV curing machine fault detection method based on data analysis, comprising: Continuously collect voltage data points of the UV curing machine during operation and divide the surrounding data segments of each voltage data point; Determine the suspected noise factor of each voltage data point based on the difference between each voltage data point and the mean of the surrounding data segments, and the difference between each voltage data point and its left and right adjacent voltage data points; determine the noise probability of each voltage data point based on the suspected noise factor of each voltage data point and the fluctuation probability of each voltage data point; The voltage data points are clustered using an agglomerative hierarchical clustering algorithm to obtain multiple clusters. In the process of recursively aggregating the multiple clusters, each two clusters are regarded as a group of clusters. One voltage data point is selected from each of the two clusters in each group to form a group of data points. The score value of each group of voltage data points in each cluster is determined based on the noise probability and distance of each group of data points. According to the score value of each group of voltage data points in each cluster, the group of voltage data points that each cluster is based on during single-link aggregation is determined to optimize the agglomerative hierarchical clustering algorithm; the optimized agglomerative hierarchical clustering algorithm is used to perform anomaly detection on the voltage data points of the UV curing machine during operation, and whether the UV curing machine has a voltage fault is determined based on the anomaly detection results.
[0008] This technical solution continuously collects a rich set of voltage data points and determines a local data range to facilitate analysis of the characteristics of the data point within its local range. By comprehensively considering the suspected noise factor and fluctuation probability of each voltage data point, the noise probability is determined. This allows for more accurate identification of noise data, avoiding misclassification of normally fluctuating voltage data points as noise data. This also improves the accuracy of the judgment of true noise data, providing more reliable input data for subsequent clustering algorithm optimization. During the recursive aggregation process of the agglomerative hierarchical clustering algorithm, the noise probability of each group of data points in each cluster and the distance between each group of data points are calculated to determine the score value. This provides a quantitative standard for the subsequent determination of the data points based on single-link aggregation. This allows for comprehensive consideration of the noise probability and distance factors of data points during clustering, avoiding incorrect clustering due to noise data, improving the accuracy of the clustering results, and making the anomaly detection results for voltage data points more accurate. Based on accurate anomaly detection results, it is possible to more accurately determine whether a UV curing machine has a voltage fault, improving the accuracy and reliability of UV curing machine fault detection.
[0009] Furthermore, the suspected noise factor of each voltage data point is determined based on the following method: Calculate the The absolute value of the difference between the voltage data point and its adjacent data point on the left, and the The absolute value of the difference between a voltage data point and its adjacent data point on the right, and the absolute value of the larger difference is recorded as ; No. The suspected noise factor of the voltage data point Where, For the The absolute value of the difference between a voltage data point and the mean of the surrounding data segments, is a hyperparameter.
[0010] This technical solution uses this comprehensive calculation method to obtain a suspected noise factor that more comprehensively and accurately reflects the likelihood that each voltage data point is noise data. It considers both the deviation of the voltage data point within a local range and the difference between the voltage data point and adjacent voltage data points, making the identification of noise data more reliable.
[0011] Furthermore, the noise probability of each voltage data point is determined based on the following method: a normalized value of the ratio of the suspected noise factor of each voltage data point to the fluctuation possibility of each voltage data point is used as the noise probability of each voltage data point.
[0012] This technical solution comprehensively considers the suspected noise factor and fluctuation probability of each voltage data point, avoiding the potential misjudgment of noise data based solely on data distribution characteristics. Specifically, considering that normal voltage fluctuations at certain stages of operation may cause a voltage data point to deviate significantly from surrounding voltage data points, but these voltage data points are not noise data. By factoring in the possibility of fluctuation, true noise data can be more accurately identified, reducing misjudgment of normal fluctuating data.
[0013] Furthermore, the fluctuation probability of each voltage data point is determined based on the following method: obtaining the time interval between the corresponding moment of each voltage data point and the start-up moment of the current UV curing machine, as well as the current data point at the corresponding moment of each voltage data point; selecting all historical voltage data points consistent with each voltage data point, and calculating the average of all historical current data points corresponding to all selected historical voltage data points as the reference current value of the voltage data point; Calculate the probability of fluctuation for each voltage data point: , where For the The fluctuation probability of each voltage data point, For the The time interval between the corresponding moment of each voltage data point and the start-up moment of this UV curing machine, For the The absolute value of the difference between the current data point at the moment corresponding to the voltage data point and the reference current value of the voltage data point, is the natural exponential function.
[0014] This technical solution takes into account that voltage fluctuations may be more frequent during the initial startup of a UV curing machine. As the voltage gradually stabilizes over time, the likelihood of voltage fluctuations decreases. By combining the time interval and current to determine the likelihood of fluctuation, the voltage fluctuation characteristics of the UV curing machine at different operating stages can be more accurately quantified.
[0015] Furthermore, the score value of each voltage data point in each cluster is calculated based on the following formula: Where, For the The cluster The score value of the group voltage data point, and Respectively The cluster The noise probability of the first and second voltage data points in a group of voltage data points, For the The cluster The distance between two voltage data points in a group of voltage data points, is a natural exponential function, where the distance between two voltage data points is determined based on the absolute value of the difference between the two voltage data points.
[0016] This technical solution determines a score by comprehensively considering the noise probability of each voltage data point within each cluster and the distance between each voltage data point. When selecting a set of voltage data points for single-link aggregation, it prioritizes those with a low noise probability and close distances. This effectively eliminates noise interference during the clustering process and improves the accuracy of the clustering results.
[0017] Furthermore, the method for determining the set of voltage data points based on which each cluster is aggregated in a single link is as follows: calculating the scores of all voltage data points of each cluster, and selecting the set of voltage data points with the highest score as the set of voltage data points based on which the cluster is aggregated in a single link.
[0018] This technical solution uses the set of voltage data points with the highest scores as the set of voltage data points on which the clusters are based during single-link aggregation. Compared with the traditional method of directly selecting the closest data point between two clusters as the distance metric, this method can more accurately reflect the true distance relationship between the two clusters because it selects a set of voltage data points with less noise influence and higher data reliability.
[0019] Furthermore, abnormal detection is performed on the voltage data points of the UV curing machine during operation, including: In the recursive aggregation process of the agglomerative hierarchical clustering algorithm, a distance threshold is preset. If the inter-cluster distance between a cluster and another cluster when they are merged is less than the distance threshold, the cluster after the merger is regarded as a normal cluster; if the inter-cluster distance between a cluster and another cluster when they are merged is greater than or equal to the distance threshold, the cluster and the other cluster are separated and regarded as two suspected abnormal clusters. A preset number threshold; if the number of voltage data points of a suspected abnormal cluster is greater than the preset number threshold, the suspected abnormal cluster is removed, and the remaining suspected abnormal cluster is regarded as an abnormal cluster, and the voltage data points contained in the abnormal cluster are regarded as abnormal voltage data points to complete the abnormal detection of the voltage data points.
[0020] Furthermore, a method for determining whether a voltage fault occurs in the UV curing machine based on the abnormality detection result is as follows: a preset proportion threshold value is set; if the proportion of the number of abnormal voltage data points to the total number of voltage data points within M minutes before the current analysis time exceeds the proportion threshold value, it is determined that a voltage fault occurs in the UV curing machine, and relevant personnel are notified to conduct an inspection; if the proportion of the number of abnormal voltage data points to the total number of voltage data points within M minutes before the current analysis time does not exceed the proportion threshold value, it is determined that no voltage fault occurs in the UV curing machine; M is a preset value.
[0021] Furthermore, the method for dividing the surrounding data segments of each voltage data point is: taking each voltage data point as the center, selecting a preset number of voltage data points on both sides of the voltage data point to form the surrounding data segments of the voltage data point.
[0022] Furthermore, the reference current value of the voltage data point is determined based on the following method: for any voltage data point, all historical voltage data points consistent with the voltage data point are selected, and the average of all historical current data points corresponding to all the selected historical voltage data points is calculated as the reference current value of the voltage data point.
[0023] The present invention has the following effects: The present invention uses a series of methods to accurately identify noise data in voltage data points. When using the agglomerative hierarchical clustering algorithm to perform anomaly detection on voltage data points, it can effectively eliminate noise interference, realize the optimization of the agglomerative hierarchical clustering algorithm, and improve the accuracy and stability of the clustering results. Based on the accurate clustering results, the voltage anomaly of the UV curing machine during operation can be accurately judged, thereby making the voltage fault detection of the UV curing machine more accurate. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 It is a schematic flow chart of the method of the present invention; Figure 2 is a schematic flow chart of the method of step S2 of the present invention; Figure 3 is a schematic flow chart of the method of step S3 of the present invention; Figure 4 It is a schematic diagram of the hierarchical clustering result tree of the present invention. DETAILED DESCRIPTION
[0025] The technical solutions in the embodiments of the present invention will be described clearly and completely below with reference to the accompanying drawings in the embodiments of the present invention.
[0026] Reference Figure 1 The present invention provides a UV curing machine fault detection method based on data analysis, including steps S1 to S5: S1: Collect voltage data points of the UV curing machine during operation.
[0027] In order to accurately obtain the voltage data points of the UV curing machine during operation, a voltage sensor is selected to perform the data acquisition task, and the acquisition time is set to one hour. During this period, voltage data points are collected at a frequency of five times per second. This frequency can more intensively capture the subtle changes in the voltage of the UV curing machine during operation, providing a sufficient and detailed data basis for subsequent analysis.
[0028] S2: Evaluate the noise probability for each voltage data point.
[0029] To accurately assess the noise probability of each voltage data point, this step requires analyzing the numerical performance and numerical variation of each voltage data point within the surrounding data segments to determine the suspected noise factor for each voltage data point. However, during UV curing machine operation, the voltage may rise in a gentle, fluctuating manner. Therefore, determining the suspected noise factor based solely on the relative size of a data point relative to its surrounding data points can easily misjudge normal voltage fluctuations as abnormal. Therefore, it is necessary to combine the suspected noise factor and fluctuation probability of each voltage data point to determine a noise probability that accurately reflects the noise situation at each voltage data point.
[0030] like Figure 2 As shown, the following steps are included: S21: Determine the suspected noise factor for each voltage data point.
[0031] Since noise data points usually have large numerical differences from their surrounding data points, when analyzing the suspected noise factor of each voltage data point, the more prominent the numerical performance of a voltage data point in its surrounding data segment, the more likely it is noise data, and the larger the suspected noise factor. The greater the numerical change between a voltage data point and its adjacent voltage data points, the more likely it is noise data. The more prominent the numerical performance of a voltage data point in the surrounding data segment, the greater the credibility, and the corresponding suspected noise factor will also be larger.
[0032] Therefore, with each voltage data point as the center, a preset number of voltage data points are selected on both sides of the voltage data point to form a data segment around the voltage data point. The preset number is 20 (empirical value). If the number of voltage data points on one side of the voltage data point is less than 20, it can be supplemented on the other side.
[0033] For the voltage data points, calculate the The absolute value of the difference between the voltage data point and its adjacent data point on the left, and the The absolute value of the difference between a voltage data point and its adjacent data point on the right, and the absolute value of the larger difference is recorded as This operation can capture the sudden change of the voltage data point.
[0034] It is specially noted that if a voltage data point has only a left adjacent voltage data point or only a right adjacent voltage data point, then the absolute value of the difference between the voltage data point and its only one adjacent voltage data point is used.
[0035] Rule No. The suspected noise factor of the voltage data point for:
[0036] In this formula, For the The absolute value of the difference between a voltage data point and the mean of the surrounding data segments, is a hyperparameter, , which exists to prevent A case of 0 occurs.
[0037] If a voltage data point differs significantly from its adjacent voltage data points, it indicates a significant deviation in value from the surrounding voltage data points, making it more likely to be noise data. In a normal voltage data point variation trend, the voltage values between adjacent voltage data points typically do not jump significantly. If a significant jump occurs, it is likely due to noise interference.
[0038] The mean of the surrounding data segments of a voltage data point represents the typical characteristics of the voltage data point in the surrounding data segments. The difference between a voltage data point and the mean of its surrounding data segments reflects the degree of deviation of the voltage data point relative to the overall level of the surrounding data segments. The greater the difference, the more the voltage data point deviates from the normal range of the surrounding data segments, and the greater the possibility that the voltage data point is noise data.
[0039] S22: Calculate the fluctuation possibility of each voltage data point.
[0040] The suspected noise factor of each voltage data point is obtained based on the numerical performance of each voltage data point in the surrounding data segment. This step further takes into account that the voltage data points of the UV curing machine will also have certain regular fluctuations, which may be confused with the noise data. Therefore, it is necessary to further combine the suspected noise factor of the voltage data point with analysis of other dimensions to more accurately reflect the noise probability of a voltage data point.
[0041] Specifically, this step takes into account that the voltage data points will fluctuate slightly after the UV curing machine is started, and then stabilize after a period of time. Furthermore, during operation, the voltage and current data points typically exhibit a certain correlation: when the voltage increases, the current also increases, and vice versa.
[0042] Therefore, this step will determine the fluctuation possibility of each voltage data point by analyzing the correlation between each voltage data point and the current data point and the time interval between the corresponding moment of the voltage data point and the current start moment of the UV curing machine.
[0043] The smaller the time interval between a voltage data point and the UV curing machine's current startup moment, the more likely it is that the corresponding moment of the voltage data point was in the early stages of the UV curing machine's current startup. This also increases the likelihood that regular voltage fluctuations were present at the corresponding moment of the voltage data point, and the lower the noise probability of the voltage data point. Furthermore, the stronger the correlation between the voltage data point and the current data point at the corresponding moment, the greater the likelihood of regular voltage fluctuations at the corresponding moment of the voltage data point. Based on these dimensions and in combination with the suspected noise factors obtained for each voltage data point, the noise probability of the voltage data point is determined.
[0044] In one embodiment, the method for obtaining the fluctuation possibility of each voltage data point is: Obtain the time interval between the corresponding moment of each voltage data point and the start-up moment of this UV curing machine, as well as the current data point at the corresponding moment of each voltage data point; Select all historical voltage data points that are consistent with each voltage data point, and calculate the mean of all historical current data points corresponding to all the selected historical voltage data points as the reference current value of the voltage data point; For example, a voltage data point , the current data point at the corresponding moment , All historical voltage data points of X5 are , , , , the current data points at the corresponding moments are , , , ; and is with All historical voltage data points will be consistent and Current data points at the corresponding time and The mean value of 120 is used as the voltage data point The reference current value.
[0045] Calculate the fluctuation probability of each voltage data point. The voltage data points have a fluctuation probability of:
[0046] In this formula, For the The fluctuation probability of each voltage data point, For the The time interval between the corresponding moment of each voltage data point and the start-up moment of this UV curing machine, Indicates the The absolute value of the difference between the current data point at the moment corresponding to the voltage data point and the reference current value of the voltage data point, is the natural exponential function. Reflects the operation stage, Reflects the degree of current deviation, the two are multiplied Constitute a comprehensive indicator to represent the The larger the product is, the farther the state of the voltage data point at the corresponding moment deviates from the steady state, and the smaller the possibility of fluctuation is, and vice versa. Build and volatility probability The negative correlation between and satisfies the above logic.
[0047] The current analysis time and this The time interval between the start-up time of the curing machine The smaller the Voltage data points are in this The greater the possibility of the curing machine starting early, the The greater the possibility that there is a regular voltage fluctuation at the corresponding moment of the voltage data point, the greater the probability that the voltage fluctuation at the corresponding moment of the voltage data point The smaller the noise probability of each voltage data point, the smaller it is, and vice versa. The smaller the The stronger the correlation between a voltage data point and its current data recorded at the same time, the The greater the possibility that there is a regular voltage fluctuation at the voltage data point, the greater the credibility. The smaller the noise probability of each voltage data point, the smaller it is, and vice versa.
[0048] S23: Determine the noise probability of each voltage data point by combining the suspected noise factor and the fluctuation possibility.
[0049] In one embodiment, the noise probability of each voltage data point is determined based on the following: The normalized value of the ratio of the suspected noise factor of each voltage data point to the fluctuation possibility of each voltage data point is used as the noise probability of each voltage data point.
[0050] For the The noise probability is calculated as follows:
[0051] In this formula, For the The noise probability of a voltage data point, Indicates the The suspected noise factor of the voltage data point is For the The fluctuation probability of each voltage data point, To normalize the function, the calculation results are uniformly mapped to the standard interval so that they have clear probabilistic meaning.
[0052] In this formula, The smaller the The more stable the voltage state of the UV curing machine is at the corresponding moment of each voltage data point, the less likely it is that the data will fluctuate. The bigger, the The larger the value, the more likely it is that it is caused by noise rather than data fluctuation. The greater the noise probability of each voltage data point, the The more voltage data points there are, the more likely they are noise data. The larger the At the corresponding moment of the voltage data point, the voltage state of the UV curing machine is relatively unstable. The greater the possibility of data fluctuation, the more likely it is to cause the first The smaller the noise probability of each voltage data point, the smaller the noise probability of each voltage data point.
[0053] When the equipment is running stably, The value of is very small (close to 0), and it will be significantly magnified as the denominator in the formula. This is completely consistent with engineering logic: in a stable state, even a tiny voltage jump should be given high attention, and the probability of it being noise should be higher.
[0054] When the device is started or the power is adjusted. The value of is large (close to 1), The amplification effect of the voltage fluctuation is greatly weakened or even suppressed, which makes the expected voltage fluctuations generated in such stages calculated The value is relatively low, which effectively avoids misjudging normal fluctuations as noise.
[0055] In this way, the present invention no longer views each data transition statically and in isolation, but instead comprehensively evaluates the data point within the dynamic context of the time and operating conditions at which it occurred. This dynamic evaluation mechanism, based on operating conditions, improves the accuracy and robustness of noise identification, laying a solid data foundation for subsequent precise cluster analysis and reliable fault diagnosis.
[0056] S3: Optimize the agglomerative hierarchical clustering algorithm based on the noise probability of each voltage data point.
[0057] Specifically, it is to optimize the basis of determining the two clusters when they are aggregated in a single link during the recursive aggregation process of the agglomerative hierarchical clustering algorithm.
[0058] like Figure 3 As shown, the following steps are included: S31: Calculate the score value of each group of voltage data points in each cluster.
[0059] Through step S2, the noise probability of each voltage data point is obtained, and the voltage data points contained in each cluster are combined to obtain multiple groups of voltage data points.
[0060] A lower noise probability for a set of voltage data points indicates a more reliable set of voltage data points, making it more likely to serve as the basis for clustering the two clusters it belongs to, and thus resulting in a higher score. The distance between a set of data points refers to the absolute value of the difference between the two data points. A smaller distance indicates a smaller difference between the two data points, and thus a higher likelihood that the data points serve as the basis for clustering the two clusters it belongs to, and thus a higher score.
[0061] In one embodiment, the score value of each set of voltage data points in each cluster is calculated based on the following formula:
[0062] Where, For the The cluster The score value of the voltage data point of the group is higher. The more suitable the group of voltage data points is as the first A set of voltage data points that clusters are based on when a single link is aggregated. and Respectively The cluster The noise probability of the first voltage data point and the noise probability of the second voltage data point in a group of voltage data points. The noise probability reflects the reliability of the voltage data point. The higher the noise probability, the greater the possibility that the voltage data point is interfered by noise, and the lower the reliability of the voltage data point; conversely, the lower the noise probability, the more reliable the voltage data point. For the The cluster The distance between two voltage data points in a group of voltage data points, is the natural exponential function.
[0063] In this formula, is the sum of the noise probabilities. When the sum of the noise probabilities is small, it means that the two noise data points are relatively reliable. The higher the score values of the two noise data points, the greater the possibility of being used as the basis for aggregation. and showed a negative correlation, The smaller it is, the closer the two noise data points are, and the greater the possibility of them being used as the basis for aggregation. and It also shows a negative correlation. Therefore, Partially constructed through negative exponential functions and The negative correlation, and and negative correlation.
[0064] The formula is used to obtain the score value of each group of voltage data points in each cluster. In this way, the likelihood of different groups of voltage data points being used as aggregation basis can be intuitively compared through the size of the score value, which facilitates the subsequent selection of voltage data points for each group of clusters in single-link aggregation.
[0065] S32: Determine, based on the score value, a set of voltage data points on which each group of clusters is based when a single link is aggregated.
[0066] The scores of all voltage data points of each cluster are calculated, and among all the voltage data points, the one with the highest score is selected as the voltage data point for the cluster during single link aggregation.
[0067] For example, A and B are a group of clusters, the voltage data point A1 in A and the voltage data point B1 in B are a group of voltage data points with a score of 0.8, the voltage data point A2 in A and the voltage data point B2 in B are a group of voltage data points with a score of 0.5, the voltage data point A3 in A and the voltage data point B3 in B are a group of voltage data points with a score of 0.6, then the group of voltage data points A1 and B1 is used as the group of voltage data points based on which A and B are aggregated in a single link.
[0068] According to this method, a set of voltage data points based on which any set of clusters are aggregated under a single link is obtained.
[0069] S33: Calculate inter-cluster distances based on a set of voltage data points and perform a clustering operation.
[0070] S331: Obtain a set of voltage data points based on which each cluster is aggregated in a single link.
[0071] For example, a set of voltage data points A1 and B1 is obtained. This set of voltage data points is a set of voltage data points based on which A and B are aggregated in a single link.
[0072] S332: Calculate the inter-cluster distance.
[0073] The absolute value of the difference between two voltage data points included in a set of voltage data points on which each cluster is based during single link aggregation is used as the distance between the two clusters included in the set of clusters.
[0074] For example, A and B are a group of clusters, the voltage data point A1 in A and the voltage data point B1 in B are a group of voltage data points, and the group of voltage data points based on which A and B are aggregated in a single link is A1 and B1. The absolute value of the difference between A1 and B1 is used as the distance between A and B.
[0075] According to this method, the distance between any two clusters included in any group of clusters, that is, the distance between any two clusters, can be obtained and recorded.
[0076] S333: Recursive aggregation and generation of clustering result tree.
[0077] During the recursive aggregation process, for each cluster, the distances between it and all other clusters are sorted. The cluster closest to it is selected and fused to form a new cluster. After fusion is complete, step S1 (continuously collecting voltage data points during UV curing machine operation) is repeated to step S3, where the distances between the fused new cluster and the remaining clusters are calculated. The new cluster is then fused with the cluster closest to it, and this process is repeated. This step-by-step fusion ultimately generates a complete hierarchical clustering tree, which illustrates the merging relationships between clusters and the hierarchical structure of the clusters throughout the clustering process.
[0078] like Figure 4 As shown: Initial state: There are 5 clusters, namely C1, C2, C3, C4, and C5.
[0079] First iteration: Taking any two clusters as a group, obtain the set of voltage data points used for each cluster during single-link aggregation. Based on this set of voltage data points, calculate the inter-cluster distances to determine the distances between each cluster and the others. The distance between C1 and C2 is the smallest, indicating the closest association, so C1 and C2 are merged into a new cluster, denoted as C12.
[0080] Second iteration: The remaining clusters are C12, C3, C4, and C5. The voltage data points used for each cluster during single-link aggregation are again obtained. Inter-cluster distances are calculated based on these voltage data points to determine the distance between each cluster and the others. At this point, the distance between C4 and C5 is the smallest, so C4 and C5 are merged into a new cluster, designated C45.
[0081] Third iteration: The remaining clusters are now C12, C3, and C45. We continue to obtain the set of voltage data points used for each cluster during single-link aggregation. We calculate inter-cluster distances based on these voltage data points to determine the distances between each cluster and the others. At this point, the distance between C3 and C45 is the smallest. Therefore, we merge C3 and C45 into a new cluster, designated C345.
[0082] Fourth round of iteration: The remaining clusters are C12 and C345. The distance between them is calculated and they are merged into a final cluster, recorded as C12345. This final cluster is also the root node of the hierarchical clustering result tree.
[0083] Through this iterative fusion process, the final hierarchical clustering result tree is formed, and the distance between each two fused clusters is recorded, showing the process of gradually merging into clusters of different levels.
[0084] S4: Use the optimized agglomerative hierarchical clustering algorithm to perform anomaly detection on the voltage data points during the operation of the UV curing machine.
[0085] Specifically, the optimized agglomerative hierarchical clustering algorithm is used to cluster the voltage data points during the operation of the UV curing machine to obtain a hierarchical clustering result tree. Each non-leaf node of the tree is a fused new cluster, which is formed by the fusion of the clusters corresponding to its child nodes. Each cluster contains multiple voltage data points.
[0086] In the agglomerative hierarchical clustering algorithm, the preset distance threshold is used to screen and judge the clustering results. This is a common method for identifying abnormal clusters: During the recursive aggregation process in the agglomerative hierarchical clustering algorithm, a distance threshold of 0.6 (empirical value) is preset. If the inter-cluster distance between a cluster and another cluster is less than 0.6 when they are merged, the merged relationship between the cluster and the other cluster is retained. The merged cluster is then treated as a single cluster for subsequent analysis, and this cluster is considered a normal cluster. Because when the inter-cluster distance is less than the distance threshold, it means that the clusters are relatively similar in characteristics or attributes, conforming to the clustering pattern of the overall data, and therefore can be retained and treated as normal clusters.
[0087] If the inter-cluster distance between a cluster and another cluster is greater than or equal to 0.6 when merged, the merged relationship between the cluster and the other cluster is severed, and the two clusters are separated and treated as two suspected anomaly clusters. Because when the inter-cluster distance is greater than or equal to the threshold, it indicates that the two clusters have significant differences in characteristics or attributes from the other clusters and do not conform to the clustering pattern of the overall data. Therefore, they are separated and treated as suspected anomaly clusters for further analysis.
[0088] This operation of identifying suspected abnormal clusters by using distance thresholds can quickly and effectively filter out data that may contain abnormalities from a large number of clustering results, which helps to further explore the causes of these anomalies.
[0089] S5: Evaluate whether there is a voltage fault in the UV curing machine based on the abnormality detection result.
[0090] During the operation of the UV curing machine, cluster analysis and anomaly detection are performed on the collected voltage data points to determine whether there is a voltage fault.
[0091] The specific steps and judgment criteria are as follows: First, identify abnormal voltage data points: a threshold of 5 (empirical value) is preset. After completing cluster analysis of the voltage data points and preliminary identification of suspected abnormal clusters, further screening is performed on each suspected abnormal cluster. If the number of voltage data points in a suspected abnormal cluster is greater than 5, it is considered a pseudo-anomaly caused by data fluctuations or other non-critical factors. If the number of voltage data points in a suspected abnormal cluster is less than or equal to 5, the remaining suspected abnormal clusters are officially identified as abnormal clusters.
[0092] Because agglomerative hierarchical clustering algorithms merge clusters based on inter-cluster distance, when a cluster has very few data points, it may be relatively isolated in the cluster space, distant from other clusters. This means that the data points in this cluster differ significantly from the majority of other data points in terms of characteristics or attributes, and may not conform to the overall data distribution pattern. If the fluctuations were random, the data points would be more evenly distributed across the clusters, rather than concentrated in a few isolated small clusters.
[0093] Next, evaluate whether there is a voltage fault in the UV curing machine.
[0094] The preset percentage threshold is 2% (empirical value). If the number of abnormal voltage data points exceeds 2% of the total number of voltage data points in the M minutes before the current analysis time (when abnormal voltage data points are currently being analyzed), the UV curing machine is deemed to have a voltage fault and personnel are notified for inspection. If the number of abnormal voltage data points does not exceed 2% of the total number of voltage data points in the M minutes before the current analysis time, the UV curing machine is deemed to have no voltage fault. M is a preset value of 5 (empirical value). This dynamic assessment method based on a time window takes into account the dynamic changes in operating status.
[0095] By monitoring recent data, real-time fluctuations in the UV curing machine's voltage can be captured. If the proportion of abnormal data points exceeds the threshold within a short period of time, it indicates that a problem may have occurred with the UV curing machine, requiring timely intervention and notification of relevant personnel for inspection. This can prevent further damage to the UV curing machine from continuing to operate, ensuring its safe and stable operation.
Claims
1. A UV curing machine fault detection method based on data analysis, characterized in that: include: Continuously collect voltage data points of the UV curing machine during operation and divide the surrounding data segments of each voltage data point; Determine the suspected noise factor of each voltage data point based on the difference between each voltage data point and the mean of the surrounding data segments, and the difference between each voltage data point and the left and right adjacent voltage data points; determining a noise probability for each voltage data point based on a suspected noise factor for each voltage data point and a fluctuation possibility for each voltage data point; The voltage data points are clustered using an agglomerative hierarchical clustering algorithm to obtain multiple clusters. In the process of recursively aggregating the multiple clusters, each two clusters are regarded as a group of clusters. One voltage data point is selected from each of the two clusters in each group to form a group of data points. The score value of each group of voltage data points in each cluster is determined based on the noise probability and distance of each group of data points. According to the score value of each group of voltage data points in each cluster, the group of voltage data points that each cluster is based on during single-link aggregation is determined to optimize the agglomerative hierarchical clustering algorithm; the optimized agglomerative hierarchical clustering algorithm is used to perform anomaly detection on the voltage data points of the UV curing machine during operation, and whether the UV curing machine has a voltage fault is determined based on the anomaly detection results.
2. The UV curing machine fault detection method according to claim 1, characterized in that: The suspected noise factor for each voltage data point is determined based on the following method: Calculate the The absolute value of the difference between the voltage data point and its adjacent data point on the left, and the The absolute value of the difference between a voltage data point and its adjacent data point on the right, and the absolute value of the larger difference is recorded as ; No. The suspected noise factor of the voltage data point , where For the The absolute value of the difference between a voltage data point and the mean of the surrounding data segments, is a hyperparameter.
3. The UV curing machine fault detection method according to claim 1, characterized in that: The noise probability of each voltage data point is determined based on the following method: The normalized value of the ratio of the suspected noise factor of each voltage data point to the fluctuation possibility of each voltage data point is used as the noise probability of each voltage data point.
4. The UV curing machine fault detection method according to claim 1, characterized in that: The fluctuation probability of each voltage data point is determined based on the following method: Obtain the time interval between the corresponding moment of each voltage data point and the start-up moment of this UV curing machine, as well as the current data point at the corresponding moment of each voltage data point; Calculate the probability of fluctuation for each voltage data point: , where For the The fluctuation probability of each voltage data point, For the The time interval between the corresponding moment of each voltage data point and the start-up moment of this UV curing machine, For the The absolute value of the difference between the current data point at the moment corresponding to the voltage data point and the reference current value of the voltage data point, is the natural exponential function.
5. The UV curing machine fault detection method according to claim 1, characterized in that: The score value of each voltage data point in each cluster is calculated based on the following formula: ; Where, For the The cluster The score value of the group voltage data point, and Respectively The cluster The noise probability of the first and second voltage data points in the group of voltage data points, For the The cluster The distance between two voltage data points in a group of voltage data points, is a natural exponential function, where the distance between two voltage data points is determined based on the absolute value of the difference between the two voltage data points.
6. The UV curing machine fault detection method according to claim 1, characterized in that: The method for determining a set of voltage data points for each cluster in single link aggregation is as follows: The scores of all voltage data points of each cluster are calculated, and the voltage data point with the highest score is selected as the voltage data point based on which the cluster is aggregated in a single link.
7. The UV curing machine fault detection method according to claim 1, characterized in that: Perform abnormal detection on voltage data points of UV curing machine during operation, including: In the recursive aggregation process of the agglomerative hierarchical clustering algorithm, a distance threshold is preset. If the inter-cluster distance between a cluster and another cluster when they are merged is less than the distance threshold, the cluster after the merger is regarded as a normal cluster; if the inter-cluster distance between a cluster and another cluster when they are merged is greater than or equal to the distance threshold, the cluster and the other cluster are separated and regarded as two suspected abnormal clusters. A preset number threshold; if the number of voltage data points of a suspected abnormal cluster is greater than the preset number threshold, the suspected abnormal cluster is removed, and the remaining suspected abnormal cluster is regarded as an abnormal cluster, and the voltage data points contained in the abnormal cluster are regarded as abnormal voltage data points to complete the abnormal detection of the voltage data points.
8. The UV curing machine fault detection method according to claim 1, characterized in that: The method to determine whether the UV curing machine has a voltage fault based on the abnormal detection results is as follows: A preset percentage threshold is set; if the ratio of the number of abnormal voltage data points to the total number of voltage data points exceeds the percentage threshold within M minutes before the current analysis time, it is determined that a voltage fault has occurred in the UV curing machine, and relevant personnel are notified to conduct an inspection; If the ratio of the number of abnormal voltage data points to the total number of voltage data points within M minutes before the current analysis time does not exceed the ratio threshold, it is determined that the UV curing machine has no voltage fault; M is a preset value.
9. The UV curing machine fault detection method according to claim 1, characterized in that: The method of dividing the surrounding data segments of each voltage data point is: Taking each voltage data point as the center, a preset number of voltage data points are selected on both sides of the voltage data point to form a data segment surrounding the voltage data point.
10. The UV curing machine fault detection method according to claim 1, characterized in that: The reference current value of the voltage data point is determined based on the following method: For any voltage data point, all historical voltage data points consistent with the voltage data point are selected, and the mean of all historical current data points corresponding to all the selected historical voltage data points is calculated as the reference current value of the voltage data point.
Citation Information
Patent Citations
Photovoltaic transformer electrical fault detection method
CN115982602A
Data anomaly detection method for machine filter cloth adhesion
CN116821833A
Mining water pump intelligent monitoring system based on multiple sensors
CN117195018A
Low-power resistor electric propulsion voltage dynamic adjusting method and system
CN117908615A
Distribution transformer fault detection method and system
CN119046854A