Working condition category analysis method and device, electronic equipment and storage medium
Patent Information
- Application Number
- CN202310354939.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-04
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2043-04-04
AI Technical Summary
[0003]基于传感器数据驱动下的工业过程监控、工况类别识别任务中,由于工业过程存在着多生产模式下动态切换过程,工艺流程复杂、智能、级联、集成,因此基于模型与机理的建模方法无法有效地监控工业过程和识别工况类别,仍然需要大量人工进行校准,导致工况类别的识别效率较低
[0087]在获取到工业时序数据后,根据工业时序数据的分布信息,对工业时序数据进行聚类分析,得到第一类簇和第一类簇所属的工况类别。基于工业时序数据的时序相关性,搜索在时间维度下满足时序变化但未在该第一类簇内的点位数据,即目标点位数据,并将目标点位数据添加至第一类簇,得到第二类簇。在空间维度,提取两个第一类簇中的共有点位数据,并在共有点位数据的数量满足预设条件的情况下,提取出共有点位数据作为新的第三类簇,并将第一类簇中除共有点位数据外的点位数据作为第四类簇。将时间维度下的第二类簇和/或空间维度下的第三类簇和第四类簇作为候选类簇,并从候选类簇中类簇区分度较大的类簇作为最终类簇,并获取最终类簇所属的工况类别,类簇区分度用于表征候选类簇的紧密程度以及与其他候选类簇的分散程度。通过上述方法可以得到更准确的最终类簇,从而对工业时序数据进行更准确的标注,提高了工况类别识别的效率和准确性。在工业过程工况类别分析的方法中,有监督的学习方法能够更好地发掘工况根原因,提升工业过程工况类别识别与预警能力。本申请实施例可以为有监督学习模型注入大量的已标注的数据样本提供有效的数据支撑。
Smart Images

Figure CN116304761B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of industrial production technology, and in particular to a method, apparatus, electronic device and storage medium for analyzing operating conditions. Background Technology
[0002] In recent years, industrial accidents have occurred frequently, causing not only economic losses to enterprises but also potentially endangering human lives. Therefore, monitoring the entire lifecycle of industrial processes and identifying operating conditions are particularly important. Currently, industrial-grade sensors can monitor and control the production process and operational status throughout the entire lifecycle of industrial processes.
[0003] In sensor-data-driven industrial process monitoring and condition classification tasks, industrial processes involve dynamic switching between multiple production modes, and are complex, intelligent, cascading, and integrated. Therefore, model- and mechanism-based modeling methods cannot effectively monitor industrial processes and identify condition categories, still requiring significant manual calibration, resulting in low efficiency in condition category identification. Furthermore, model- and mechanism-based modeling methods suffer from complex modeling, unreasonable knowledge linkage, and low generalization ability, often leading to frequent false alarms during monitoring due to external fluctuations and human error, resulting in low accuracy in condition category identification. Summary of the Invention
[0004] To address the aforementioned technical problems, this application provides a method, apparatus, electronic device, and storage medium for analyzing operating conditions.
[0005] According to a first aspect of this application, a method for analyzing operating conditions is provided, comprising:
[0006] Acquire industrial time-series data during the industrial production process, and perform cluster analysis on the industrial time-series data based on the distribution information of the industrial time-series data to obtain the first cluster and the operating condition category to which the first cluster belongs;
[0007] Target point data that is temporally correlated with the point data in the first cluster is obtained from the point data outside each first cluster, and the target point data is added to the first cluster to obtain the second cluster;
[0008] From the two first-category clusters, filter out the common point data that belong to both first-category clusters;
[0009] If the number of shared location data meets the preset conditions, the shared location data is regarded as the third cluster, and the location data after removing the shared location data from the two first clusters is regarded as the fourth cluster.
[0010] The set of the third and fourth clusters and / or the second cluster are used as candidate clusters, and the cluster discrimination degree of each candidate cluster is calculated according to a preset formula. The cluster discrimination degree is used to characterize the compactness of the candidate clusters and the degree of dispersion from other candidate clusters.
[0011] Candidate clusters with a cluster discrimination score greater than the discrimination score threshold are taken as the final clusters, and the working condition category to which the final clusters belong is obtained.
[0012] The process involves deleting point data belonging to the final cluster from the industrial time series data to update the industrial time series data, returning the distribution information based on the industrial time series data, and performing cluster analysis on the industrial time series data until no point data is contained in the industrial time series data.
[0013] Optionally, after obtaining the first cluster and the operating condition category to which the first cluster belongs, the method further includes:
[0014] Select target clusters from the first clusters whose density is greater than the density threshold;
[0015] The target point data obtained from point data outside each of the first clusters, which has a temporal correlation with the point data in the first cluster, includes:
[0016] Target point data that is temporally correlated with the point data in the target clusters is obtained from point data outside each target cluster;
[0017] From the two first-category clusters, filter the common location data that belong to both first-category clusters, including:
[0018] From the two target clusters, filter out the common point data that belong to both target clusters;
[0019] The step of removing the common point data from the two first clusters and then defining them as the fourth cluster includes:
[0020] The point data after removing the common point data from the two target clusters are respectively used as the fourth cluster.
[0021] Optionally, after acquiring industrial time-series data from the industrial production process, the method further includes:
[0022] The industrial time-series data is subjected to dimensionality reduction processing to obtain dimensionality-reduced data;
[0023] The clustering analysis of industrial time-series data based on distribution information includes:
[0024] Based on the distribution information of the industrial time-series data, cluster analysis is performed on the dimensionality-reduced data.
[0025] Optionally, the target point data obtained from point data outside each of the first clusters, which has a temporal correlation with the point data in the first cluster, includes:
[0026] Based on the data collection time of each point in each of the first clusters, determine the number of consecutive point data and the number of subsequent consecutive disconnected point data corresponding to the time when each point data is located.
[0027] The target time is determined based on the number of consecutive data points corresponding to the time of each data point and the number of subsequent consecutive disconnected data points.
[0028] For each target time, the target search points adjacent to the target time are used as target point data according to the time dimension.
[0029] Optionally, determining the target time based on the number of consecutive data points corresponding to the time of each data point and the number of subsequent consecutive disconnected data points includes:
[0030] In the set of consecutive point data corresponding to the time of each point data, the number of consecutive point data that is less than the first preset quantile is selected as the candidate consecutive point data.
[0031] The number of subsequent consecutive disconnected point data corresponding to the time when the number of candidate consecutive point data is located is taken as the first candidate number of subsequent consecutive disconnected point data.
[0032] Select a second candidate number of consecutive disconnection points that is less than the second preset quantile from the set of data of the first candidate subsequent consecutive disconnection points.
[0033] The time corresponding to the number of subsequent consecutive disconnection points in the second candidate is determined as the target time; or...
[0034] In the set of subsequent consecutive disconnection point data corresponding to the time of each point data, the number of subsequent consecutive disconnection point data that is less than the third preset quantile is selected as the target number of subsequent consecutive disconnection point data.
[0035] The time corresponding to the number of consecutive disconnected data points of the target is determined as the target time.
[0036] Optionally, the step of filtering common location data belonging to both first clusters includes:
[0037] For two first-class clusters, calculate the Euclidean distance between each data point in one first-class cluster and each data point in the other first-class cluster;
[0038] Two points whose Euclidean distance is less than the distance threshold are identified as shared points of the two first clusters.
[0039] Optionally, the method for determining whether the number of shared location data meets the preset conditions includes:
[0040] Perform statistical analysis on the first cluster to determine the number of long-tail data points in the first cluster;
[0041] The ratio of the number of common data points to the number of long-tail data points is determined as the cluster common purity.
[0042] If the common purity of the cluster is greater than the purity threshold, then the number of common point data is determined to meet the preset condition.
[0043] If the common purity of the cluster is less than or equal to the purity threshold, then the number of common point data is determined to be insufficient.
[0044] Optionally, calculating the cluster distinguishability of each candidate cluster according to a preset formula includes:
[0045] The local density percentage of each candidate cluster is determined based on the local density algorithm.
[0046] Calculate the silhouette coefficient and CH index for each of the candidate clusters;
[0047] The cluster discrimination of each candidate cluster is determined based on the local density ratio, the contour coefficient, and the CH index.
[0048] According to a second aspect of this application, a working condition category analysis device is provided, comprising:
[0049] The industrial time-series data acquisition module is used to acquire industrial time-series data during the industrial production process.
[0050] The clustering analysis module is used to perform clustering analysis on industrial time series data based on the distribution information of industrial time series data, and obtain the first cluster and the operating condition category to which the first cluster belongs;
[0051] The target point data acquisition module is used to acquire target point data that has a temporal correlation with the point data in the first cluster from the point data outside each of the first clusters;
[0052] The first cluster update module is used to add the target point data to the first cluster to obtain the second cluster;
[0053] The common location data determination module is used to filter common location data that simultaneously belong to both of the two first type clusters from the two first type clusters;
[0054] The third cluster determination module is used to identify the common point data as a third cluster when the number of common point data meets a preset condition.
[0055] The fourth cluster determination module is used to determine the point data after removing the common point data from the two first clusters as the fourth cluster;
[0056] The cluster discrimination determination module is used to take the set of the second cluster and / or the third cluster and the fourth cluster as candidate clusters, and calculate the cluster discrimination of each candidate cluster according to a preset formula. The cluster discrimination is used to characterize the density of the candidate clusters and the degree of dispersion from other candidate clusters.
[0057] The final cluster determination module is used to select candidate clusters with a cluster discrimination degree greater than the discrimination degree threshold as the final clusters and obtain the working condition category to which the final clusters belong.
[0058] The industrial time series data update module is used to delete the point data belonging to the final cluster in the industrial time series data in order to update the industrial time series data.
[0059] The return module is used to return the clustering analysis module to perform the steps of clustering analysis on the industrial time series data based on the distribution information of the industrial time series data, until the industrial time series data does not contain any point data.
[0060] Optionally, the working condition category analysis device further includes:
[0061] The target cluster selection module is used to select target clusters with a density greater than a density threshold from the first clusters;
[0062] The target point data acquisition module is specifically used to acquire target point data that has a temporal correlation with the point data in the target cluster from the point data outside each target cluster;
[0063] The common location data determination module is specifically used to filter common location data that belong to both target clusters from the two target clusters.
[0064] The fourth cluster determination module is specifically used to take the point data after removing the common point data from the two target clusters as the fourth cluster.
[0065] Optionally, the working condition category analysis device further includes:
[0066] The dimensionality reduction module is used to perform dimensionality reduction processing on the industrial time-series data to obtain dimensionality-reduced data;
[0067] The clustering analysis module is specifically used to perform clustering analysis on the dimensionality-reduced data based on the distribution information of the industrial time-series data, to obtain a first cluster and the operating condition category to which the first cluster belongs.
[0068] Optionally, the target point data acquisition module is specifically used to determine the number of consecutive point data corresponding to the time when each point data is collected in each of the first clusters and the number of subsequent consecutive disconnected point data based on the collection time of each point data; determine the target time based on the number of consecutive point data corresponding to the time when each point data is collected and the number of subsequent consecutive disconnected point data; and for each target time, select the target search number of point data adjacent to the target time as target point data according to the time dimension.
[0069] Optionally, the target point data acquisition module is specifically used to determine the target time by means of the following steps: the number of consecutive point data corresponding to the time when each point data is located and the number of subsequent consecutive disconnected point data.
[0070] In the set of consecutive point data corresponding to the time of each point data, the number of consecutive point data that is less than the first preset quantile is selected as the candidate consecutive point data.
[0071] The number of subsequent consecutive disconnected point data corresponding to the time when the number of candidate consecutive point data is located is taken as the first candidate number of subsequent consecutive disconnected point data.
[0072] Select a second candidate number of consecutive disconnection points that is less than the second preset quantile from the set of data of the first candidate subsequent consecutive disconnection points.
[0073] The time corresponding to the number of subsequent consecutive disconnection points in the second candidate is determined as the target time; or...
[0074] In the set of subsequent consecutive disconnection point data corresponding to the time of each point data, the number of subsequent consecutive disconnection point data that is less than the third preset quantile is selected as the target number of subsequent consecutive disconnection point data.
[0075] The time corresponding to the number of consecutive disconnected data points of the target is determined as the target time.
[0076] Optionally, the common point data determination module is specifically used to calculate the Euclidean distance between each point data in one of the first clusters and each point data in the other first cluster for the two first clusters; and determine the two point data whose corresponding Euclidean distance is less than the distance threshold as the common point data of the two first clusters.
[0077] Optionally, the third cluster determination module is specifically used to determine whether the number of common location data meets a preset condition through the following steps:
[0078] Perform statistical analysis on the first cluster to determine the number of long-tail data points in the first cluster;
[0079] The ratio of the number of common data points to the number of long-tail data points is determined as the cluster common purity.
[0080] If the common purity of the cluster is greater than the purity threshold, then the number of common point data is determined to meet the preset condition.
[0081] If the common purity of the cluster is less than or equal to the purity threshold, then the number of common point data is determined to be insufficient.
[0082] Optionally, the cluster discrimination determination module is specifically used to take the set of the second cluster and / or the third cluster and the fourth cluster as candidate clusters, and determine the local density ratio of each candidate cluster according to the local density algorithm; calculate the contour coefficient and CH index of each candidate cluster; and determine the cluster discrimination of each candidate cluster according to the local density ratio, the contour coefficient and the CH index.
[0083] According to a third aspect of this application, an electronic device is provided, comprising: a processor configured to execute a computer program stored in a memory, wherein the computer program, when executed by the processor, implements the method described in the first aspect.
[0084] According to a fourth aspect of this application, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the method described in the first aspect.
[0085] According to a fifth aspect of this application, a computer program product is provided that, when the computer program product is run on a computer, causes the computer to perform the method described in the first aspect.
[0086] The technical solution provided in this application has the following advantages compared with the prior art:
[0087] After acquiring industrial time-series data, cluster analysis is performed on the data based on its distribution information to obtain a first cluster and the operating condition category to which it belongs. Based on the temporal correlation of the industrial time-series data, location data that meet the temporal variation criteria but are not within the first cluster (i.e., target location data) are searched and added to the first cluster to obtain a second cluster. In the spatial dimension, common location data between the two first clusters is extracted. If the number of common location data meets a preset condition, this common location data is extracted as a new third cluster, and the location data in the first cluster excluding the common location data is used as a fourth cluster. The second cluster in the temporal dimension and / or the third and fourth clusters in the spatial dimension are used as candidate clusters. The cluster with the higher cluster discrimination among the candidate clusters is selected as the final cluster, and the operating condition category to which the final cluster belongs is obtained. Cluster discrimination is used to characterize the density of candidate clusters and their dispersion from other candidate clusters. The above methods yield more accurate final clusters, enabling more precise labeling of industrial time-series data and improving the efficiency and accuracy of operational condition category identification. In industrial process operational condition category analysis, supervised learning methods are better able to uncover the root causes of operational conditions, enhancing the ability to identify and warn of industrial process operational condition categories. The embodiments of this application can provide effective data support by injecting a large number of labeled data samples into the supervised learning model. Attached Figure Description
[0088] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0089] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0090] Figure 1 This is a flowchart of a working condition category analysis method in an embodiment of this application;
[0091] Figure 2 This is a schematic diagram of the clustering results in an embodiment of this application;
[0092] Figure 3 This is a schematic diagram of the clustering results in the time dimension in an embodiment of this application;
[0093] Figure 4 This is a schematic diagram of the clustering results after adding target point data in the time dimension in an embodiment of this application.
[0094] Figure 5 This is another flowchart of the working condition category analysis method in the embodiments of this application;
[0095] Figure 6 This is a schematic diagram of a working condition category analysis device in an embodiment of this application;
[0096] Figure 7 This is a schematic diagram of the structure of an electronic device in an embodiment of this application. Detailed Implementation
[0097] To better understand the above-mentioned objectives, features, and advantages of this application, the solution of this application will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described in these embodiments can be combined with each other.
[0098] Many specific details are set forth in the following description in order to provide a full understanding of this application, but this application may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some embodiments of this application, and not all embodiments.
[0099] To address the challenges of numerous monitoring and control instruments, noise from a large number of interfering variables, high costs of manual annotation, and difficulties in algorithm annotation during industrial production, this application provides a method, apparatus, electronic device, and storage medium for analyzing operating conditions. Based on the characteristics of industrial time-series data in the time and / or spatial dimensions of industrial production processes, it accurately classifies the operating conditions of each process stage of the industrial time-series data, thereby improving the efficiency and accuracy of operating condition identification.
[0100] See Figure 1 , Figure 1 This is a flowchart of a working condition category analysis method in an embodiment of this application, which may include the following steps:
[0101] Step S102: Obtain industrial time-series data during the industrial production process.
[0102] In industrial production processes, equipment can continuously collect data from various points along the process to obtain industrial time-series data. Each data point has a corresponding collection time. For example, pressure, temperature, and production load can be collected within a preset period (e.g., 6 months, 8 months, or 10 months). The data points can include various operating condition categories, such as pressure, temperature, and production load. For each operating condition category, data points can be collected at least once every preset time interval (e.g., 10 seconds).
[0103] Step S104: Based on the distribution information of industrial time series data, perform cluster analysis on the industrial time series data to obtain the first cluster and the operating condition category to which the first cluster belongs.
[0104] Optionally, after acquiring industrial time-series data, the data can be cleaned and aligned before performing cluster analysis to reduce computational load and improve efficiency. For example, missing and out-of-limit values can be filtered out from the industrial time-series data, and strongly correlated variables and abnormal operating conditions can be removed. Then, the data can be standardized and denoised.
[0105] In some embodiments, density clustering analysis learning methods that determine classification through data distribution can be used to determine the labels of industrial time-series data, that is, to determine the operating condition category to which the industrial time-series data belongs. For example, the dbscan (Density-Based Spatial Clustering of Applications with Noise) clustering algorithm can be used to perform cluster analysis on the industrial time-series data to obtain the first cluster and the operating condition category to which the first cluster belongs, thereby completing the basic full labeling of the operating condition category of the industrial time-series data.
[0106] See Figure 2 , Figure 2 This is a schematic diagram of the clustering results in an embodiment of this application. It can be seen that after clustering the industrial time series data, multiple first-class clusters are obtained. The density of the data points contained in each first-class cluster is different, and the degree of dispersion between the first-class clusters and other first-class clusters is also different.
[0107] Step S106: Obtain target point data that has a temporal correlation with the point data in the first cluster from the point data outside each first cluster, and add the target point data to the first cluster to obtain the second cluster.
[0108] Industrial time-series data typically exhibits a certain degree of temporal correlation. This application embodiment can search for target data points that satisfy temporal changes but are not within a specific first cluster based on these correlations. In other words, a data point may belong to a certain first cluster in the time dimension, but was not clustered into that first cluster during the aforementioned clustering analysis. In this case, the data point can be added to that first cluster to improve the accuracy of the clustering results in the time dimension.
[0109] In one optional implementation, the number of consecutive data points corresponding to the time of data collection for each data point in each first cluster and the number of subsequent consecutively disconnected data points can be determined based on the data collection time of each data point. As mentioned earlier, a preset time period can be set when collecting data points. If the next data point is collected within the preset time period after collecting data from one data point, then that data point is continuous with the previous data point; if the next data point is collected after the preset time period, then that data point is disconnected from the previous data point. Therefore, the number of consecutive data points corresponding to the time of data collection for each data point can be determined, that is, the total number of currently continuous data points. The number of subsequent consecutively disconnected data points refers to the total number of consecutively disconnected data points after the time of data collection for each data point. The number of consecutive data points corresponding to the time of data collection for each data point and the number of subsequent consecutively disconnected data points can be found in Table 1.
[0110] Table 1
[0111]
[0112]
[0113] It can be seen that when the point data is continuous, the number of consecutive point data points gradually increases. When the point data is interrupted, the number of consecutive point data points is recalculated. When subsequent point data is continuous, the number of subsequent consecutive interrupted point data points is 0. When subsequent point data is interrupted, the number of subsequent consecutive interrupted point data points is obtained based on the number of subsequent consecutive interrupted point data points.
[0114] Next, the target time is determined based on the number of consecutive data points corresponding to the current time of each data point and the number of subsequent consecutive disconnected data points. For each target time, the target data points can be selected from the adjacent data points of the target time according to the time dimension. In other words, data points with temporal relevance to the target time are searched. For example, the target data points can be selected from the data points before the target time (target search count / 2) and the data points after the target time (target search count / 2) according to the time dimension.
[0115] Optionally, the target search quantity can be a fixed value set based on experience. Alternatively, the target search quantity can also be a value determined based on the number of subsequent consecutive disconnection point data points corresponding to the target time. For example, the target search quantity could be twice the number of subsequent consecutive disconnection point data points corresponding to the target time.
[0116] In some embodiments, the target time can be determined in the following manner:
[0117] From the set of consecutive data points corresponding to the time of each data point, the number of consecutive data points less than a first preset quantile (e.g., the 5th percentile) is selected as the candidate consecutive data point number. The number of subsequent consecutive disconnected data points corresponding to the time of the candidate consecutive data point number is selected as the first candidate subsequent consecutive disconnected data point number. From the set of the first candidate subsequent consecutive disconnected data point numbers, a second candidate subsequent consecutive disconnected data point number less than a second preset quantile is selected. The time corresponding to the second candidate subsequent consecutive disconnected data point number is determined as the target time. In other words, a smaller consecutive data point number can be selected first, and then a smaller number of subsequent consecutive disconnected data points can be selected from the number of subsequent consecutive disconnected data points corresponding to the time of the smaller consecutive data point number. The time corresponding to the selected smaller number of subsequent consecutive disconnected data points is then used as the target time.
[0118] In some embodiments, the target time can also be determined in the following ways:
[0119] From the set of subsequent consecutive disconnected data points corresponding to the time of each data point, the number of subsequent consecutive disconnected data points less than the third preset quantile is selected as the target number of subsequent consecutive disconnected data points. The time corresponding to the target number of subsequent consecutive disconnected data points is determined as the target time. That is, the time corresponding to the smaller number of subsequent consecutive disconnected data points can also be directly selected as the target time.
[0120] Both methods described above determine the target time based on the statistical distribution of location data, thus improving the accuracy of target time determination and making the resulting second cluster more suitable for industrial operating condition classification. See also... Figure 3 and Figure 4 , Figure 3 This is a schematic diagram of a first-class cluster. Figure 4 This is a schematic diagram of the second type of cluster. It can be seen that the point data in the second type of cluster are more distinguishable and better meet the requirements of industrial operating condition classification, that is, the point data are represented as interval segments.
[0121] Step S108: Filter out common location data that belong to both first-class clusters from the two first-class clusters.
[0122] It should be noted that a long-tail effect may exist within the first-class clusters, and long-tail connections may exist between two first-class clusters. In many cases, the long-tail portions can be merged into one cluster. Therefore, for any two first-class clusters, the Euclidean distance between each data point in one first-class cluster and each data point in the other first-class cluster can be calculated. The smaller the Euclidean distance, the closer the two data points are. Data points whose Euclidean distance is less than a distance threshold can be identified as shared data points between the two first-class clusters.
[0123] Step S110: If the number of common location data meets the preset conditions, the common location data is taken as the third cluster, and the location data after removing the common location data from the two first clusters are taken as the fourth cluster.
[0124] In this embodiment of the application, the greater the number of shared location data points, the greater the likelihood of grouping the shared location data points into one cluster. Therefore, when the number of shared location data points meets the preset conditions, the shared location data points can be extracted from the first cluster as the third cluster, and the remaining location data points can be used as the third cluster.
[0125] Optionally, the number of shared location data points can be determined to meet preset conditions in the following ways:
[0126] Perform statistical analysis on the first cluster to determine the number of long-tail data points within it. For example, you can count the number of data points belonging to the long tail in all first clusters and take the average of all counts as the number of long-tail data points in the first cluster. Alternatively, you can first remove the larger and smaller values from all counts and then take the average as the number of long-tail data points in the first cluster.
[0127] Next, the ratio of the number of shared data points to the number of long-tail data points is determined as the cluster shared purity. A higher cluster shared purity indicates a higher degree of coupling between the two first-class clusters, and a higher probability that the shared data points will merge into one cluster. If the cluster shared purity is greater than a purity threshold, the number of shared data points is determined to meet the preset condition; if the cluster shared purity is less than or equal to the purity threshold, the number of shared data points is determined not to meet the preset condition. The purity threshold can be, for example, 0.6, and can be set according to the actual scenario; no limitation is made here.
[0128] Step S112: Select the set of the third and fourth clusters and / or the second cluster as candidate clusters, and calculate the cluster discrimination of each candidate cluster according to the preset formula.
[0129] It should be noted that the first cluster can be processed only in the time dimension to obtain the second cluster, which is then used as a candidate cluster. Alternatively, the first cluster can be processed only in the spatial dimension to obtain the third and fourth clusters, which are then used as candidate clusters. Or, the first cluster can be processed simultaneously in both the time and spatial dimensions to obtain the third and fourth clusters, which are then used as candidate clusters. It is understood that the more candidate clusters there are, the better the final clusters can be selected from them. In this embodiment, the cluster discrimination index of the candidate clusters can be calculated using a preset formula. The cluster discrimination index characterizes the density of the candidate clusters and their dispersion from other candidate clusters; that is, the cluster discrimination index characterizes the clustering effect of the candidate clusters.
[0130] Optionally, the local density percentage of each candidate cluster can be determined using a local density algorithm. For example, the local density percentage of each candidate cluster can be calculated using the SNN (Shell Nearest Neighbor) algorithm. During the prediction process, the SNN algorithm can predict the number of target point data points in the local density of the candidate cluster based on a preset K value. The ratio of the number of target point data points to the number of all point data points in the candidate cluster is the local density percentage.
[0131] Calculate the silhouette coefficient and CH index for each candidate cluster. The silhouette coefficient is an evaluation method for clustering performance, with values ranging from -1 to 1. A value closer to 1 indicates better cohesion and separation. The CH index measures the density within a cluster by calculating the sum of squared distances between each point in the cluster and its cluster center, and the separation of the dataset by calculating the sum of squared distances between each cluster center and the dataset center. The CH index is derived as the ratio of separation to density; a higher CH index indicates denser clusters and more dispersed clusters, resulting in better clustering.
[0132] The cluster discrimination of each candidate cluster is determined based on the local density proportion, silhouette coefficient, and CH index. For example, the cluster discrimination of a candidate cluster can be determined by the sum of the product of the local density proportion and the silhouette coefficient and the CH index. Alternatively, the cluster discrimination of a candidate cluster can be determined directly by the sum of the local density proportion, silhouette coefficient, and CH index.
[0133] Step S114: Select the candidate clusters whose cluster discrimination is greater than the discrimination threshold as the final clusters, and obtain the working condition category to which the final clusters belong.
[0134] As mentioned above, a higher cluster discrimination index indicates a better clustering effect for the candidate clusters. When the cluster discrimination index of a candidate cluster is greater than the discrimination index threshold, the candidate cluster can be used as the final cluster. Based on the working condition category to which the first cluster belongs, obtained in step S104 above, the working condition category to which the final cluster belongs can be obtained.
[0135] Step S116: Delete the point data belonging to the final cluster in the industrial time series data to update the industrial time series data.
[0136] The data points in the final clusters that have been clustered are only a part of the industrial time series data. The data points in the final clusters can be deleted from the industrial time series data. Then, the clustering analysis of the remaining industrial time series data can be continued using the method described above until the clustering of all industrial time series data is completed.
[0137] Step S118: Determine whether the industrial time series data contains location data.
[0138] The industrial time series data in this step is the updated data. If the industrial time series data contains location data, return to step S104 and continue to perform cluster analysis on the remaining industrial time series data. If the industrial time series data does not contain location data, it indicates that the clustering of all industrial time series data has been completed, and the process ends.
[0139] The operating condition category analysis method of this application can perform cluster analysis on industrial time-series data based on the distribution information of industrial time-series data to obtain a first cluster and the operating condition category to which the first cluster belongs. Based on the temporal correlation of industrial time-series data, it searches for point data that meet the temporal change in the time dimension but are not in the first cluster, i.e., target point data, and adds the target point data to the first cluster to obtain a second cluster. In the spatial dimension, it extracts the common point data in the two first clusters, and if the number of common point data meets a preset condition, it extracts the common point data as a new third cluster, and the point data in the first cluster other than the common point data as a fourth cluster. The second cluster in the time dimension and / or the third and fourth clusters in the spatial dimension are used as candidate clusters, and the cluster with the higher cluster discrimination among the candidate clusters is used as the final cluster, and the operating condition category to which the final cluster belongs is obtained. The cluster discrimination is used to characterize the density of the candidate clusters and their dispersion from other candidate clusters. The above methods yield more accurate final clusters, enabling more precise labeling of industrial time-series data and improving the efficiency and accuracy of operational condition category identification. In industrial process operational condition category analysis, supervised learning methods are better able to uncover the root causes of operational conditions, enhancing the ability to identify and warn of industrial process operational condition categories. The embodiments of this application can provide effective data support by injecting a large number of labeled data samples into the supervised learning model.
[0140] See Figure 5 , Figure 5 This is another flowchart of the working condition category analysis method in the embodiments of this application, which may include the following steps:
[0141] Step S502: Obtain industrial time-series data during the industrial production process.
[0142] Step S504: Perform dimensionality reduction processing on the industrial time series data to obtain dimensionality-reduced data.
[0143] In this embodiment, dimensionality reduction algorithms can be used to reduce the dimensionality of industrial time-series data. By reducing the dimensionality of the data, the complexity of the data can be reduced, and the efficiency of data processing can be improved. For example, t-SNE and UMAP generative models can be used to reduce the dimensionality of industrial time-series data, while preserving the distribution information of the industrial time-series data itself during the dimensionality reduction process. Among them, t-SNE (t-distributed stochastic neighbor embedding) is a machine learning algorithm for dimensionality reduction. It is a non-linear dimensionality reduction algorithm suitable for reducing high-dimensional data to 2 or 3 dimensions. UMAP can retain the features of the original data to the greatest extent while significantly reducing the feature dimensionality.
[0144] Step S506: Based on the distribution information of industrial time series data, perform cluster analysis on the dimensionality-reduced data to obtain the first cluster and the operating condition category to which the first cluster belongs.
[0145] Step S508: Select target clusters from the first clusters whose density is greater than the density threshold.
[0146] Understandably, the number of first-order clusters obtained through cluster analysis can be multiple, including those with good clustering results and those with poor results. Performing time and / or spatial processing on all first-order clusters would increase computational cost. Therefore, it's advisable to first select the target clusters with good clustering results from the first-order clusters, and then perform time and / or spatial processing on these target clusters.
[0147] The clustering effect of the first type of cluster can be determined by the density. The higher the density, the better the clustering effect. Therefore, the first type of cluster with a density greater than the density threshold can be used as the target cluster.
[0148] Step S510: Obtain target point data that has a temporal correlation with the point data in the target cluster from the point data outside each target cluster, and add the target point data to the first cluster to obtain the second cluster.
[0149] Step S512: Filter the common point data that belong to both target clusters from the two target clusters.
[0150] Step S514: If the number of common location data meets the preset conditions, the common location data is taken as the third cluster, and the location data after removing the common location data from the two target clusters are taken as the fourth cluster.
[0151] Step S516: The set of the third and fourth clusters and / or the second cluster are taken as candidate clusters, and the cluster discrimination degree of each candidate cluster is calculated according to the preset formula. The cluster discrimination degree is used to characterize the tightness of the candidate clusters and the degree of dispersion from other candidate clusters.
[0152] Step S518: Select the candidate clusters whose cluster discrimination is greater than the discrimination threshold as the final clusters, and obtain the working condition category to which the final clusters belong.
[0153] Step S520: Delete the point data belonging to the final cluster in the industrial time series data in order to update the industrial time series data.
[0154] Step S522: Determine whether the industrial time series data contains location data.
[0155] The industrial time series data in this step is the updated data. If the industrial time series data contains location data, return to step S504 to continue dimensionality reduction processing on the remaining industrial time series data. If the industrial time series data does not contain location data, it indicates that the clustering of all industrial time series data has been completed, and the process ends.
[0156] The embodiments of this application and Figure 1 For details related to the embodiments, please refer to the following: Figure 1 The descriptions in the embodiments are sufficient and will not be repeated here.
[0157] The operating condition category analysis method in this application, after acquiring industrial time-series data, can first perform dimensionality reduction processing on the industrial time-series data to obtain dimensionality-reduced data, thereby reducing data complexity and improving data processing efficiency. After obtaining a first cluster through cluster analysis of the dimensionality-reduced data, a target cluster with better clustering effect can be selected from the first cluster. Time and / or spatial dimension processing is then performed on the target cluster to obtain candidate clusters. This further improves data processing efficiency. Finally, the final cluster is selected from the candidate clusters. The above-described cyclic operating condition category analysis method based on time-series correlation and / or statistical distribution improves the accuracy of operating condition category identification by increasing the accuracy of the obtained final clusters. This process does not require manual annotation, thus improving efficiency.
[0158] Corresponding to the above method embodiments, this application also provides a working condition category analysis device, see [link to relevant documentation]. Figure 6 The operating condition category analysis device 600 includes:
[0159] The industrial time-series data acquisition module 602 is used to acquire industrial time-series data during the industrial production process;
[0160] The clustering analysis module 604 is used to perform clustering analysis on industrial time series data based on the distribution information of industrial time series data, and obtain the first cluster and the operating condition category to which the first cluster belongs;
[0161] The target point data acquisition module 606 is used to acquire target point data that has a temporal correlation with the point data in the first cluster from the point data outside each first cluster;
[0162] The first cluster update module 608 is used to add the target point data to the first cluster to obtain the second cluster;
[0163] The common location data determination module 610 is used to filter common location data that belong to both first-class clusters from two first-class clusters;
[0164] The third cluster determination module 612 is used to identify the common point data as the third cluster when the number of common point data meets the preset conditions.
[0165] The fourth cluster determination module 614 is used to determine the point data after removing the common point data from the two first clusters as the fourth cluster;
[0166] The cluster discrimination determination module 616 is used to take the set of the second cluster and / or the third cluster and the fourth cluster as candidate clusters, and calculate the cluster discrimination of each candidate cluster according to a preset formula. The cluster discrimination is used to characterize the tightness of the candidate clusters and the degree of dispersion from other candidate clusters.
[0167] The final cluster determination module 618 is used to select candidate clusters with a cluster discrimination degree greater than the discrimination degree threshold as the final clusters and obtain the working condition category to which the final clusters belong.
[0168] The industrial time series data update module 620 is used to delete the point data belonging to the final cluster in the industrial time series data in order to update the industrial time series data.
[0169] Return module 622 is used to return the clustering analysis module to perform the steps of clustering analysis on the industrial time series data based on the distribution information of the industrial time series data, until the industrial time series data does not contain any point data.
[0170] Optionally, the operating condition category analysis device 600 also includes:
[0171] The target cluster selection module is used to select target clusters with a density greater than a density threshold from the first cluster.
[0172] The target point data acquisition module 606 is specifically used to acquire target point data that has a temporal correlation with the point data in the target cluster from the point data outside each target cluster;
[0173] The common location data determination module 610 is specifically used to filter common location data that belong to both target clusters from two target clusters.
[0174] The fourth cluster determination module 614 is specifically used to determine the point data after removing the common point data from the two target clusters as the fourth cluster.
[0175] Optionally, the operating condition category analysis device 600 also includes:
[0176] The dimensionality reduction module is used to reduce the dimensionality of industrial time-series data to obtain dimensionality-reduced data.
[0177] The clustering analysis module 604 is specifically used to perform clustering analysis on dimensionality-reduced data based on the distribution information of industrial time-series data, to obtain the first cluster and the operating condition category to which the first cluster belongs.
[0178] Optionally, the target point data acquisition module 606 is specifically used to determine the number of consecutive point data corresponding to the time when each point data is collected in each first cluster and the number of subsequent consecutive disconnected point data based on the collection time of each point data; determine the target time based on the number of consecutive point data corresponding to the time when each point data is collected and the number of subsequent consecutive disconnected point data; and for each target time, select the target search number of point data adjacent to the target time as target point data according to the time dimension.
[0179] Optionally, the target point data acquisition module 606 is specifically used to determine the target time by means of the following steps: the number of consecutive point data corresponding to the time when each point data is located and the number of subsequent consecutive disconnected point data.
[0180] In the set of consecutive point data corresponding to the time of each point data, the number of consecutive point data that is less than the first preset quantile is selected as the candidate consecutive point data.
[0181] The number of subsequent consecutive disconnected point data points corresponding to the time when the candidate consecutive point data points are located is taken as the first candidate subsequent consecutive disconnected point data points.
[0182] Select a number of second candidate subsequent consecutive disconnection point data that is less than the second preset quantile from the set of data consisting of the number of data of the first candidate subsequent consecutive disconnection points.
[0183] The time corresponding to the number of subsequent consecutive disconnection points in the second candidate is determined as the target time; or...
[0184] In the set of subsequent consecutive disconnection point data corresponding to the time of each point data, the number of subsequent consecutive disconnection point data that is less than the third preset quantile is selected as the target number of subsequent consecutive disconnection point data.
[0185] The time corresponding to the number of consecutive disconnected data points of the target is determined as the target time.
[0186] Optionally, the common point data determination module 610 is specifically used to calculate the Euclidean distance between each point data in one first-class cluster and each point data in the other first-class cluster for two first-class clusters; and determine the two point data whose corresponding Euclidean distance is less than the distance threshold as the common point data of the two first-class clusters.
[0187] Optionally, the third cluster determination module 612 is specifically used to determine whether the number of common point data meets the preset conditions through the following steps:
[0188] Statistical analysis was performed on the first type of cluster to determine the number of long-tail data points in the first type of cluster.
[0189] The ratio of the number of common data points to the number of long-tail data points is determined as the cluster common purity.
[0190] If the common purity of the cluster is greater than the purity threshold, then the number of common point data is determined to meet the preset conditions.
[0191] If the purity of a cluster is less than or equal to the purity threshold, then the number of shared data points does not meet the preset conditions.
[0192] Optionally, the cluster discrimination determination module 616 is specifically used to take the set of the second cluster and / or the third cluster and the fourth cluster as candidate clusters, and determine the local density ratio of each candidate cluster according to the local density algorithm; calculate the silhouette coefficient and CH index of each candidate cluster; and determine the cluster discrimination of each candidate cluster according to the local density ratio, silhouette coefficient and CH index.
[0193] The specific details of each module or unit in the above-mentioned device have been described in detail in the corresponding methods, so they will not be repeated here.
[0194] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of this application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0195] In an exemplary embodiment of this application, an electronic device is also provided, including: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to execute the operating condition category analysis method described above in this exemplary embodiment.
[0196] Figure 7 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application. It should be noted that... Figure 7 The electronic device 700 shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0197] like Figure 7 As shown, the electronic device 700 includes a central processing unit (CPU) 701, which can perform various appropriate actions and processes based on a program stored in a read-only memory (ROM) 702 or a program loaded from a storage section 708 into a random access memory (RAM) 703. The RAM 703 also stores various programs and data required for system operation. The CPU 701, ROM 702, and RAM 703 are interconnected via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0198] The following components are connected to I / O interface 705: an input section 706 including a keyboard, mouse, etc.; an output section 707 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a local area network (LAN) card, modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to I / O interface 705 as needed. A removable medium 711, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on drive 710 as needed so that computer programs read from it can be installed into storage section 708 as needed.
[0199] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 709, and / or installed from removable medium 711. When the computer program is executed by central processing unit 701, it performs the various functions defined in the apparatus of this application.
[0200] In this embodiment of the application, a computer-readable storage medium is also provided, on which a computer program is stored, which, when executed by a processor, implements the above-described working condition category analysis method.
[0201] It should be noted that the computer-readable storage medium shown in this application can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this application, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable storage medium can be transmitted using any suitable medium, including but not limited to: wireless, wire, optical fiber, radio frequency, etc., or any suitable combination thereof.
[0202] In this embodiment of the application, a computer program product is also provided, which, when run on a computer, causes the computer to execute the above-described working condition category analysis method.
[0203] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0204] The above description is merely a specific embodiment of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for analyzing operating conditions, characterized in that, The method includes: Industrial time-series data during industrial production is acquired. Based on the distribution information of the industrial time-series data, cluster analysis is performed on the industrial time-series data to obtain a first cluster and the operating condition category to which the first cluster belongs. The industrial time-series data is continuously collected point data, which includes pressure, temperature, and production load. Target point data that is temporally correlated with the point data in the first cluster is obtained from the point data outside each first cluster, and the target point data is added to the first cluster to obtain the second cluster; From the two first-category clusters, filter out the common point data that belong to both first-category clusters; If the number of shared location data meets the preset conditions, the shared location data is regarded as the third cluster, and the location data after removing the shared location data from the two first clusters is regarded as the fourth cluster. The set of the third and fourth clusters and / or the second cluster are used as candidate clusters, and the cluster discrimination degree of each candidate cluster is calculated according to a preset formula. The cluster discrimination degree is used to characterize the compactness of the candidate clusters and the degree of dispersion from other candidate clusters. Candidate clusters with a cluster discrimination score greater than the discrimination score threshold are taken as the final clusters, and the working condition category to which the final clusters belong is obtained. The steps of deleting the point data belonging to the final cluster in the industrial time series data to update the industrial time series data, returning the distribution information based on the industrial time series data, and performing cluster analysis on the industrial time series data, continue until the industrial time series data no longer contains point data. The method further includes: after obtaining the first cluster and the working condition category to which the first cluster belongs, selecting a target cluster with a density greater than a density threshold from the first cluster; The target point data obtained from point data outside each of the first clusters, which has a temporal correlation with the point data in the first cluster, includes: Target point data that is temporally correlated with the point data in the target clusters is obtained from point data outside each target cluster; From the two first-category clusters, filter the common location data that belong to both first-category clusters, including: From the two target clusters, filter out the common point data that belong to both target clusters; The step of removing the common point data from the two first clusters and then defining them as the fourth cluster includes: The point data after removing the common point data from the two target clusters are respectively used as the fourth cluster.
2. The method according to claim 1, characterized in that, After acquiring industrial time-series data from the industrial production process, the method further includes: The industrial time-series data is subjected to dimensionality reduction processing to obtain dimensionality-reduced data; The clustering analysis of industrial time-series data based on distribution information includes: Based on the distribution information of the industrial time-series data, cluster analysis is performed on the dimensionality-reduced data.
3. The method according to claim 1 or 2, characterized in that, The target point data obtained from point data outside each of the first clusters, which has a temporal correlation with the point data in the first cluster, includes: Based on the data collection time of each point in each of the first clusters, determine the number of consecutive point data and the number of subsequent consecutive disconnected point data corresponding to the time when each point data is located. The target time is determined based on the number of consecutive data points corresponding to the time of each data point and the number of subsequent consecutive disconnected data points. For each target time, the target search points adjacent to the target time are used as target point data according to the time dimension.
4. The method according to claim 3, characterized in that, The step of determining the target time based on the number of consecutive data points corresponding to the time of each data point and the number of subsequent consecutive disconnected data points includes: In the set of consecutive point data corresponding to the time of each point data, the number of consecutive point data that is less than the first preset quantile is selected as the candidate consecutive point data. The number of subsequent consecutive disconnected point data corresponding to the time when the number of candidate consecutive point data is located is taken as the first candidate number of subsequent consecutive disconnected point data. Select a second candidate number of consecutive disconnection points that is less than the second preset quantile from the set of data of the first candidate subsequent consecutive disconnection points. The time corresponding to the number of subsequent consecutive disconnection points in the second candidate is determined as the target time; or... In the set of subsequent consecutive disconnection point data corresponding to the time of each point data, the number of subsequent consecutive disconnection point data that is less than the third preset quantile is selected as the target number of subsequent consecutive disconnection point data. The time corresponding to the number of consecutive disconnected data points of the target is determined as the target time.
5. The method according to claim 1 or 2, characterized in that, The step of filtering common location data belonging to both of the first clusters includes: For two first-class clusters, calculate the Euclidean distance between each data point in one first-class cluster and each data point in the other first-class cluster; Two points whose Euclidean distance is less than the distance threshold are identified as shared points of the two first clusters.
6. The method according to claim 1 or 2, characterized in that, The methods for determining whether the number of shared location data meets the preset conditions include: Perform statistical analysis on the first cluster to determine the number of long-tail data points in the first cluster; The ratio of the number of common point data to the number of long-tail point data is determined as the cluster common purity. If the common purity of the cluster is greater than the purity threshold, then the number of common point data is determined to meet the preset condition. If the common purity of the cluster is less than or equal to the purity threshold, then it is determined that the number of common point data does not meet the preset condition.
7. The method according to claim 1 or 2, characterized in that, The calculation of the cluster discrimination of each candidate cluster according to the preset formula includes: Based on the local density algorithm, the local density ratio of each candidate cluster is determined; Calculate the silhouette coefficient and CH index for each of the candidate clusters; The cluster discrimination of each candidate cluster is determined based on the local density ratio, the contour coefficient, and the CH index.
8. A working condition category analysis device, characterized in that, The device includes: An industrial time-series data acquisition module is used to acquire industrial time-series data during the industrial production process; the industrial time-series data is continuously collected point data, and the point data includes pressure, temperature, and production load; The clustering analysis module is used to perform clustering analysis on industrial time series data based on the distribution information of industrial time series data, and obtain the first cluster and the operating condition category to which the first cluster belongs; The target point data acquisition module is used to acquire target point data that has a temporal correlation with the point data in the first cluster from point data outside each of the first clusters; The first cluster update module is used to add the target point data to the first cluster to obtain the second cluster; The common location data determination module is used to filter common location data that simultaneously belong to both of the two first type clusters from the two first type clusters; The third cluster determination module is used to identify the common point data as a third cluster when the number of common point data meets a preset condition. The fourth cluster determination module is used to determine the point data after removing the common point data from the two first clusters as the fourth cluster; The cluster discrimination determination module is used to take the set of the second cluster and / or the third cluster and the fourth cluster as candidate clusters, and calculate the cluster discrimination of each candidate cluster according to a preset formula. The cluster discrimination is used to characterize the density of the candidate clusters and the degree of dispersion from other candidate clusters. The final cluster determination module is used to select candidate clusters with a cluster discrimination degree greater than the discrimination degree threshold as the final clusters and obtain the working condition category to which the final clusters belong. The industrial time series data update module is used to delete the point data belonging to the final cluster in the industrial time series data in order to update the industrial time series data. The return module is used to return the clustering analysis module to perform the steps of clustering analysis on the industrial time series data based on the distribution information of the industrial time series data, until the industrial time series data does not contain any point data. The working condition category analysis device further includes: The target cluster selection module is used to select target clusters with a density greater than a density threshold from the first clusters; The target point data acquisition module is specifically used to acquire target point data that has a temporal correlation with the point data in the target cluster from the point data outside each target cluster; The common location data determination module is specifically used to filter common location data that belong to both target clusters from the two target clusters. The fourth cluster determination module is specifically used to take the point data after removing the common point data from the two target clusters as the fourth cluster.
9. An electronic device, characterized in that, include: A processor for executing a computer program stored in a memory, wherein the computer program, when executed by the processor, implements the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Information acquisition method and device, equipment and storage medium
CN113971426A
Process industrial energy consumption state evaluation method based on clustering
CN114565209A