An industrial data edge perception fusion management system based on cloud edge collaboration
Patent Information
- Application Number
- CN202610993363.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-04
- Publication Date
- 2026-09-29
AI Technical Summary
但是由于不同种类工业数据之间数据的取值范围、局部密度、变化趋势等数据类间差异较为显著,在利用聚类算法对不同种类数据进行融合处理时,无法反映不同种类数据内部密度结构各异且与量级冲突的工业数据
1、本申请针对工业数据中不同类别在取值范围、局部密度、变化趋势等方面差异显著的问题,通过各类工业数据聚类过程中各数据点相对聚类中心的位置分布情况分析其置信度,并引入局部密度计算,确定加权局部密度,能够有效刻画各类数据内部的密度结构差异,克服了传统聚类算法在密度不均与量级冲突场景下的局限性,显著提升了对工业数据的分类精度;
Smart Images

Figure CN122845448A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of cloud-edge collaboration technology, specifically to an industrial data edge perception fusion management system based on cloud-edge collaboration. Background Technology
[0002] In industrial sectors such as manufacturing, power, energy, and transportation, edge computing technology, combined with multi-protocol data access, data fusion processing, and cloud-edge collaboration mechanisms, enables real-time collection, on-site processing, and unified management of industrial field data, effectively improving industrial data processing efficiency and system responsiveness.
[0003] In the entire processing flow of industrial data based on cloud-edge collaboration, due to the diverse data types of industrial big data, it is necessary to accurately classify the massive amounts of industrial big data to facilitate the system's processing or management of each type of data. However, because the differences in value range, local density, and trend of different types of industrial data are quite significant, clustering algorithms cannot reflect the industrial data with varying internal density structures and conflicting magnitudes when used to fuse different types of data. Summary of the Invention
[0004] To address the aforementioned technical problems, the purpose of this application is to provide an industrial data edge-aware fusion management system based on cloud-edge collaboration, and the specific technical solution adopted is as follows: This application proposes an industrial data edge-aware fusion management system based on cloud-edge collaboration. The system includes: a multi-protocol data acquisition module, an edge computing module, a data fusion and modeling module, a management module, a data visualization module, and a security protection module. The system includes: a multi-protocol data acquisition module for unified data access from different devices; an edge computing module for clustering various types of data, analyzing the clustering results and data distribution during the iteration process, calculating composite distances by combining the data's inherent values, optimizing classification algorithms, and performing data classification; a data fusion and modeling module for performing fusion operations on the classified data and building device operation models, including health assessment models for device operation; a management module for centralized management of edge nodes and data resources; a data visualization module for displaying device status in various formats; and a security protection module for providing data transmission encryption, access control, and multi-node disaster recovery mechanisms to ensure system operation security, data privacy, and high availability.
[0005] Preferably, the edge computing module includes: a clustering analysis module, a near-far neighbor classification module, a local density evaluation module, and a data classification module; The clustering analysis module is used to calculate the confidence feature value of each data point based on the difference in distance between each data point and each data point in its cluster during each iteration when clustering various types of data collected by the multi-protocol data acquisition module. The near-far neighbor classification module is used to classify high-confidence near neighbors and high-confidence far neighbors based on the confidence feature value, and set corresponding label values; The local density evaluation module is used to calculate the weighted local density of each data point based on the classification results of the nearest and far neighbors of each data point in the neighborhood of each data point. The data classification module is used to calculate the composite distance between any two time points based on the data point values, the label values, and the weighted local density; and to use the composite distance as the distance metric in the clustering algorithm to classify each time point.
[0006] Preferably, the process of calculating the confidence feature value of each data point based on the difference in distance between each data point and each data point in its cluster during each iteration is as follows: Obtain the average distance from all points within the cluster to the cluster center in each iteration. Based on the difference between the distance from each data point to its cluster center and the average distance, as well as the difference between the distance from each data point to other cluster centers, determine the confidence feature value of each data point.
[0007] Preferably, the formula for calculating the confidence feature value is: In the formula For data points Confidence feature value; This represents the number of clustering iterations. This represents the actual distance from data point k to its cluster center in the i-th iteration. It is the average distance from all points in the cluster containing data point k to the cluster center in the i-th iteration, used to characterize the overall density range in this iteration; This is a preset, extremely small positive number used to prevent overflow caused by a denominator of 0; M is the number of clusters automatically determined by AP clustering. Let be the distance from data point k to the j-th cluster center in the i-th iteration.
[0008] Preferably, the classification method for high-confidence nearest neighbors and high-confidence distant neighbors is as follows: Data points with confidence feature values less than or equal to a preset confidence threshold are designated as high-confidence nearest neighbors, while data points with confidence feature values greater than the preset confidence threshold are designated as high-confidence distant neighbors.
[0009] Preferably, the specific setting method of the label value is as follows: set the label value of the high confidence nearest neighbor to 1, and set the label value of the high confidence distant neighbor to 0.
[0010] Preferably, the local density assessment module specifically includes: Within the neighborhood of each data point, the number of high-confidence nearest neighbors and the number of high-confidence distant neighbors are counted separately, and then weighted and summed to obtain the weighted local density of each data point.
[0011] Preferably, the weighted summation uses a weight value of 1.
[0012] Preferably, the weighted summation process is as follows: The sum of the absolute values of the confidence feature values of the high-confidence nearest neighbors of each data point is recorded as the overall nearest neighbor confidence score, and the sum of the absolute values of the confidence feature values of the high-confidence distant neighbors of each data point is recorded as the overall distant neighbor confidence score. The sum of the overall nearest neighbor confidence score and the overall distant neighbor confidence score is used as the denominator, the ratio of the overall nearest neighbor confidence score to the denominator is used as the weight of the number of high-confidence nearest neighbors, and the ratio of the overall distant neighbor confidence score to the denominator is used as the weight of the number of high-confidence distant neighbors. A weighted sum of the number of high-confidence nearest neighbors and the number of high-confidence distant neighbors is then performed.
[0013] Preferably, the process of calculating the composite distance between any two time points based on the data point values, the label values, and the weighted local density is as follows: The data of all types at each time point are arranged into a sequence, denoted as the data sequence; the label values of the data of all types at each time point are used to generate the multi-label vector of the data sequence at each time point; the weighted local density of the data of all types at each time point is used to generate the density vector of the data sequence at each time point. The composite distance between any two time points is determined based on the distance between the data sequences at any two time points, the distance between the multi-label vectors, and the distance between the density vectors.
[0014] Compared with the prior art, the beneficial effects of this application are: 1. This application addresses the problem that different categories of industrial data have significant differences in value range, local density, and trend of change. It analyzes the confidence level of each data point relative to the cluster center during the clustering process of various types of industrial data, and introduces local density calculation to determine the weighted local density. This can effectively characterize the density structure differences within various types of data, overcome the limitations of traditional clustering algorithms in scenarios with uneven density and conflicting magnitudes, and significantly improve the classification accuracy of industrial data. 2. By generating binary multi-label vectors and density vectors for the data sequence at each sampling time, and constructing a composite distance based on the distance between the multi-label vectors and density vectors at two times and the distance between the original data sequences, local density information is incorporated while preserving the features of the original data, thus achieving accurate temporal clustering of the device's operating status and facilitating the identification of various status categories such as healthy, abnormal, and overloaded states. 3. Based on the data classification results, the system can group, clean, standardize, and extract features from the data according to equipment status categories, construct equipment operation models (such as health assessment models), and support online updates of the models. This fusion strategy effectively improves data utilization efficiency and model generalization ability. Attached Figure Description
[0015] Figure 1 This is a block diagram of an industrial data edge-aware fusion management system based on cloud-edge collaboration, according to an embodiment of this application. Figure 2 This is a flowchart of the edge computing module in an embodiment of this application. Detailed Implementation
[0016] The following will describe in detail, with reference to the accompanying drawings in the embodiments of the present invention, a specific solution of an industrial data edge perception fusion management system based on cloud-edge collaboration provided in this application.
[0017] Please see Figure 1-2 The diagram illustrates a block diagram of an industrial data edge-aware fusion management system based on cloud-edge collaboration, according to an embodiment of this application. The system includes: a multi-protocol data acquisition module, an edge computing module, a data fusion and modeling module, a management module, a data visualization module, and a security protection module. The multi-protocol data acquisition module is used to enable unified data access for industrial equipment of different brands and types (such as sensors, PLCs, and industrial robots).
[0018] The edge computing module is used to classify (e.g., by device type, parameter type, urgency, etc.), filter, preprocess, and perform rule calculations on the collected data, thereby reducing cloud load and improving the system's real-time response capability.
[0019] The data fusion and modeling module, based on the classification results of the edge computing module, cleans, standardizes, and correlates multi-source heterogeneous data to build equipment operation models (such as health assessment models), providing a unified and high-quality data foundation for intelligent decision-making.
[0020] The management module enables centralized management of edge nodes and data resources, including device configuration, status monitoring, alarm policies, user permissions, log auditing, etc., and supports system-level operation and maintenance and control.
[0021] The data visualization module displays equipment operating status, historical data, health scores, and anomaly alarms in the form of charts, dashboards, and trend curves, supporting real-time monitoring and analysis decision-making in industrial scenarios.
[0022] The security protection module provides data transmission encryption, access control, and multi-node disaster recovery mechanisms to ensure system operation security, data privacy, and high availability.
[0023] Specifically, the multi-protocol data acquisition module includes: For different brands and types of industrial equipment within the park, corresponding protocol adapters are deployed. For example: Sensor devices: connect via protocol adapters such as Modbus RTU / TCP, Zigbee, and LoRa; PLC-type devices: connected via protocol adapters such as OPC UA, S7, and EtherNet / IP; Industrial robot equipment: connected via dedicated industrial Ethernet protocol or vendor SDK adapter.
[0024] Each protocol adapter is responsible for parsing raw data from different communication protocols into intermediate data in a unified format, including the device identifier, timestamp, data type, value, and unit of measurement for each data item. This enables unified data access for heterogeneous devices.
[0025] Furthermore, all categories of industrial data are resampled at the same sampling frequency to ensure that the data types contained at each sampling moment are the same during subsequent processing. The resampled industrial data are then standardized. Specifically, this application uses a maximum-minimum normalization method to standardize the industrial data for each category. In particular, if the maximum and minimum values of a certain category of industrial data are the same, the normalized value of that category of industrial data is set to 1. Further, a data sequence consisting of the standardized values of all categories of data after resampling is generated at each sampling moment. In this embodiment, the resampled data sampling frequency is set to 1Hz.
[0026] The edge computing module includes: clustering analysis module, near and far neighbor classification module, local density evaluation module, and data classification module.
[0027] The clustering analysis module is used to calculate the confidence feature value of each data point based on the difference in distance between each data point and each data point in its cluster during each iteration when clustering various types of data collected by the multi-protocol data acquisition module.
[0028] Specifically, firstly, clustering algorithms such as AP clustering or DBSCAN clustering, which do not require setting the number of clusters, are used to classify each type of industrial data. Preferably, in this embodiment, AP clustering is chosen to classify each type of industrial data. The damping coefficient during clustering is typically set between 0.5 and 1; in this embodiment, the damping coefficient is set to 0.7, and the preference parameter is set to the median similarity. AP clustering is a well-known technique, and its specific process will not be elaborated upon. Implementers can set the damping coefficient and preference parameter according to their actual situation; this application does not impose specific restrictions.
[0029] For any data point k in this type of industrial data, if the attraction of data point k to other data points remains at a high level during the AP clustering iteration process, and the other data points with a high degree of belonging to data point k remain in a small fluctuation state during the iteration process, it indicates that data point k belongs to a high-confidence cluster center. Similarly, during the iteration process, if the data point k satisfies the following conditions: it becomes the center of a cluster less often, it stays in the same cluster more often, and its metric distance to the cluster center within the same cluster is always kept at a low level and is always less than its metric distance to the centers of other clusters, it indicates that the deviation when using the data point k to represent the center point is smaller, and the data point k belongs to a high-quality nearest neighbor.
[0030] Based on this, in a preferred embodiment of this application, the distribution characteristics of each data point in the classification results of each type of industrial data are analyzed, and the confidence feature value of each data point is calculated to quantify the state of each data point under the current clustering distribution. The expression is as follows: In the formula For data points Confidence feature value; This represents the number of iterations for AP clustering. This represents the actual distance from data point k to its cluster center in the i-th iteration. It is the average distance from all points in the cluster containing data point k to the cluster center in the i-th iteration, used to characterize the overall density range in this iteration; This is a preset, extremely small positive number used to prevent overflow caused by a denominator of 0; M is the number of clusters automatically determined by AP clustering. Let be the distance from data point k to the j-th cluster center in the i-th iteration.
[0031] in, The middle part compares the actual distance with the average distance by performing a difference operation. If the numerical comparison result is large and positive, it means that the data point k is much farther from the cluster center than the average level of remoteness and can be classified as an outlier. If the numerical comparison result is small or even negative, it means that the data point k is closer to the cluster center and can be classified as a core point. The denominator is used to smooth the difference. This step extracts the absolute value of the distance difference. This ensures that the summation proceeds in the positive direction. If data point k is closer to its own cluster and farther from other clusters, the denominator will have a larger value, making... It tends to converge and become smaller.
[0032] The nearest and far neighbor classification module is used to classify high-confidence nearest neighbors and high-confidence far neighbors based on the confidence feature value, and set corresponding label values.
[0033] Specifically, a confidence feature value is calculated for each data point. Within each cluster, the confidence feature values of all data points are sorted in descending order; the lower the ranking, the more reasonable it is to be considered a high-confidence nearest neighbor. The average of all confidence feature values is selected as the confidence threshold. Data points with confidence feature values less than or equal to the confidence threshold are considered high-confidence nearest neighbors, while those greater than the threshold are considered high-confidence distant neighbors. Furthermore, a label value is set for each data point; specifically, the label value for high-confidence nearest neighbors is set to 1, and the label value for high-confidence distant neighbors is set to 0.
[0034] The local density evaluation module is used to calculate the weighted local density of each data point based on the classification results of the nearest and far neighbors of each data point in its neighborhood.
[0035] Specifically, because different types of industrial data differ significantly in local density and magnitude, relying solely on... The value itself is insufficient to accurately characterize the local density distribution around the data point. In a preferred embodiment of this application, the number of neighboring data points for each data point is first set according to the size of each cluster. Specifically, the number of neighboring data points K for each data point is 1.5% of the number of data points in the cluster. The neighboring data points of data point k are the K data points closest to data point k. If K is not an integer, then K is rounded up to obtain the final value of K.
[0036] Then, the number of high-confidence nearest neighbors and high-confidence distant neighbors of data point k are counted separately; the sum of the absolute values of the confidence feature values of the high-confidence nearest neighbors of each data point is recorded as the overall nearest neighbor confidence. The sum of the absolute values of the confidence feature values of the high-confidence distant neighbors of each data point is denoted as the overall distant neighbor confidence score. Furthermore, based on the above two values, the nearest neighbor weight and the far neighbor weight of data point k are determined, and the expression is: , . , These figures respectively reflect the proportions of the sum of confidence scores of nearest neighbors and the sum of confidence scores of far neighbors in the total local confidence score. This is a preset minimum positive additional constant, used to prevent the denominator from being zero in extremely rare cases.
[0037] Next, the weighted local density of data point k is calculated, expressed as: In the formula, The weighted local density of data point k measures the degree to which the point is influenced by high-confidence nearest / far neighbor points in its neighborhood, and is an important feature for subsequent multi-label classification. This represents the number of high-confidence nearest neighbors of data point k in its neighborhood. This represents the number of high-confidence distant neighbors of data point k in its neighborhood. The nearest neighbor weights of data point k, Let K be the weights of the far neighbors of data point k, and satisfy the following conditions: .
[0038] This embodiment, through the aforementioned confidence-weighted processing, can extract more refined density characterization features for situations where there are large differences in local density among different categories in industrial data.
[0039] The data classification module is used to calculate the composite distance between any two time points based on the data point values, the label values, and the weighted local density; and to use the composite distance as the distance metric in the clustering algorithm to classify each time point.
[0040] Specifically, since the above results are currently only at the level of individual data points, they have not yet classified the data sequences at each sampling time from a holistic time series perspective. In industrial settings, data is collected continuously in chronological order, and the data sequences at different times not only exhibit similarity in category labels but are also influenced by the local density distribution of similar data points. To achieve accurate identification and assessment of equipment status, in a preferred embodiment of this application, the data sequence at each time point is treated as a whole, and a binary multi-label vector is generated for each time point's data sequence using the high-confidence nearest / far-neighbor label values of each data point. By utilizing the weighted local density of each data point, a density vector is generated for the data sequence at each time step. The expression for the composite distance between any two time points is: In the formula, for Time and Composite distance between data sequences at different times; for Time and The normalized value of the Hamming distance between the multi-label vectors at time step; for Time and The normalized value of the Euclidean distance between the density vectors at time step; for Time and The normalized value of the Euclidean distance between data sequences at different times; , , These are the weighting coefficients for the three distance components mentioned above, used to balance the contributions of labels, local density, and original data in the composite distance, satisfying... The values are set to 0.3, 0.3, and 0.4 respectively. The above normalization uses maximum normalization, that is, the Hamming distance between the multi-label vectors at all times is used as the input of the maximum normalization method to obtain the normalized value of the Hamming distance. Similarly, the normalized values of the Euclidean distance between density vectors and the Euclidean distance between data sequences can be obtained. The specific process will not be described in detail.
[0041] Next, the device status categories are set. In this embodiment, they are divided into three categories: healthy device status, abnormal device status, and overloaded device status. In other embodiments of this application, the implementer can set the device status categories according to the actual situation.
[0042] Based on the distance metric and the number of categories, a time-series clustering algorithm is used to classify the data sequence at all times. The composite distance between any two times is used as the distance metric in the time-series clustering algorithm to cluster all times. Using the above distance matrix, combined with the preset number of state categories (healthy, abnormal, overload), a time-series clustering algorithm (such as K-medoids, hierarchical clustering) is run to assign each time t to a certain state category, thus completing the data classification.
[0043] Specifically, for the three clusters obtained by clustering, the average value of each industrial data at all times within each cluster is calculated to obtain the centroid vector of each cluster; the centroid vector corresponding to the equipment data during historical normal operation is obtained and denoted as the baseline vector. Specifically, in this embodiment, the average value of various industrial data under normal equipment operation within the past week is used to construct the baseline vector; the difference between the centroid vector of each cluster and the baseline vector is calculated, and the cluster corresponding to the centroid vector with the largest difference is classified as overloaded, the next largest as abnormal, and the cluster corresponding to the centroid vector with the smallest difference is classified as healthy. In a preferred embodiment of this application, the difference between the centroid vector and the baseline vector can be quantitatively evaluated by calculating the Euclidean distance between them. The larger the Euclidean distance, the greater the difference.
[0044] The data fusion and modeling module performs the following operations: (1) Group and aggregate the data according to the status category, and put the data of different times and different parameter types (such as voltage, current, temperature and power) under the same category into the same dataset.
[0045] (2) Receive the classification data results and clean them, including removing outliers, filling missing values with linear interpolation, and eliminating inconsistencies caused by local environmental differences in industrial data.
[0046] (3) For each state category, extract statistical features and time-series features. The statistical features include mean, variance, rate of change, and peak value. The time-series features include trend slope and autocorrelation coefficient. Construct the equipment operation feature vector for that category through association rule mining or weighted fusion methods. For example, fuse the "healthy" category to form a benchmark health model and fuse the "abnormal" category to form a fault feature library.
[0047] (4) On this basis, machine learning methods (such as Bayesian fusion, cluster center fusion or neural network fusion) are used to synthesize multi-source and multi-time data within the same category into a unified state representation model, and the model is updated online as new data is added.
[0048] The management module utilizes the aforementioned device operation model to achieve centralized management of edge nodes and data resources, including real-time status monitoring, alarm policy triggering, device configuration optimization, and log auditing.
[0049] The data visualization module displays the integrated health score and anomaly warnings in the form of charts, dashboards, and trend curves. This forms a complete closed loop from "data classification → fusion by category → status modeling → intelligent management," achieving the overall goal of the industrial data edge-aware fusion management system.
[0050] In summary, this application addresses the issue of significant differences in value range, local density, and trend among different categories of industrial data. By analyzing the confidence level of each data point relative to the cluster center during the clustering process of various types of industrial data, and introducing local density calculation to determine the weighted local density, it can effectively characterize the density structure differences within various types of data. This overcomes the limitations of traditional clustering algorithms in scenarios with uneven density and conflicting magnitudes, and significantly improves the classification accuracy of industrial data. By generating binary multi-label vectors and density vectors for the data sequence at each sampling time, and constructing a composite distance based on the distance between the multi-label vectors and density vectors at two times and the distance between the original data sequences, local density information is incorporated while preserving the features of the original data, thus achieving accurate temporal clustering of the device's operating status and facilitating the identification of various status categories such as health, abnormality, and overload. Based on the data classification results, the system can group, clean, standardize, and extract features from the data according to equipment status categories, construct equipment operation models (such as health assessment models), and support online updates of the models. This fusion strategy effectively improves data utilization efficiency and model generalization ability.
Claims
1. An industrial data edge-aware fusion management system based on cloud-edge collaboration, characterized in that, The system includes: a multi-protocol data acquisition module, an edge computing module, a data fusion and modeling module, a management module, a data visualization module, and a security protection module; The system includes: a multi-protocol data acquisition module for unified data access from different devices; an edge computing module for clustering various types of data, analyzing the clustering results and data distribution during the iteration process, calculating composite distances by combining the data's inherent values, optimizing classification algorithms, and performing data classification; a data fusion and modeling module for performing fusion operations on the classified data and building device operation models, including health assessment models for device operation; a management module for centralized management of edge nodes and data resources; a data visualization module for displaying device status in various formats; and a security protection module for providing data transmission encryption, access control, and multi-node disaster recovery mechanisms to ensure system operation security, data privacy, and high availability.
2. The industrial data edge-aware fusion management system based on cloud-edge collaboration according to claim 1, characterized in that, The edge computing module includes: a clustering analysis module, a near-far neighbor classification module, a local density evaluation module, and a data classification module; The clustering analysis module is used to calculate the confidence feature value of each data point based on the difference in distance between each data point and each data point in its cluster during each iteration when clustering various types of data collected by the multi-protocol data acquisition module. The near-far neighbor classification module is used to classify high-confidence near neighbors and high-confidence far neighbors based on the confidence feature value, and set corresponding label values; The local density evaluation module is used to calculate the weighted local density of each data point based on the classification results of the nearest and far neighbors of each data point in the neighborhood of each data point. The data classification module is used to calculate the composite distance between any two time points based on the data point values, the label values, and the weighted local density; and to use the composite distance as the distance metric in the clustering algorithm to classify each time point.
3. The industrial data edge-aware fusion management system based on cloud-edge collaboration according to claim 2, characterized in that, The process of calculating the confidence feature value of each data point based on the difference in distance between each data point and each data point in its cluster during each iteration is as follows: Obtain the average distance from all points within the cluster to the cluster center in each iteration. Based on the difference between the distance from each data point to its cluster center and the average distance, as well as the difference between the distance from each data point to other cluster centers, determine the confidence feature value of each data point.
4. The industrial data edge-aware fusion management system based on cloud-edge collaboration according to claim 3, characterized in that, The formula for calculating the confidence level feature value is as follows: In the formula For data points Confidence feature value; This represents the number of clustering iterations. This represents the actual distance from data point k to its cluster center in the i-th iteration. It is the average distance from all points in the cluster containing data point k to the cluster center in the i-th iteration, used to characterize the overall density range in this iteration; This is a preset, extremely small positive number used to prevent overflow caused by a denominator of 0; M is the number of clusters automatically determined by AP clustering. Let be the distance from data point k to the j-th cluster center in the i-th iteration.
5. The industrial data edge-aware fusion management system based on cloud-edge collaboration according to claim 2, characterized in that, The classification method for high-confidence nearest neighbors and high-confidence distant neighbors is as follows: Data points with confidence feature values less than or equal to a preset confidence threshold are designated as high-confidence nearest neighbors, while data points with confidence feature values greater than the preset confidence threshold are designated as high-confidence distant neighbors.
6. The industrial data edge-aware fusion management system based on cloud-edge collaboration according to claim 2, characterized in that, The specific method for setting the label values is as follows: set the label value of high-confidence nearest neighbors to 1, and set the label value of high-confidence distant neighbors to 0.
7. The industrial data edge-aware fusion management system based on cloud-edge collaboration according to claim 2, characterized in that, The local density assessment module specifically includes: Within the neighborhood of each data point, the number of high-confidence nearest neighbors and the number of high-confidence distant neighbors are counted separately, and then weighted and summed to obtain the weighted local density of each data point.
8. The industrial data edge-aware fusion management system based on cloud-edge collaboration according to claim 7, characterized in that, The weighted summation uses a weight of 1.
9. The industrial data edge-aware fusion management system based on cloud-edge collaboration according to claim 8, characterized in that, The weighted summation process is as follows: The sum of the absolute values of the confidence feature values of the high-confidence nearest neighbors of each data point is recorded as the overall nearest neighbor confidence, and the sum of the absolute values of the confidence feature values of the high-confidence distant neighbors of each data point is recorded as the overall distant neighbor confidence. The sum of the overall nearest neighbor confidence score and the overall far neighbor confidence score is used as the denominator. The ratio of the overall nearest neighbor confidence score to the denominator is used as the weight of the number of high-confidence nearest neighbors, and the ratio of the overall far neighbor confidence score to the denominator is used as the weight of the number of high-confidence far neighbors. The weighted sum of the number of high-confidence nearest neighbors and the number of high-confidence far neighbors is then calculated.
10. The industrial data edge-aware fusion management system based on cloud-edge collaboration according to claim 2, characterized in that, The process of calculating the composite distance between any two time points based on the data point values, the label values, and the weighted local density is as follows: The data of all types at each time point are arranged into a sequence, denoted as the data sequence; the label values of the data of all types at each time point are used to generate the multi-label vector of the data sequence at each time point; the weighted local density of the data of all types at each time point is used to generate the density vector of the data sequence at each time point. The composite distance between any two time points is determined based on the distance between the data sequences at any two time points, the distance between the multi-label vectors, and the distance between the density vectors.