Distributed data processing method, system and equipment and storage medium

By performing sharded data compression and data covariate analysis on industrial data in distributed data processing, the threshold value is dynamically determined, which solves the problem that threshold setting in distributed data processing is difficult to take into account global accuracy and local sensitivity, and improves the accuracy of abnormal culling.

CN120215843AActive Publication Date: 2025-06-27GUIZHOU BUSINESS SCHOOL
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510702743.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-06-27
Estimated Expiration
2045-05-29

AI Technical Summary

Technical Problem

In distributed data processing, the heterogeneity and non-stationary properties of industrial data make it difficult to take into account the global accuracy and local sensitivity of static thresholds, resulting in false positives or missed detection, and frequent manual intervention is required.

Method used

By sharding data compression of industrial data packets in distributed systems, data covariates and distribution density chaos are extracted, data covariance threshold values ​​are dynamically determined, potential abnormal data are divided, and de-elimination data packets are constructed.

Benefits of technology

It realizes the setting of a dynamic adaptive threshold that takes into account global accuracy and local sensitivity in distributed data processing, improves the accuracy of abnormal culling, reduces false detection and missed detection, and avoids the roughness of traditional global threshold settings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120215843A_ABST
    Figure CN120215843A_ABST
Patent Text Reader

Abstract

The invention provides a distributed data processing method, system and device and a storage medium, and relates to the technical field of distributed data processing, and the method comprises the steps: obtaining a to-be-processed industrial data packet in a distributed system, carrying out the fragmentation data compression of the to-be-processed industrial data packet, and obtaining a plurality of data compression fragments; extracting an industrial data covariant corresponding to each data compression fragment, thereby determining a distribution density confusion degree of each data compression fragment, and further determining a data covariant threshold value of each data compression fragment; determining potential abnormal data of each data compression fragment according to the corresponding data covariant threshold value, and determining a de-isomerism data packet in industrial data distributed processing through the potential abnormal data of each data compression fragment; and storing the de-dissimilarity data packet to complete distributed processing of the industrial data. According to the method and the device, the dynamic self-adaptive threshold value considering the global accuracy and the local sensitivity can be set, so that the accuracy of exception elimination in distributed data processing is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of distributed data processing. More specifically, this application relates to a distributed data processing method, system, device, and storage medium. Background Art

[0002] Distributed data processing refers to a method of dividing large-scale data into several subsets, which are processed in parallel by multiple computing nodes and then the results are aggregated. With the advent of the big data era, the amount of data has shown an explosive growth. The traditional single-machine processing mode has been difficult to meet the requirements of real-time performance, high availability, and scalability. Therefore, distributed data processing technology has become an important foundation to support various applications. From basic data storage (such as the distributed file system HDFS) to computing frameworks (such as MapReduce, Spark, Flink), and then to data transmission systems (such as Kafka, Pulsar), the distributed data processing system has been continuously enriched and improved.

[0003] In the prior art, industrial data in distributed data processing often has high heterogeneity. The data distributions of different devices, production lines, or time windows are significantly different. Fixed thresholds are difficult to balance global accuracy and local sensitivity, resulting in false alarms or missed detections. Moreover, data streams often exhibit non-stationary characteristics (such as distribution drift caused by equipment aging and working condition switching). Static thresholds cannot be adaptively adjusted and require frequent manual intervention. In a large-scale distributed environment, the statistical characteristics between data shards may have significant differences. Uniform thresholds are likely to cause resource waste or key anomalies to be missed. Many methods rely on static thresholds set manually and are difficult to cope with the dynamic changes in the statistical characteristics of different data segments. Outliers are easily misjudged or missed. Therefore, how to set dynamic adaptive thresholds that balance global accuracy and local sensitivity to improve the accuracy of anomaly rejection in distributed data processing has become a difficult problem faced by the industry. Summary of the Invention

[0004] This application provides a distributed data processing method, system, device, and storage medium, which can set dynamic adaptive thresholds that balance global accuracy and local sensitivity to improve the accuracy of anomaly rejection in distributed data processing.

[0005] In a first aspect, this application provides a distributed data processing method. The processing method includes the following steps:

[0006] Obtain industrial data packets to be processed in a distributed system, perform shard data compression on the industrial data packets to be processed, and then obtain a plurality of data compression shards;

[0007] Extract the industrial data covariates corresponding to each data compression slice, determine the distribution density chaos degree of each data compression slice according to the corresponding industrial data covariates, and determine the data covariance threshold value of each data compression slice through the corresponding industrial data covariates and distribution density chaos degree;

[0008] Divide the industrial data in each data compression slice respectively according to the corresponding data covariance threshold value, and then obtain the potential abnormal data of each data compression slice. Determine the abnormal data removal packet in the distributed processing of industrial data through the potential abnormal data of each data compression slice;

[0009] Store the abnormal data removal packet to complete the distributed processing of industrial data.

[0010] In this embodiment, performing slice data compression on the industrial data packet to be processed, and then obtaining multiple industrial slice data specifically includes:

[0011] Slice the data in the industrial data packet to be processed to obtain multiple slice data;

[0012] Determine the data variability and data level of each slice data;

[0013] Convert each slice data into a corresponding data compression slice according to the corresponding data variability and data level, and then obtain multiple data compression slices.

[0014] In this embodiment, extracting the industrial data covariates corresponding to each data compression slice specifically includes:

[0015] Obtain the covariate extraction parameters and spatial spectrum distribution functions corresponding to each data compression slice;

[0016] Determine the industrial data covariates corresponding to each data compression slice respectively according to the corresponding covariate extraction parameters and spatial spectrum distribution functions.

[0017] In this embodiment, determining the distribution density chaos degree of each data compression slice according to the corresponding industrial data covariates specifically includes:

[0018] Determine the distribution density of each data compression slice according to the industrial data covariates corresponding to each data compression slice;

[0019] Determine the distribution density chaos degree of each data compression slice according to the corresponding distribution density.

[0020] In this embodiment, determining the data covariance threshold value of each data compression slice through the corresponding industrial data covariates and distribution density chaos degree specifically includes:

[0021] For each data compression slice, determine the weighted chaos degree of the data compression slice based on the industrial data covariation and distribution density chaos degree corresponding to the data compression slice;

[0022] Determine the data covariation threshold value of the data compression slice according to the weighted chaos degree, and then obtain the data covariation threshold values of each data compression slice.

[0023] In this embodiment, dividing the industrial data in each data compression slice according to the corresponding data covariation threshold value respectively, and then obtaining the potential abnormal data of each data compression slice specifically includes:

[0024] For each data compression slice, obtain the data covariation threshold value of the data compression slice;

[0025] Take the compressed data points in the data compression slice that are greater than the data covariation threshold value as potential abnormal data points, and then obtain multiple potential abnormal data points in the data compression slice;

[0026] Construct the potential abnormal data of the data compression slice through all potential abnormal data points, and then obtain the potential abnormal data of each data compression slice.

[0027] In this embodiment, determining the de - abnormal data packet in the distributed processing of industrial data through the potential abnormal data of each data compression slice specifically includes:

[0028] For the potential abnormal data of each data compression slice, determine the data difference degree between each potential abnormal data point in the potential abnormal data of the data compression slice;

[0029] Determine the abnormal detection limit value corresponding to the potential abnormal data of the data compression slice;

[0030] Determine the de - abnormal data slice corresponding to the data compression slice according to all data difference degrees and the abnormal detection limit value, and then obtain the de - abnormal data slices corresponding to each data compression slice;

[0031] Construct the de - abnormal data packet in the distributed processing of industrial data based on all de - abnormal data slices.

[0032] In a second aspect, the present application provides a distributed data processing system for executing a distributed data processing method. The processing system includes:

[0033] A slice compression module, configured to obtain the industrial data packet to be processed in the distributed system, perform slice - by - slice data compression on the industrial data packet to be processed, and then obtain multiple data compression slices;

[0034] A threshold module, which is used to extract the industrial data covariants corresponding to each data compression slice, determine the distribution density chaos degree of each data compression slice according to the corresponding industrial data covariants, and determine the data covariant threshold value of each data compression slice through the corresponding industrial data covariants and the distribution density chaos degree;

[0035] A de - anomaly module, which is used to divide the industrial data in each data compression slice respectively according to the corresponding data covariant threshold value, so as to obtain the potential abnormal data of each data compression slice, and determine the de - anomaly data packet in the distributed processing of industrial data through the potential abnormal data of each data compression slice;

[0036] A storage module, which is used to store the de - anomaly data packet to complete the distributed processing of industrial data.

[0037] In a third aspect, the present application provides a computer device, which includes a memory and a processor. The memory is used to store a computer program, and the processor is used to call and run the computer program from the memory, so that the computer device executes the above - mentioned distributed data processing method.

[0038] In a fourth aspect, the present application provides a computer - readable storage medium, in which instructions or codes are stored. When the instructions or codes run on a computer, the computer is enabled to execute the above - mentioned distributed data processing method.

[0039] The technical solutions provided by the disclosed embodiments of the present application have the following beneficial effects:

[0040] By obtaining the industrial data packets to be processed in the distributed system, performing slice - by - slice data compression on the industrial data packets to be processed, and then obtaining multiple data compression slices; extracting the industrial data covariants corresponding to each data compression slice, determining the distribution density chaos degree of each data compression slice according to the corresponding industrial data covariants, and determining the data covariant threshold value of each data compression slice through the corresponding industrial data covariants and the distribution density chaos degree; dividing the industrial data in each data compression slice respectively according to the corresponding data covariant threshold value, so as to obtain the potential abnormal data of each data compression slice, and determining the de - anomaly data packet in the distributed processing of industrial data through the potential abnormal data of each data compression slice; storing the de - anomaly data packet to complete the distributed processing of industrial data.

[0041] It can be seen that in this application, first, the industrial data packets to be processed are subjected to fragmented data compression, and multiple data compression fragments are obtained, which can effectively reduce the storage and transmission overhead of the data, while retaining important information and improving the processing efficiency of the system. Then, by extracting the industrial data covariants corresponding to each data compression fragment and combining the distribution density chaos degree to determine the data covariance threshold value, the distribution characteristics and potential anomalies of the data can be captured more accurately. Setting an adaptable data covariance threshold value for each data compression fragment can improve the accuracy of abnormal data detection and can flexibly cope with the changes of different data compression fragments. Finally, based on the data covariance threshold value, the industrial data in the data compression fragments are divided, and potential abnormal data are identified, which can effectively eliminate outliers accurately from large-scale data, improve the quality and accuracy of the data. By analyzing each fragment piece by piece and combining the data characteristics, potential abnormal data are screened out targeted, avoiding the roughness of traditional global threshold setting. Moreover, the setting of the dynamic adaptive threshold (i.e., the abnormal detection limit value) takes into account both global accuracy and local sensitivity, and can flexibly adjust the threshold according to the actual situation of different data fragments, thereby improving the accuracy of abnormal elimination in distributed data processing.

[0042] In summary, the technical solution adopted in this application can set a dynamic adaptive threshold that takes into account both global accuracy and local sensitivity to improve the accuracy of abnormal elimination in distributed data processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0044] Figure 1 is an exemplary flowchart of the distributed data processing method provided by the present application;

[0045] Figure 2 is an exemplary flowchart of determining the distribution density chaos degree of each data compression fragment provided by the present application;

[0046] Figure 3 is an exemplary flowchart of determining the potential abnormal data of each data compression fragment provided by the present application;

[0047] Figure 4 is a module structure diagram of the distributed data processing system provided by the present application;

[0048] Figure 5It is a schematic structural diagram of a computer device for implementing a distributed data processing method provided by this application. Specific embodiments

[0049] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.

[0050] The embodiments of this application provide a distributed data processing method, system, device, and storage medium. The core is to obtain industrial data packets to be processed in a distributed system, perform shard data compression on the industrial data packets to be processed, and then obtain multiple data compression shards; extract the industrial data covariants corresponding to each data compression shard, determine the distribution density chaos degree of each data compression shard according to the corresponding industrial data covariants, and determine the data covariant threshold value of each data compression shard through the corresponding industrial data covariants and distribution density chaos degree; divide the industrial data in each data compression shard according to the corresponding data covariant threshold value, and then obtain the potential abnormal data of each data compression shard, and determine the abnormal-removed data packet in industrial data distributed processing through the potential abnormal data of each data compression shard; store the abnormal-removed data packet to complete the distributed processing of industrial data. By adopting the above solution, a dynamic adaptive threshold that takes into account both global accuracy and local sensitivity can be set to improve the accuracy of abnormal data elimination in distributed data processing.

[0051] Embodiment 1. To better understand the above technical solution, the above technical solution will be described in detail below in conjunction with the accompanying drawings of the specification and specific implementation manners. Refer to Figure 1 As shown, this figure is an exemplary flowchart of a distributed data processing method shown in this embodiment of this application. The processing method includes the following steps:

[0052] In step S1, obtain industrial data packets to be processed in a distributed system, perform shard data compression on the industrial data packets to be processed, and then obtain multiple data compression shards.

[0053] Specifically, the industrial data packets to be processed can be obtained from the memory of the distributed system. Among them, the industrial data packets to be processed contain multiple pieces of industrial data.

[0054] In this embodiment, the specific implementation of performing shard data compression on the industrial data packets to be processed to obtain multiple data compression shards can be achieved by the following steps:

[0055] Fragment the data in the industrial data packet to be processed to obtain multiple data fragments;

[0056] Determine the data variability and data level of each data fragment;

[0057] Convert each data fragment into a corresponding data compression fragment according to the corresponding data variability and data level, thereby obtaining multiple data compression fragments.

[0058] It should be noted that in this application, the data compression fragment is a data fragment obtained by fragmenting and compressing the industrial data in the industrial data packet to be processed; specifically, first, obtain the data packet to be processed from the memory of the distributed system, and according to the preset fragmentation strategy, fragment the industrial data in the industrial data packet to be processed to obtain multiple data fragments. Among them, the preset fragmentation strategy can be set through historical experience; then, the data variability and data level of each data fragment can be determined. Among them, the data variability represents the abnormal fluctuation degree of the overall value of all industrial data points in the data fragment, and the data level is the overall value level of all industrial data points in the corresponding data fragment. For each data fragment, the average value of all industrial data points in the data fragment can be calculated, and the calculation result can be used as the data level of the data fragment. The standard deviation of all industrial data points in the data fragment can be calculated, and the calculation result can be used as the data variability of the data fragment. Through the above method, the data variability and data level of each data fragment can be obtained; finally, each data fragment can be converted into a corresponding data compression fragment according to the corresponding data variability and data level, that is, for each industrial data point of the data fragment, the industrial data point can be converted according to the data variability and data level of the data fragment to obtain the compressed data point corresponding to the industrial data point. In actual implementation, the compressed data point can be determined according to the following formula:

[0059]

[0060] Among them, represents the industrial data point corresponding compressed data point, represents the t-th industrial data point in the data fragment, represents the data level, represents the data variability, represents the low-dimensional space, represents taking a specific value in the low-dimensional space. The specific value in the low-dimensional space can be calibrated as a constant according to historical experience. Through the above steps, all compressed data points can be obtained, and the fragment composed of all compressed data points is used as the data compression fragment corresponding to the data fragment, thereby obtaining multiple data compression fragments.

[0061] It should be noted that fragmenting and compressing the industrial data packets to be processed and obtaining multiple data compression fragments can effectively reduce the storage and transmission overhead of the data, while retaining important information and improving the processing efficiency of the system. The compression of each data fragment can adjust the compression strategy according to its data variability and data level, so as to optimize the data storage method and improve the data transmission speed.

[0062] In step S2, extract the industrial data covariates corresponding to each data compression fragment, determine the distribution density chaos degree of each data compression fragment according to the corresponding industrial data covariates, and determine the data covariance threshold value of each data compression fragment through the corresponding industrial data covariates and distribution density chaos degree.

[0063] In this embodiment, extracting the industrial data covariates corresponding to each data compression fragment can be specifically implemented by the following steps:

[0064] Obtain the covariate extraction parameters and spatial spectrum distribution functions corresponding to each data compression fragment;

[0065] Determine the industrial data covariates corresponding to each data compression fragment according to the corresponding covariate extraction parameters and spatial spectrum distribution functions respectively.

[0066] It should be noted that in this application, the covariate extraction parameters are a set of numerical values used to quantitatively describe the behavioral characteristics extracted from the original data; the spatial spectrum distribution function is a function representing the distribution of covariates on the spatial spectrum; the industrial data covariates are feature vectors used to represent important information in the industrial data.

[0067] Specifically, first, the mean, variance, and skewness of each data compression fragment can be calculated, and the corresponding mean, variance, and skewness can be used as the specific numerical values in the covariate extraction parameters; then, the data compression fragment can be subjected to a fast Fourier transform to obtain its spectrum, and thus the square of the absolute value of its spectrum can be calculated, which can be used as the spatial spectrum distribution function. Through the above method, the covariate extraction parameters and spatial spectrum distribution functions corresponding to each data compression fragment can be obtained; finally, for each data compression fragment, the feature vector composed of the covariate extraction parameters and spatial spectrum distribution functions corresponding to the data compression fragment can be used as the industrial data covariates corresponding to the data compression fragment. Through the above method, the industrial data covariates corresponding to each data compression fragment can be obtained.

[0068] Preferably, in this embodiment, referring to Figure 2 As shown, this figure is an exemplary flowchart for determining the distribution density chaos degree of each data compression fragment in the embodiment of this application. In this embodiment, determining the distribution density chaos degree of each data compression fragment according to the corresponding industrial data covariates can be specifically implemented by the following steps:

[0069] In step S21, the distribution density of each data compression slice is determined based on the industrial data covariates corresponding to each data compression slice;

[0070] In step S22, the distribution density chaos degree of each data compression slice is determined according to the corresponding distribution density.

[0071] Specifically, first, the distribution density of each data compression slice can be determined based on the industrial data covariates corresponding to each data compression slice. Here, the distribution density refers to the degree of aggregation of the compressed data points in each data compression slice in its feature space. In actual implementation, the distribution density can be determined by the following formula:

[0072]

[0073] where, represents the distribution density of data compression slice x; represents the total number of features in the industrial data covariates corresponding to data compression slice x; represents the value of the m-th feature in the industrial data covariates; then, the distribution density chaos degree of each data compression slice can be determined according to the corresponding distribution density. Here, the distribution density chaos degree represents the degree of chaos of the information contained in the data compression slice relative to the historical information. In actual implementation, the historical distribution density data of this data compression slice can be obtained, so that the ratio of the distribution density of this data compression slice to the average value of all distribution densities in the historical distribution density data can be calculated, and the calculation result is used as the distribution density chaos degree of this data compression slice. Through the above steps, the distribution density chaos degrees of all data compression slices can be obtained.

[0074] In this embodiment, to determine the data covariance threshold value of each data compression slice through the corresponding industrial data covariates and distribution density chaos degree, the following steps can be specifically adopted:

[0075] For each data compression slice, the weighted chaos degree of the data compression slice is determined through the industrial data covariates and distribution density chaos degree corresponding to the data compression slice;

[0076] Based on the weighted chaos degree, the data covariance threshold value of the data compression slice is determined, and then the data covariance threshold values of all data compression slices are obtained.

[0077] In specific implementation, first, each feature in the industrial data covariate corresponding to the data compression shard is multiplied by the distribution density confusion degree, and the obtained results are summed up, so that the sum result is used as the weighted confusion degree of the data compression shard. It should be noted that in this application, the weighted confusion degree is a comprehensive index used to reflect the sensitivity of the importance of actual data; then, the ratio of the weighted confusion degree to the number of data points in the data compression shard can be calculated, so that the calculation result is used as the data covariance threshold value of the data compression shard. Performing the above steps on all data compression shards can obtain the data covariance threshold values of each data compression shard. It should be noted that in this application, the data covariance threshold value is a threshold parameter used to measure the degree of data abnormality in the data compression shard.

[0078] It should be noted that by extracting the industrial data covariates corresponding to each data compression shard and combining the distribution density confusion degree to determine the data covariance threshold value, the distribution characteristics and potential anomalies of the data can be captured more accurately. By dynamically evaluating the characteristics of each data compression shard, an adaptable data covariance threshold value can be set for each data compression shard, which can improve the accuracy of abnormal data detection, can flexibly cope with the changes of different data compression shards, enable the system to ensure the overall data accuracy when processing large-scale distributed data, and can sensitively detect abnormal data in local areas, thereby significantly improving the accuracy of abnormal data removal and reducing the occurrence of false detection and missed detection.

[0079] In step S3, the industrial data in each data compression shard is divided according to the corresponding data covariance threshold value, and then the potential abnormal data of each data compression shard is obtained. The de-anomaly data packet in the distributed processing of industrial data is determined through the potential abnormal data of each data compression shard.

[0080] Preferably, in this embodiment, with reference to Figure 3 As shown, this figure is an exemplary flowchart for determining the potential abnormal data of each data compression shard in the embodiment of this application. In this embodiment, the industrial data in each data compression shard is divided according to the corresponding data covariance threshold value, and then the potential abnormal data of each data compression shard can be specifically obtained by the following steps:

[0081] In step S31, for each data compression shard, the data covariance threshold value of the data compression shard is obtained;

[0082] In step S32, the compressed data points in the data compression shard that are greater than the data covariance threshold value are used as potential abnormal data points, and then multiple potential abnormal data points in the data compression shard are obtained;

[0083] In step S33, potential abnormal data of the data compression shard is constructed from all potential abnormal data points, and then potential abnormal data of each data compression shard is obtained.

[0084] When specifically implemented, first, the data covariance threshold value of the data compression shard can be relied on; then, the compressed data points in the data compression shard are compared with the data covariance threshold value of the data compression shard, and the compressed data points in the data compression shard that are greater than the data covariance threshold value are used as potential abnormal data points. The data covariance threshold value is calculated by combining the covariance variable of the shard and the distribution density confusion degree, and it is the "fluctuation tolerance boundary" of the shard data in the statistical sense. The data within this boundary is considered reasonable fluctuation, that is, the compressed data points in the data compression shard that are greater than the data covariance threshold value are used as potential abnormal data points. It should be noted that in this application, a potential abnormal data point refers to a data point in the data compression shard that significantly deviates from the overall data point distribution pattern and has the possibility of being abnormal but has not been completely determined to be abnormal; finally, the data set composed of all potential abnormal data points is used as the potential abnormal data of the data compression shard. Repeating the above steps can obtain the potential abnormal data of each data compression shard.

[0085] In this embodiment, the abnormal data removal packet in the industrial data distributed processing can be determined by the potential abnormal data of each data compression shard, and the following steps can be specifically adopted to implement it:

[0086] For the potential abnormal data of each data compression shard, determine the data difference degree between each potential abnormal data point in the potential abnormal data of the data compression shard;

[0087] Determine the abnormal detection limit value corresponding to the potential abnormal data of the data compression shard;

[0088] According to all the data difference degrees and the abnormal detection limit value, determine the abnormal data removal shard corresponding to the data compression shard, and then obtain the abnormal data removal shard corresponding to each data compression shard;

[0089] Construct an abnormal data removal packet in the industrial data distributed processing based on all the abnormal data removal shards.

[0090] In specific implementation, first, for each potential abnormal data point in the potentially abnormal data of the data compression shard, the Euclidean distance between each potential abnormal data point can be calculated, and the corresponding Euclidean distance can be used as the data difference degree between the potential abnormal data points. Here, the data difference degree is an index for measuring the degree of data difference between potential abnormal data points. Then, the data mean and data standard deviation of the potentially abnormal data of the data compression shard can be calculated, and the sum of the data mean and three times the data standard deviation can be used as the abnormal detection limit corresponding to the potentially abnormal data of this data compression shard. Here, the abnormal detection limit is a threshold for determining whether a potential abnormal data point is abnormal.

[0091] In addition, in specific implementation, first, each potential abnormal data point in the potentially abnormal data of the data compression shard can be compared with the corresponding abnormal detection limit, and the potential abnormal data points greater than the abnormal detection limit can be used as abnormal data points, so as to obtain the first set of abnormal data points of this data compression shard. Then, the data difference degree between each potential abnormal data point can be used as the input of the local outlier factor algorithm, and all the abnormal data points in the potentially abnormal data of this data compression shard can be screened out through the local outlier factor algorithm, and all the screened abnormal data points can be used as the second set of abnormal data points. Finally, all the abnormal data points included in the first set of abnormal data points and the second set of abnormal data points can be removed in this data compression shard, and the data compression shard after removing the abnormal data points can be used as the corresponding abnormal-removed data shard. By repeating the above steps, the abnormal-removed data shards corresponding to each data compression shard can be obtained, and the data packet composed of all the abnormal-removed data shards can be used as the abnormal-removed data packet in the distributed processing of industrial data. It should be noted that in this application, the abnormal-removed data packet refers to the industrial data packet obtained after abnormal data removal.

[0092] It should be noted that dividing the industrial data in the data compression shard according to the data covariance threshold and identifying the potentially abnormal data can effectively and accurately remove the outliers from the large-scale data, improving the quality and accuracy of the data. By analyzing each shard piece by piece and combining the data characteristics, the potentially abnormal data can be screened out targeted, avoiding the coarseness of the traditional global threshold setting. Moreover, the setting of the dynamic adaptive threshold (i.e., the abnormal detection limit) takes into account both the global accuracy and local sensitivity, and can flexibly adjust the threshold according to the actual situation of different data shards, thereby improving the accuracy of abnormal data removal.

[0093] In step S4, the abnormal-removed data packet is stored to complete the distributed processing of industrial data.

[0094] In this embodiment, storing the abnormal-removed data packet means storing the abnormal-removed data packet back in the distributed system, and thus the distributed processing of industrial data can be completed.

[0095] It can be seen that in this application, first, the industrial data packets to be processed are subjected to sliced data compression to obtain multiple data compression slices, which can effectively reduce the storage and transmission overhead of the data, while retaining important information and improving the processing efficiency of the system; then, by extracting the industrial data covariants corresponding to each data compression slice and combining the distribution density chaos degree to determine the data covariance threshold value, the distribution characteristics and potential anomalies of the data can be captured more accurately. Setting an adaptable data covariance threshold value for each data compression slice can improve the accuracy of abnormal data detection and can flexibly handle the changes of different data compression slices; finally, dividing the industrial data in the data compression slices according to the data covariance threshold value and identifying potential abnormal data can effectively eliminate outliers accurately from large-scale data, improve the quality and accuracy of the data. By analyzing each slice and combining the data characteristics, potential abnormal data can be screened out specifically, avoiding the coarseness of the traditional global threshold setting. Moreover, the setting of the dynamic adaptive threshold (i.e., the abnormal detection limit) takes into account both global accuracy and local sensitivity and can flexibly adjust the threshold according to the actual situation of different data slices, thereby improving the accuracy of abnormal elimination in distributed data processing.

[0096] In summary, the technical solution adopted in this application can set a dynamic adaptive threshold that takes into account both global accuracy and local sensitivity to improve the accuracy of abnormal elimination in distributed data processing.

[0097] Embodiment 2. This application provides a distributed data processing system. Refer to Figure 4 As shown, this figure is a module structure diagram of the processing system according to this embodiment of this application. The processing system includes:

[0098] A slicing compression module 100, configured to obtain the industrial data packets to be processed in the distributed system, perform sliced data compression on the industrial data packets to be processed, and thus obtain multiple data compression slices;

[0099] A threshold module 200, configured to extract the industrial data covariants corresponding to each data compression slice, determine the distribution density chaos degree of each data compression slice according to the corresponding industrial data covariants, and determine the data covariance threshold value of each data compression slice through the corresponding industrial data covariants and the distribution density chaos degree;

[0100] An outlier removal module 300, configured to divide the industrial data in each data compression slice respectively according to the corresponding data covariance threshold value, and thus obtain the potential abnormal data of each data compression slice, and determine the outlier removal data packet in the distributed processing of industrial data through the potential abnormal data of each data compression slice;

[0101] A storage module 400 is used to store the de - duplicated data packets, thereby completing the distributed processing of industrial data.

[0102] The above has introduced in detail the examples of the distributed data processing method and system provided by the embodiments of the present application. It can be understood that, correspondingly, in order to implement the above functions, the device includes the corresponding hardware structure and / or software module for executing each function. Those skilled in the art should easily realize that, combining the units and algorithm steps of each example described in the embodiments disclosed in this article, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0103] Embodiment 3: The present application further provides a computer device, which includes a memory and a processor. The memory is used to store a computer program, and the processor is used to call and run the computer program from the memory, so that the computer device executes the above - mentioned distributed data processing method.

[0104] In this embodiment, referring to Figure 5 , the dotted line in this figure indicates that the unit or the module is optional. This figure is a schematic structural diagram of a computer device for a distributed data processing method provided by the embodiments of the present application. The above - mentioned distributed data processing method in the above - mentioned embodiments can be implemented by Figure 5 the computer device shown. The computer device includes at least one processor 501, a memory 502, and at least one communication unit 505. The computer device can be a terminal device, a server, or a chip.

[0105] The processor 501 can be a general - purpose processor or a special - purpose processor. For example, the processor 501 can be a central processing unit (CPU). The CPU can be used to control the computer device, execute software programs, and process the data of software programs. The computer device can also include a communication unit 505 for realizing the input (receiving) and output (sending) of signals.

[0106] For example, the computer device can be a chip. The communication unit 505 can be the input and / or output circuit of the chip, or the communication interface of the chip. The chip can be a component of a terminal device, a network device, or other devices.

[0107] For another example, the computer device may be a terminal device or a server, the communication unit 505 may be a transceiver of the terminal device or the server, or the communication unit 505 may be a transceiver circuit of the terminal device or the server.

[0108] The computer device 500 may include one or more memories 502, on which a program 504 is stored. The program 504 can be run by the processor 501 to generate instructions 503, so that the processor 501 executes the method described in the above method embodiments according to the instructions 503. Optionally, data (such as a target audit model) may also be stored in the memory 502. Optionally, the processor 501 may also read the data stored in the memory 502. This data may be stored at the same storage address as the program 504, or it may be stored at a different storage address from the program 504.

[0109] The processor 501 and the memory 502 may be provided separately or integrated together. For example, they may be integrated on a system on chip (SOC) of the terminal device.

[0110] It should be understood that the steps of the above method embodiments can be completed by a logic circuit in hardware form or instructions in software form in the processor 501. The processor 501 may be a central processing unit, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices. For example, discrete gate, transistor logic devices, or discrete hardware components.

[0111] Those skilled in the art should understand that the embodiments of the present application may be provided as a method, a system, or a computer program product. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0112] Embodiment 4, the present application also provides a computer-readable storage medium, in which instructions or code are stored. When the instructions or code run on a computer, the computer is caused to execute a distributed data processing method as described above.

[0113] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present application.

[0114] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.

Claims

1. A distributed data processing method, characterized in that, The processing method includes the following steps: Obtain the industrial data packets to be processed in the distributed system, perform fragmented data compression on the industrial data packets to be processed, and then obtain multiple data compression fragments; Extract the industrial data covariants corresponding to each data compression fragment, determine the distribution density chaos degree of each data compression fragment according to the corresponding industrial data covariants, and determine the data covariant threshold value of each data compression fragment through the corresponding industrial data covariants and distribution density chaos degree; Divide the industrial data in each data compression fragment respectively according to the corresponding data covariant threshold value, and then obtain the potential abnormal data of each data compression fragment, and determine the abnormal data removal packet in the distributed processing of industrial data through the potential abnormal data of each data compression fragment; Store the abnormal data removal packet to complete the distributed processing of industrial data.

2. The distributed data processing method according to claim 1, characterized in that, Performing fragmented data compression on the industrial data packets to be processed, and then obtaining multiple data compression fragments specifically includes: Fragment the data in the industrial data packets to be processed to obtain multiple fragmented data; Determine the data variability and data level of each fragmented data; Convert each fragmented data into a corresponding data compression fragment according to the corresponding data variability and data level, and then obtain multiple data compression fragments.

3. The distributed data processing method according to claim 1, characterized in that Extracting the industrial data covariants corresponding to each data compression fragment specifically includes: Obtain the covariant extraction parameters and spatial spectrum distribution functions corresponding to each data compression fragment; Determine the industrial data covariants corresponding to each data compression fragment respectively according to the corresponding covariant extraction parameters and spatial spectrum distribution functions.

4. A distributed data processing method according to claim 1, characterized in that Determining the distribution density chaos degree of each data compression fragment according to the corresponding industrial data covariants specifically includes: Determine the distribution density of each data compression fragment according to the industrial data covariants corresponding to each data compression fragment; Determine the distribution density chaos degree of each data compression fragment according to the corresponding distribution density.

5. A distributed data processing method according to claim 1, characterized in that, Determining the data covariant threshold value of each data compression fragment through the corresponding industrial data covariants and distribution density chaos degree specifically includes: For each data compression fragment, determine the weighted chaos degree of the data compression fragment through the industrial data covariants and distribution density chaos degree corresponding to the data compression fragment; Determine the data covariant threshold value of the data compression fragment according to the weighted chaos degree, and then obtain the data covariant threshold values of each data compression fragment.

6. A distributed data processing method according to claim 1, characterized in that, Dividing the industrial data in each data compression fragment respectively according to the corresponding data covariant threshold value, and then obtaining the potential abnormal data of each data compression fragment specifically includes: For each data compression fragment, obtain the data covariant threshold value of the data compression fragment; Take the compressed data points in the data compression fragment that are greater than the data covariant threshold value as potential abnormal data points, and then obtain multiple potential abnormal data points in the data compression fragment; Construct the potential abnormal data of the data compression fragment through all potential abnormal data points, and then obtain the potential abnormal data of each data compression fragment.

7. A distributed data processing method according to claim 1, characterized in that, Determining the abnormal data removal packet in the distributed processing of industrial data through the potential abnormal data of each data compression fragment specifically includes: For the potential abnormal data of each data compression slice, determine the data difference degree between each potential abnormal data point in the potential abnormal data of the data compression slice; Determine the abnormal detection limit value corresponding to the potential abnormal data of the data compression slice; Based on all the data difference degrees and the abnormal detection limit value, determine the abnormal data removal slice corresponding to the data compression slice, and then obtain the abnormal data removal slices corresponding to each data compression slice; Construct an abnormal data removal packet in industrial data distributed processing based on all the abnormal data removal slices.

8. A distributed data processing system for performing a distributed data processing method according to any one of claims 1 to 7, characterized in that, The processing system includes: A slice compression module, configured to obtain an industrial data packet to be processed in a distributed system, perform slice data compression on the industrial data packet to be processed, and then obtain a plurality of data compression slices; A threshold module, configured to extract the industrial data covariants corresponding to each data compression slice, determine the distribution density confusion degree of each data compression slice according to the corresponding industrial data covariants, and determine the data covariance threshold values of each data compression slice through the corresponding industrial data covariants and the distribution density confusion degree; An abnormal data removal module, configured to divide the industrial data in each data compression slice respectively according to the corresponding data covariance threshold values, and then obtain the potential abnormal data of each data compression slice, and determine the abnormal data removal packet in industrial data distributed processing through the potential abnormal data of each data compression slice; A storage module, configured to store the abnormal data removal packet, and complete the distributed processing of industrial data.

9. A computer device, characterized in that, The computer device includes a memory and a processor. The memory is used to store a computer program, and the processor is used to call and run the computer program from the memory, so that the computer device executes a distributed data processing method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, Instructions or codes are stored in the computer-readable storage medium. When the instructions or codes run on a computer, the computer is caused to execute a distributed data processing method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Abnormal data detection method and device, equipment and storage medium

    CN111931860A

  • Abnormality detection method, electronic equipment and storage medium

    CN115080289A

  • Internet of Things information platform and implementation method thereof

    CN120017670A

  • Method and apparatus for detecting abnormality using time-series data

    KR1020170084445A

  • Scalable and real-time anomaly detection

    US20190349247A1