High-dimension parallel processing method and system based on heterogeneous data set

By preprocessing, clustering, and extracting high-dimensional features from heterogeneous datasets, and combining outlier detection with a distributed computing framework of deep neural networks, the cumbersome and inefficient problems in large-scale data processing are solved, achieving efficient parallel processing and outlier correction.

CN120951188APending Publication Date: 2025-11-14UNIV OF SCI & TECH BEIJING
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510818465.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing large-scale data processing methods rely on complex preprocessing steps, resulting in cumbersome and time-consuming processing. Potential redundant information and noise in high-dimensional data affect the accuracy and stability of the processing results, leading to insufficient model generalization ability and low data processing efficiency.

Method used

By collecting multiple heterogeneous data, preprocessing them, and then using the UK-means clustering algorithm to generate datasets with similar structures, high-dimensional features are extracted. Outlier detection algorithms are then used to identify abnormal data, and a distributed computing framework based on deep neural networks is constructed for parallel processing.

Benefits of technology

It improves computational efficiency and processing power, simplifies the processing procedure, enhances the model's generalization ability, and can more accurately identify and handle complex data anomalies, saving processing time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120951188A_ABST
    Figure CN120951188A_ABST
Patent Text Reader

Abstract

The invention provides a high-dimension parallel processing method and system based on a heterogeneous data set, and relates to the technical field of data processing, and the method comprises the steps: collecting a plurality of heterogeneous data; preprocessing the heterogeneous data; through a U-k-means clustering algorithm, the preprocessed heterogeneous data is classified, and a plurality of heterogeneous data sets containing similar structures are generated; extracting high-dimensional features of each heterogeneous data set; according to the high-dimensional features, an outlier detection ranking is calculated through an outlier detection algorithm; determining abnormal data according to the outlier detection ranking; based on the deep neural network, a distributed computing framework is constructed, and abnormal data are processed in parallel. By means of the method, the calculation efficiency and the processing capacity can be effectively improved, meanwhile, complex data anomalies can be recognized and processed more accurately, the processing process is simplified, the processing time is further saved, and the generalization capacity of the model is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a high-dimensional parallel processing method and system based on heterogeneous datasets. Background Technology

[0002] Heterogeneous datasets refer to data from different sources, of different types, or with different structures. High dimensionality means that the feature dimension in the dataset is very high, that is, each data point contains a large number of features or attributes. Parallel processing methods for high-dimensional heterogeneous datasets improve the efficiency of processing large-scale datasets by utilizing parallel computing.

[0003] With the advent of the big data era, the types and dimensions of data have increased dramatically. Parallel computing and high-dimensional data processing can significantly improve processing efficiency and overcome the bottleneck of single-machine processing. Parallel computing can not only extract valuable information, but also cope with the complexity of high-dimensional data, thus providing strong support for large-scale data analysis, anomaly detection and prediction.

[0004] Existing large-scale data processing methods typically rely on complex preprocessing steps, which can lead to cumbersome and time-consuming processes. Secondly, due to the heterogeneity of data, effectively integrating data from different sources and formats remains a challenge. Although parallel computing can improve processing efficiency, balancing the allocation of computing resources and avoiding communication bottlenecks between nodes remains difficult. Furthermore, potential redundancy and noise in high-dimensional data can also affect the accuracy and stability of the processing results, leading to insufficient model generalization ability and low data processing efficiency. Summary of the Invention

[0005] To address the technical problems that existing large-scale data processing methods typically rely on complex preprocessing steps, which can lead to cumbersome and time-consuming processing, as well as the potential redundancy and noise in high-dimensional data that may affect the accuracy and stability of the processing results, resulting in insufficient model generalization ability and low data processing efficiency, this invention provides a high-dimensional parallel processing method and system based on heterogeneous datasets.

[0006] The technical solutions provided by the embodiments of the present invention are as follows: First aspect: This invention provides a high-dimensional parallel processing method based on heterogeneous datasets, comprising: S1: Collect multiple heterogeneous data sets; S2: Preprocess the heterogeneous data; S3: Using the UK-means clustering algorithm, the preprocessed heterogeneous data is classified to generate multiple heterogeneous datasets with similar structures; S4: Extract high-dimensional features from each heterogeneous dataset; S5: Based on high-dimensional features, determine the outlier detection ranking using an outlier detection algorithm; S6: Identify outlier data based on outlier detection ranking; S7: Based on deep neural networks, a distributed computing framework is built to process abnormal data in parallel.

[0007] The second aspect: This invention provides a high-dimensional parallel processing system based on heterogeneous datasets, comprising: processor; The memory stores computer-readable instructions, which, when executed by the processor, implement the high-dimensional parallel processing method based on heterogeneous datasets as described in the first aspect.

[0008] Third aspect: The present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the high-dimensional parallel processing method based on heterogeneous datasets as described in the first aspect.

[0009] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following: In this embodiment of the invention, multiple heterogeneous data sets are collected and preprocessed to ensure data quality. Then, the preprocessed heterogeneous data is classified using the uk-means clustering algorithm, generating multiple heterogeneous datasets with similar structures. Simultaneously, high-dimensional features of each heterogeneous dataset are extracted to capture deeper patterns. Based on these high-dimensional features, an outlier detection algorithm is used to calculate an outlier ranking, thereby identifying anomalous data. Finally, a distributed computing framework is constructed based on a deep neural network to process the anomalous data in parallel, achieving efficient computation and anomalous data correction. Through parallel processing and high-dimensional feature extraction, computational efficiency and processing power are effectively improved, while complex data anomalies can be identified and processed more accurately, simplifying the processing procedure, further saving processing time, and enhancing the model's generalization ability. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 A flowchart illustrating a high-dimensional parallel processing method based on heterogeneous datasets provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the structure of a high-dimensional parallel processing system based on heterogeneous datasets, provided in an embodiment of the present invention. Detailed Implementation

[0012] The technical solution of the present invention will now be described with reference to the accompanying drawings.

[0013] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.

[0014] In the embodiments of this invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning. Similarly, the terms "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning.

[0015] In this embodiment of the invention, sometimes a subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meaning they express is the same.

[0016] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0017] Reference manual attached Figure 1 The diagram illustrates a flowchart of a high-dimensional parallel processing method based on heterogeneous datasets provided by an embodiment of the present invention.

[0018] This invention provides a high-dimensional parallel processing method based on heterogeneous datasets. This method can be implemented using a high-dimensional parallel processing device based on heterogeneous datasets, which can be a terminal or a server. The processing flow of the high-dimensional parallel processing method based on heterogeneous datasets may include the following steps: S1: Collect multiple heterogeneous data.

[0019] Heterogeneous data refers to data from different sources, formats, or structures. By integrating heterogeneous data from different data sources, a wider range of information and domains can be covered. This approach enhances the flexibility and accuracy of the system when dealing with complex problems because it combines the advantages of multiple data types to comprehensively capture the potential information within the data.

[0020] In one possible implementation, heterogeneous data specifically includes structured data, semi-structured data, and unstructured data.

[0021] Structured data refers to data stored in a fixed format and strictly defined model, typically organized in rows and columns, facilitating processing and access in a database. Common structured data formats include tabular data in relational databases (such as MySQL and Oracle). Semi-structured data refers to data that, while lacking a strict structure and schema, still possesses some form of labeling or identifier to aid in understanding and extracting information. Unstructured data refers to data without a fixed format or predefined model, typically lacking available tags or labels. Its storage and analysis are more complex due to the lack of a clear structure.

[0022] S2: Preprocess the heterogeneous data.

[0023] It should be noted that the data preprocessing methods provide a good foundation for subsequent analysis and modeling, ensuring the high quality and consistency of the data.

[0024] In one possible implementation, preprocessing specifically includes: data cleaning, handling of missing data values, handling of outliers, data format conversion, and standardization.

[0025] Data cleaning refers to identifying and correcting errors, duplicates, missing values, or inconsistencies in the data. Missing value handling involves using appropriate imputation strategies (such as mean imputation, interpolation, etc.) or deleting records containing missing values. Outlier handling involves identifying and correcting extreme anomalies in the dataset. Data format conversion refers to converting data from different sources into a unified format for subsequent processing and analysis. Standardization involves normalizing or standardizing the data to ensure that all feature values ​​have a consistent scale, thereby avoiding unnecessary impacts of certain features on model training.

[0026] S3: Using the UK-means clustering algorithm, the preprocessed heterogeneous data is classified to generate multiple heterogeneous datasets with similar structures.

[0027] Uk-means is an extension algorithm based on K-means clustering. It combines the membership degree and mixing ratio of clusters and iteratively adjusts the membership degree of cluster centers and clusters to divide the data into multiple clusters, with data points within each cluster having high similarity.

[0028] It's worth noting that the UK-means clustering algorithm can efficiently classify preprocessed heterogeneous datasets and generate multiple datasets with similar structures. It not only adjusts the clustering results by calculating cluster membership and mixing ratios, but also adaptively adjusts the number and shape of clusters, avoiding the dependence on the number of clusters found in traditional K-means. Furthermore, UK-means makes the clustering process more flexible and accurate by dynamically adjusting cluster centers and the learning rate.

[0029] In one possible implementation, S3 specifically includes: S301: Initialize the number of clusters, learning rate, and merging ratio, and set the stopping condition threshold and number of iterations.

[0030] S302: Calculate the membership degree of each data point in the heterogeneous dataset based on the number of clusters, learning rate, stopping threshold, and number of iterations.

[0031] in, Indicates the first t During the nth iteration i The data to the first k The membership degree of each cluster center, where ln represents the logarithmic function. x i Indicates the first i One data point, a k Indicates the first k The cluster center of each cluster, , c Indicates the total number of clusters. Indicates the first k The mixing ratio of individual clusters, This represents the learning rate.

[0032] S303: Update the learning rate.

[0033] S304: Update cluster membership and mixing ratio based on the updated learning rate:

[0034] in, Indicates the first t At the +1st iteration, the... i The data to the firstk Membership degree of each cluster center , n This represents the total number of data points in the dataset. Indicates the first t During the nth iteration k The cluster center of each cluster, Indicates the first t During the nth iteration k The mixing ratio of individual clusters, Indicates the first t At the +1st iteration, the... k The mixing ratio of individual clusters, This represents a parameter that controls the updating of the cluster ratio. Indicates the first t During the nth iteration s The mixing ratio of individual clusters.

[0035] S305: Adjust the cluster center based on the updated membership:

[0036] in, Indicates the first t At the +1st iteration, the... k The cluster center of each cluster.

[0037] S306: Discard clusters whose proportion is less than the mixing proportion, and adjust the number of clusters:

[0038] in, Indicates the first t The number of clusters after +1 iterations Indicates the first t The number of clusters in the next iteration.

[0039] S307: Determine if the change in cluster centers is less than the stopping condition threshold; if so, extract multiple datasets containing similar structures based on the clustering results; otherwise, return to step S302 until the maximum number of iterations is reached.

[0040] It's worth noting that the UK-means clustering algorithm offers greater flexibility and adaptability compared to the traditional K-means algorithm. It can dynamically adjust the number of clusters, handle complex data structures, reduce the risk of local optima, adapt to data changes, and provide more accurate clustering results. For large-scale datasets and real-time data analysis tasks, the UK-means algorithm also offers significant computational advantages.

[0041] In one possible implementation, the optimal cluster is determined using an improved greedy algorithm: Initialize the cluster center.

[0042] Calculate the distance between each data point and the cluster center:

[0043] Among them, dist( x i , C k ) represents a data point x i With cluster center C k The distance between them d Indicates the number of dimensions of the data points. x i,j Indicates the first i The first data point j 1 eigenvalue, C k,j Indicates cluster center C k The j Each feature value.

[0044] Based on the distance, select the cluster corresponding to the smallest distance:

[0045] Here, cluster affiliation indicates the cluster to which the data point belongs, and argmin represents the independent variable that takes the minimum value.

[0046] Update the cluster center location:

[0047] in, Indicates the first k The new center of the cluster, S k Indicates the first k The number of data points contained in a cluster.

[0048] The process continues until the change in the cluster center is less than a preset threshold or the maximum number of iterations is reached.

[0049] S4: Extract high-dimensional features from each heterogeneous dataset.

[0050] High-dimensional features refer to multiple attributes or variables contained in a dataset. These attributes or variables are numerous and usually appear when the dimensionality of the dataset is high. By extracting high-dimensional features from heterogeneous datasets, we can deeply mine important information in the data, especially by analyzing event statistics and environmental variables, and comprehensively capture the dynamic behavior and potential patterns of the system.

[0051] In one possible implementation, S4 specifically includes: S401: Divide heterogeneous datasets into events and environment variables.

[0052] In this context, an event typically refers to a key change or behavior within the data. In time series data, event extraction is usually related to specific patterns, thresholds, or significant turning points in the data, while environmental variables refer to external or internal conditions related to the data source.

[0053] S402: Calculate the number of times and duration of events to extract basic statistical information.

[0054] S403: Identify the abnormal distribution of frequency or duration:

[0055] in, Z Representing data points x Its location within its distribution This represents the mean. It represents the standard deviation.

[0056] S404: Calculate the statistical characteristics of environmental variables to extract deeper information from the data.

[0057] S405: Based on basic statistical information and deep data information, high-dimensional features are extracted by calculating the correlation between different events.

[0058] It's worth noting that by calculating the frequency and duration of events, as well as the statistical characteristics of environmental variables, we can more effectively understand the distribution characteristics and trends of the data, and discover potential anomalies and anomalous patterns. This provides a solid data foundation for subsequent cluster analysis and anomaly detection, helping to optimize the accuracy and precision of the model. Furthermore, this method exhibits good adaptability when dealing with high-dimensional datasets, capable of handling large-scale, heterogeneous, and complex data, providing more valuable information for data mining and decision-making.

[0059] In one possible implementation, the statistical characteristics specifically include: mean, variance, skewness, and kurtosis; The formula for calculating the mean is:

[0060] in, This represents the mean. n This indicates the total number of data points.

[0061] The formula for calculating variance is:

[0062] in, Indicates variance.

[0063] The formula for calculating skewness is:

[0064] in, Skewness Indicates skewness.

[0065] The formula for calculating kurtosis is:

[0066] in, Kurtosis Indicates kurtosis.

[0067] S5: Based on the high-dimensional features, determine the outlier detection ranking using an outlier detection algorithm.

[0068] Outlier detection algorithms are used to identify data points in a dataset that are significantly different from other data points. Outliers (or anomalies) may be erroneous data, noise, or, in some cases, represent valuable special events. Outlier detection ranking refers to sorting the data points in a dataset according to their probability (or degree of outlier) as outliers.

[0069] It should be noted that outlier detection algorithms, especially those combining local reachability density and local outlier factor calculations, can effectively identify outliers in the data. This method is particularly suitable for situations where the data distribution is uneven and the density varies greatly.

[0070] In one possible implementation, S5 specifically includes: S501: Calculate the distance between each data point:

[0071] in, Representing data points x i and x j The distance between them x i,m Representing data points x i In the m Feature values ​​in each dimension x j,m Representing data points x j In the m Feature values ​​in each dimension.

[0072] S502: Determine the nearest neighbor based on the distance.

[0073] S503: Calculate the local reachability density based on the nearest neighbor:

[0074] in, lrd ( x i ) represents a data point x i Locally achievable density, N k ( x i ) represents a data point x i The k A group of neighbors, r k ( x j ) indicates the first j The reachability distance of a neighbor.

[0075] S504: Determine the local outlier factor by combining local reachability density and neighbor local reachability density.

[0076] in, LOF ( x i ) represents a data point x i Local outlier lrd ( x j ) represents neighboring points x j The locally achievable density.

[0077] S505: Sort outliers to identify abnormal data.

[0078] It should be noted that by calculating local density, it is possible to flexibly identify data points that may not be significant globally but exhibit anomalous behavior in local areas, thus capturing truly anomalous data more accurately. Furthermore, the outlier detection process can process multiple data points in parallel, improving the ability to handle large-scale datasets and providing a reliable basis for subsequent anomaly analysis, cleanup, and decision support.

[0079] S6: Identify outlier data based on the ranking.

[0080] It's important to note that ranking data points allows us to identify the most anomalous ones based on their anomaly scores, enabling more precise anomaly detection. This process can help quickly identify potential problems from large amounts of data, such as data errors, fraudulent activities, or system malfunctions.

[0081] S7: Based on deep neural networks, a distributed computing framework is built to process abnormal data in parallel.

[0082] Deep neural networks are artificial neural networks containing multiple hidden layers. They can learn the non-linear characteristics of data and exhibit powerful capabilities in handling complex patterns. Distributed computing frameworks refer to technical architectures that use multiple computing nodes to process tasks collaboratively. This allows data to be processed in parallel across multiple machines, thereby improving the efficiency of handling large-scale datasets.

[0083] It should be noted that by combining deep neural networks and distributed computing frameworks, it is possible to effectively process large amounts of abnormal data in parallel.

[0084] In one possible implementation, S7 specifically includes: S701: Reads abnormal data from distributed files and converts the abnormal data into a DataFrame format supported by the deep learning framework.

[0085] S702: Divide the transformed abnormal data into mini-batches of a preset size.

[0086] Mini-batch is a method of dividing a dataset into smaller batches for training. Each mini-batch contains a certain number of data points. During training, one mini-batch is used for forward and backward propagation to update the parameters of the neural network.

[0087] S703: Load the pre-trained bidirectional GRU model, where the update mechanism of the bidirectional GRU model is as follows:

[0088] in, r t Indicates time step t Reset door at time, This represents the Sigmoid activation function. x t Indicates time step t Input error data at that time z t Indicates time step t Timely update gate, h t Indicates time step t The hidden state at that time h t-1 Indicates time step t The hidden state when -1, Indicates time step t Candidate hidden state at time, W r , W z andW h Both represent weight matrices. b r , b z and b h Both represent bias terms.

[0089] It should be noted that bidirectional GRU is an extension of GRU (Gated Recurrent Unit), which enhances the model's learning ability by simultaneously considering both forward and backward input sequences. GRU is widely used in time series and sequential data, and bidirectional GRU can capture both future and past contextual information, improving the processing performance of sequential data.

[0090] S704: Distributes Mini-batch data to various computing nodes via the parameter server for gradient calculation.

[0091] The parameter server is a distributed system used to store and manage the parameters of deep learning models. During training, the parameter server coordinates parameter updates among the computing nodes to ensure consistency and efficiency in training.

[0092] S705: Determine the global gradient based on the gradient calculation results.

[0093] S706: Based on the early stopping strategy, determine whether the bidirectional GRU model has reached the convergence condition; if yes, proceed to step S707; otherwise, proceed to step S704.

[0094] S707: Determine the optimal bidirectional GRU model based on the global gradient and store the optimal bidirectional GRU model in HDFS.

[0095] S708: Transforms anomalous data and a trained bidirectional GRU model into an RDD using Spark programs, and distributes them to various nodes for parallel processing.

[0096] It's worth noting that using mini-batch and distributed computing significantly improves training speed and effectively utilizes computational resources. In particular, the bidirectional GRU model better handles time-series data, capturing dependencies between data points and thus improving anomaly detection accuracy. The parameter server mechanism ensures synchronization between computing nodes, enabling efficient model training in a distributed environment and avoiding the bottleneck of single-node computation.

[0097] In one possible implementation, S7 is followed by: S8: Calculate the similarity between outlier and normal data based on Mahalanobis distance:

[0098] in, Representing data points x and y Mahalanobis distance, S Represents the covariance matrix. -1 Represents the inverse matrix. T This indicates transpose.

[0099] S9: Mark outlier data based on similarity.

[0100] Similarity is used to measure the degree of similarity between two data points or objects. In anomaly detection, labeling usually means marking data points to indicate whether they belong to an anomaly category.

[0101] It's worth noting that similarity-based calculations, particularly using Mahalanobis distance to measure the similarity between anomalous and normal data, offer significant advantages. Mahalanobis distance considers the covariance structure of the data, effectively identifying relationships between data points, and is especially suitable for anomaly detection in high-dimensional datasets. It can more accurately identify anomalous data points that are significantly different from normal data in high-dimensional datasets.

[0102] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following: In this embodiment of the invention, multiple heterogeneous data sets are collected and preprocessed to ensure data quality. Then, the preprocessed heterogeneous data is classified using the uk-means clustering algorithm, generating multiple heterogeneous datasets with similar structures. Simultaneously, high-dimensional features of each heterogeneous dataset are extracted to capture deeper patterns. Based on these high-dimensional features, an outlier detection algorithm is used to calculate an outlier ranking, thereby identifying anomalous data. Finally, a distributed computing framework is constructed based on a deep neural network to process the anomalous data in parallel, achieving efficient computation and anomalous data correction. Through parallel processing and high-dimensional feature extraction, computational efficiency and processing power are effectively improved, while complex data anomalies can be identified and processed more accurately, simplifying the processing procedure, further saving processing time, and enhancing the model's generalization ability.

[0103] Reference manual attached Figure 2 The diagram shows a schematic of the structure of a high-dimensional parallel processing system based on heterogeneous datasets provided by the present invention.

[0104] The present invention also provides a high-dimensional parallel processing system 20 based on heterogeneous datasets, applied to the above-mentioned high-dimensional parallel processing method based on heterogeneous datasets, comprising: Processor 201.

[0105] The memory 202 stores computer-readable instructions, which, when executed by the processor 201, implement the high-dimensional parallel processing method based on heterogeneous datasets as described in the method embodiment.

[0106] The high-dimensional parallel processing system 20 based on heterogeneous datasets provided by the present invention can execute the above-mentioned high-dimensional parallel processing method based on heterogeneous datasets and achieve the same or similar technical effects. To avoid repetition, the present invention will not elaborate further.

[0107] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following: In this embodiment of the invention, multiple heterogeneous data sets are collected and preprocessed to ensure data quality. Then, the preprocessed heterogeneous data is classified using the uk-means clustering algorithm, generating multiple heterogeneous datasets with similar structures. Simultaneously, high-dimensional features of each heterogeneous dataset are extracted to capture deeper patterns. Based on these high-dimensional features, an outlier detection algorithm is used to calculate an outlier ranking, thereby identifying anomalous data. Finally, a distributed computing framework is constructed based on a deep neural network to process the anomalous data in parallel, achieving efficient computation and anomalous data correction. Through parallel processing and high-dimensional feature extraction, computational efficiency and processing power are effectively improved, while complex data anomalies can be identified and processed more accurately, simplifying the processing procedure, further saving processing time, and enhancing the model's generalization ability.

[0108] It should be understood that the processor in the embodiments of the present invention can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.

[0109] It should also be understood that the memory in the embodiments of the present invention can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced synchronous DRAM (ESDRAM), synchronous linked DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0110] The above embodiments can be implemented, in whole or in part, by software, hardware (such as circuits), firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the flow or function according to the embodiments of the present invention is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. Computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. A computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. Available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media. Semiconductor media can be solid-state drives.

[0111] It should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. A and B can be singular or plural. Additionally, the character " / " in this article generally indicates an "or" relationship between the preceding and following related objects, but it can also represent an "and / or" relationship. Please refer to the context for a more accurate understanding.

[0112] In this invention, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of a single item or a plurality of items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be a single item or multiple items.

[0113] It should be understood that, in various embodiments of the present invention, the sequence number of each process does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0114] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0115] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the devices, apparatuses, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0116] In the embodiments provided by this invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0117] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0118] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0119] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0120] This invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the high-dimensional parallel processing method based on heterogeneous datasets as described in the method embodiments.

[0121] The present invention provides a computer-readable storage medium that can implement the steps and effects of the high-dimensional parallel processing method based on heterogeneous datasets in the above-described method embodiments. To avoid repetition, the present invention will not repeat the steps.

[0122] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following: In this embodiment of the invention, multiple heterogeneous data sets are collected and preprocessed to ensure data quality. Then, the preprocessed heterogeneous data is classified using the uk-means clustering algorithm, generating multiple heterogeneous datasets with similar structures. Simultaneously, high-dimensional features of each heterogeneous dataset are extracted to capture deeper patterns. Based on these high-dimensional features, an outlier detection algorithm is used to calculate an outlier ranking, thereby identifying anomalous data. Finally, a distributed computing framework is constructed based on a deep neural network to process the anomalous data in parallel, achieving efficient computation and anomalous data correction. Through parallel processing and high-dimensional feature extraction, computational efficiency and processing power are effectively improved, while complex data anomalies can be identified and processed more accurately, simplifying the processing procedure, further saving processing time, and enhancing the model's generalization ability.

[0123] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

[0124] The following points need to be explained: (1) The accompanying drawings of the embodiments of the present invention only involve the structures involved in the embodiments of the present invention. Other structures can refer to the general design.

[0125] (2) For clarity, the thickness of layers or regions is enlarged or reduced in the drawings used to describe embodiments of the invention, i.e., these drawings are not drawn to scale. It is understood that when an element such as a layer, film, region or substrate is referred to as being “above” or “below” another element, the element may be “directly” located “above” or “below” the other element or there may be intermediate elements.

[0126] (3) Where there is no conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other to obtain new embodiments.

[0127] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. The scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A high-dimensional parallel processing method based on heterogeneous datasets, characterized in that, include: S1: Collect multiple heterogeneous data sets; S2: Preprocess each of the heterogeneous data; S3: Use the Uk-means clustering algorithm to classify the preprocessed heterogeneous data and generate multiple heterogeneous datasets with similar structures; S4: Extract high-dimensional features from each of the heterogeneous datasets; S5: Based on the high-dimensional features, determine the outlier detection ranking using an outlier detection algorithm; S6: Based on the outlier detection ranking, identify abnormal data; S7: Based on deep neural networks, construct a distributed computing framework to process the abnormal data in parallel.

2. The high-dimensional parallel processing method based on heterogeneous datasets according to claim 1, characterized in that, The heterogeneous data specifically includes: structured data, semi-structured data, and unstructured data.

3. The high-dimensional parallel processing method based on heterogeneous datasets according to claim 1, characterized in that, The preprocessing specifically includes: data cleaning, handling of missing data values, handling of outliers, data format conversion, and standardization.

4. The high-dimensional parallel processing method based on heterogeneous datasets according to claim 1, characterized in that, S3 specifically includes: S301: Initialize the number of clusters, learning rate, and merging ratio, and set the stopping condition threshold and number of iterations; S302: Calculate the membership degree of each data point in the heterogeneous dataset based on the number of clusters, the learning rate, the stopping condition threshold, and the number of iterations. in, Indicates the first t During the nth iteration i The data to the first k The membership degree of each cluster center, where ln represents the logarithmic function. x i Indicates the first i One data point, a k Indicates the first k The cluster center of each cluster, , c Indicates the total number of clusters. Indicates the first k The mixing ratio of individual clusters, Indicates the learning rate; S303: Update the learning rate; S304: Update cluster membership and mixing ratio based on the updated learning rate: in, Indicates the first t At the +1st iteration, the... i The data to the first k Membership degree of each cluster center , n This represents the total number of data points in the dataset. Indicates the first t During the nth iteration k The cluster center of each cluster, Indicates the first t During the nth iteration k The mixing ratio of individual clusters, Indicates the first t At the +1st iteration, the... k The mixing ratio of individual clusters, This represents a parameter that controls the updating of the cluster ratio. Indicates the first t During the nth iteration s The mixing ratio of each cluster; S305: Adjust the cluster center according to the updated membership degree: in, Indicates the first t At the +1st iteration, the... k The cluster center of each cluster; S306: Discard clusters with a proportion less than the mixing proportion, and adjust the number of clusters: in, Indicates the first t The number of clusters after +1 iterations Indicates the first t The number of clusters in the next iteration; S307: Determine whether the change in the cluster center is less than the stopping condition threshold; if so, extract multiple datasets containing similar structures based on the clustering results; otherwise, return to step S302 until the maximum number of iterations is reached.

5. The high-dimensional parallel processing method based on heterogeneous datasets according to claim 1, characterized in that, S4 specifically includes: S401: Divide the heterogeneous dataset into events and environment variables; S402: Calculate the number of times the event occurred and the duration to extract basic statistical information; S403: Determine the abnormal distribution of the number of times or the duration: in, Z Representing data points x Its location within its distribution This represents the mean. Indicates standard deviation; S404: Calculate the statistical characteristics of the environmental variables to extract deeper information from the data; S405: Based on the basic statistical information and the deep data information, extract the high-dimensional features by calculating the correlation between different events.

6. The high-dimensional parallel processing method based on heterogeneous datasets according to claim 5, characterized in that, The statistical characteristics specifically include: mean, variance, skewness, and kurtosis.

7. The high-dimensional parallel processing method based on heterogeneous datasets according to claim 1, characterized in that, S5 specifically includes: S501: Calculate the distance between each data point: in, Representing data points x i and x j The distance between them x i,m Representing data points x i In the m Feature values ​​in each dimension x j,m Representing data points x j In the m Feature values ​​in each dimension; S502: Determine the nearest neighbor point based on the distance; S503: Calculate the local reachability density based on the nearest neighbor points: in, lrd ( x i ) represents a data point x i Locally achievable density, N k ( x i ) represents a data point x i The k A group of neighbors, r k ( x j ) indicates the first j The reachability distance of each neighbor; S504: Combine the local reachability density and the neighboring local reachability density to determine the local outlier factor: in, LOF ( x i ) represents a data point x i Local outlier lrd ( x j ) represents neighboring points x j Locally achievable density; S505: Sort the local outlier factors to identify abnormal data.

8. The high-dimensional parallel processing method based on heterogeneous datasets according to claim 1, characterized in that, Specifically, S7 includes: S701: Read the abnormal data from the distributed file and convert the abnormal data into a DataFrame format supported by the deep learning framework; S702: Divide the transformed abnormal data into a Mini-batch of a preset size; S703: Load the pre-trained bidirectional GRU model, wherein the update mechanism of the bidirectional GRU model is as follows: in, r t Indicates time step t Reset door at time, This represents the Sigmoid activation function. x t Indicates time step t Input error data at that time z t Indicates time step t Timely update gate, h t Indicates time step t The hidden state at that time h t-1 Indicates time step t The hidden state when -1, Indicates time step t Candidate hidden state at time, W r , W z and W h Both represent weight matrices. b r , b z and b h Both represent bias terms; S704: Distribute the Mini-batch to each computing node through the parameter server for gradient calculation; S705: Determine the global gradient based on the gradient calculation results; S706: Based on the early stopping strategy, determine whether the bidirectional GRU model has reached the convergence condition; if yes, proceed to step S707; otherwise, proceed to step S704. S707: Determine the optimal bidirectional GRU model based on the global gradient, and store the optimal bidirectional GRU model in HDFS; S708: Transforms anomalous data and a trained bidirectional GRU model into an RDD using Spark programs, and distributes them to various nodes for parallel processing.

9. The high-dimensional parallel processing method based on heterogeneous datasets according to claim 1, characterized in that, Following S7, the following also includes: S8: Calculate the similarity between the abnormal data and the normal data based on Mahalanobis distance: in, Representing data points x and y Mahalanobis distance, S Represents the covariance matrix. -1 Represents the inverse matrix. T Indicates transpose; S9: The abnormal data is marked according to the similarity.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the high-dimensional parallel processing method based on heterogeneous datasets as described in any one of claims 1 to 9.