A Multilevel Heterogeneous Data Processing Method and Device Based on Big Data Processing Technology
Through a multi-level heterogeneous data processing method based on big data processing technology, data cleaning, fusion and clustering are used to clean, fusion and cluster data, the problems of complex data integration and inaccurate decision-making in traditional methods are solved, and data quality and decision-making accuracy are improved.
Patent Information
- Application Number
- CN202411136132.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-19
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2044-08-19
AI Technical Summary
Traditional methods are difficult to effectively process multi-level and multi-type multi-level heterogeneous data, resulting in complex data integration, loss of information, and incomplete and accurate decisions.
Through the multi-level heterogeneous data processing method, multiple sensor devices are used to obtain data and represent it as a data table, and the similarity between data is calculated, and the data is cleaned and fusion is used to perform data making fusion and clustering, and standard multi-level heterogeneous data is generated.
Improve data quality and accuracy, reduce manual intervention, improve processing efficiency, and extract key information through data fusion and feature fusion, enhance the comprehensive expressiveness and decision-making accuracy of the data, and be able to identify potential patterns and trends.
Smart Images

Figure CN119089114B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a multi-level heterogeneous data processing method and device based on big data processing technology. Background Art
[0002] Multi-level heterogeneous data refers to data from different sources and types, which may differ in structure, format, attributes and hierarchy. It includes structured data, semi-structured data and unstructured data, which usually need to be integrated and utilized through advanced processing and analysis technologies.
[0003] Traditional methods often rely on manual cleaning or rule-based methods, which are slow and error-prone. In addition, the formats and attributes of different data sources in traditional methods are inconsistent, making data integration and comparison complex and cumbersome. In addition, traditional methods may not be able to effectively handle multi-level and multi-type data fusion, and are prone to information loss. In addition, the decisions made by traditional methods based on a single data source or simple processing may not be comprehensive and accurate enough. Traditional methods may not be able to capture complex patterns or hidden trends in data clustering and pattern recognition. Summary of the Invention
[0004] The purpose of the present invention is to address the problems existing in the background technology and propose a multi-level heterogeneous data processing method and device based on big data processing technology.
[0005] The technical solution of the present invention is a multi-level heterogeneous data processing method based on big data processing technology, comprising:
[0006] Acquire multi-level heterogeneous data through multiple sensor devices, represent the multi-level heterogeneous data in a data table format, and acquire attributes of the heterogeneous data based on the type of the sensor device;
[0007] Calculating similarities between the multi-level heterogeneous data based on the attributes of the heterogeneous data, and performing data cleaning on the multi-level heterogeneous data using a trained neural network-based data cleaning model to obtain standard multi-level heterogeneous data;
[0008] Performing data fusion on the standard multi-level heterogeneous data to obtain a first data fusion result, and performing feature fusion on the standard multi-level heterogeneous data to obtain a second data fusion result;
[0009] performing decision fusion on the standard multi-level heterogeneous data based on the first data fusion result and the second data fusion result to obtain a third data fusion result;
[0010] The third data fusion result is clustered to obtain a plurality of clusters corresponding to the standard multi-level heterogeneous data, and the standard multi-level heterogeneous data is divided based on the plurality of clusters to obtain a plurality of pattern data sets.
[0011] Preferably, the similarity calculation formula between the multi-level heterogeneous data is as follows:
[0012] ;
[0013] in, Representing heterogeneous data and heterogeneous data The similarity between Representing heterogeneous data and heterogeneous data The normalized Manhattan distance between , Represents heterogeneous data dimensions, Represents the attributes of heterogeneous data, and Represent heterogeneous data The corresponding data table's attributes and heterogeneous data The corresponding attributes of the corresponding data table.
[0014] Preferably, the neural network-based data cleaning model includes an input layer, multiple hidden layers and an output layer, the number of neurons in the input layer corresponds to the number of attributes of the multi-level heterogeneous data, the node weights between the input layer and the hidden layer correspond to the similarity between the multi-level heterogeneous data, and the output layer is used to output standard multi-level heterogeneous data.
[0015] Preferably, the training method of the neural network-based data cleaning model includes:
[0016] Construct a node weight matrix between the levels of the neural network-based data cleaning model. The node weight matrix is as follows:
[0017] ;
[0018] in, represents the node weight matrix, Represents the input layer of the data cleaning model based on neural network nodes and The weight vector between hidden layers, Indicates the number of hidden layers in the neural network-based data cleaning model;
[0019] Based on the node weight matrix, a node energy function of the multi-level heterogeneous data is obtained, and the node energy function is as follows:
[0020] ;
[0021] in, Indicates the first Node data, Indicates the first Node data, Indicates the first The bias coefficient of the node data, Indicates the first The bias coefficient of the node data, Indicates the first Node data and Node energy between node data;
[0022] The data distribution probability of the nodes in the input layer and the nodes in the hidden layer is obtained based on the node energy function. The data distribution probability calculation formula is as follows:
[0023] ;
[0024] in, represents the distribution factor, Represents the probability of data distribution;
[0025] After activating the nodes in the hidden layer of the neural network-based data cleaning model based on the above operations, the node data in the input layer performs back propagation on the neural network-based data cleaning model to obtain a trained neural network-based data cleaning model.
[0026] Preferably, performing data fusion on the standard multi-level heterogeneous data to obtain a first data fusion result includes:
[0027] The total mean square error (TMSE) of the sensor device is calculated. The total mean square error calculation formula is as follows:
[0028] ;
[0029] in, represents the total mean square error, represents an unbiased estimate of multi-level heterogeneous data, express The corresponding weight, express The corresponding weight, Indicates the Metadata collected by sensor devices, Indicates the Metadata collected by sensor devices;
[0030] The optimal weighting factor of the sensor device is calculated based on the total mean square error of the standard multi-level heterogeneous data. The optimal weighting factor calculation formula of the sensor device is as follows:
[0031] ;
[0032] in, Indicates the The optimal weighting factor for each sensor device;
[0033] The first data fusion result is calculated based on the historical data mean of the sensor device, the optimal weighting factor, and the minimum total mean square error. The minimum total mean square error calculation formula is as follows:
[0034] ;
[0035] in, represents the minimum total mean square error;
[0036] The calculation formula of the first data fusion result is as follows:
[0037] ;
[0038] in, represents the first data fusion result, Indicates the The average value of historical data of sensor devices.
[0039] Preferably, performing feature fusion on the standard multi-level heterogeneous data to obtain a second data fusion result includes:
[0040] Based on the variance maximization criterion, the feature matrix composed of the p-dimensional feature vectors of the standard multi-level heterogeneous data is linearly combined by the principal component analysis method to obtain p principal components of the feature. The p principal components of the feature are expressed as follows:
[0041] ;
[0042] in, represents the pth principal component of the feature, represents the pth eigenvector of the feature matrix, represents the coefficients of the feature matrix;
[0043] Calculate the projection of the p principal components of the feature in space. The projection calculation formula in space is as follows:
[0044] ;
[0045] in, and They represent the projection in space, and They represent the vectors corresponding to the projections in space, Indicates the feature vectors, Indicates the feature vectors;
[0046] A feature component is obtained based on the projection of the p principal components of the feature in space, and the feature vector and the feature component are fused to obtain a second data fusion result. The calculation formula of the second data fusion result is as follows:
[0047] ;
[0048] in, represents the second data fusion result, and Respectively represent The eigenvector and The characteristic components of the eigenvector.
[0049] Preferably, the third data fusion result calculation formula is as follows:
[0050] ;
[0051] in, Represents the third data fusion result.
[0052] Preferably, clustering the third data fusion result to obtain a plurality of clusters corresponding to the standard multi-level heterogeneous data includes:
[0053] The separation degree of the third data fusion result under the q-dimensional feature is analyzed by the data clustering silhouette coefficient, and the data clustering silhouette coefficient is as follows:
[0054] ;
[0055] in, represents the interval of clustering stability in q-dimensional features, represents the interval of cluster transition state in q-dimensional features, The number of cluster sequences representing the stable state, It represents the data clustering silhouette coefficient of the heterogeneous data points in the q-dimensional feature in the third data fusion result. The total number of cluster sequences representing stable states, The total dimension of the features representing the third data fusion result;
[0056] Set a sliding window and calculate the average value of the sliding window feature value. The sliding window setting formula is as follows:
[0057] ;
[0058] in, Represents the internal difference of p-dimensional features, represents a sliding window, Indicates the width of the sliding window;
[0059] The calculation formula for the average value of the sliding window feature value is as follows:
[0060] ;
[0061] in, Represents the average value of the sliding window feature;
[0062] The data range within the sliding window is limited based on the data clustering silhouette coefficient. The data range expression is as follows:
[0063] ;
[0064] in, Indicates the range value of the p-dimensional feature within the sliding window.
[0065] Preferably, clustering the third data fusion result to obtain a plurality of clusters corresponding to the standard multi-level heterogeneous data further includes:
[0066] The feature difference is calculated based on the average value of the sliding window feature value, and the cluster center is set based on the feature difference. The feature difference calculation formula is as follows:
[0067] ;
[0068] in, represents the eigenvalue of the initial cluster p-point sliding window, Represents the feature difference between sliding windows;
[0069] The clustering conditions that the cluster centers need to meet are set based on the feature difference and the data range in the sliding window. The clustering conditions are as follows:
[0070] ;
[0071] in, represents the data clustering standard threshold;
[0072] The clustering result is calculated based on the clustering condition. The clustering result calculation formula is as follows:
[0073] ;
[0074] in, Indicates data space category To The Euclidean distance of the data sample points, represents the data sample point, represents the sample point density function, represents the clustering results, Indicates the effective radius of the cluster center.
[0075] The technical solution of the present invention is a multi-stage heterogeneous data processing device based on big data processing technology, which is applicable to the multi-stage heterogeneous data processing method based on big data processing technology, and is characterized by comprising:
[0076] A data acquisition module, configured to acquire multi-level heterogeneous data through multiple sensor devices, represent the data in a data table format, and acquire attributes of the heterogeneous data based on the type of sensor device;
[0077] A data cleaning module, configured to calculate similarities between the multi-level heterogeneous data based on attributes of the heterogeneous data, and to perform data cleaning on the multi-level heterogeneous data using a trained neural network-based data cleaning model to obtain standard multi-level heterogeneous data;
[0078] a data fusion module, the data fusion module being used to perform data fusion on the standard multi-level heterogeneous data to obtain a first data fusion result, and to perform feature fusion on the standard multi-level heterogeneous data to obtain a second data fusion result;
[0079] a decision fusion module, configured to perform decision fusion on the standard multi-level heterogeneous data based on the first data fusion result and the second data fusion result to obtain a third data fusion result;
[0080] A data clustering module is used to cluster the third data fusion result to obtain multiple clusters corresponding to the standard multi-level heterogeneous data, and divide the standard multi-level heterogeneous data based on the multiple clusters to obtain multiple pattern data sets.
[0081] Compared with the prior art, the above technical solution of the present invention has the following beneficial technical effects:
[0082] 1. The present invention uses a trained neural network model to perform data cleaning on multi-level heterogeneous data, which can automatically process noise and errors in the data, thereby improving data quality and accuracy, reducing manual intervention and improving processing efficiency. It also standardizes multi-level heterogeneous data into standard data, making subsequent analysis and processing more consistent and comparable, thereby helping to unify data formats and attributes and reducing problems caused by data inconsistency. In addition, through data fusion and feature fusion methods, data from different sources and types are comprehensively processed. This multi-level fusion method can extract key information from the data from different angles and improve the comprehensive expressiveness of the data.
[0083] 2. The present invention combines data fusion results and feature fusion results to perform decision fusion, which can obtain more comprehensive analysis results, thereby improving the accuracy of decision-making and reducing the deviation that may be caused by a single data source. The fused data can be clustered to divide the data into different clusters. Through this method, potential patterns and trends in the data can be discovered, so as to perform more targeted analysis and application. Moreover, through the analysis of the clustering results, multiple pattern data sets can be identified. These patterns can be used in various application scenarios, such as anomaly detection, trend prediction, personalized recommendations, etc., providing data support for business decisions and improving data-driven decision-making capabilities. BRIEF DESCRIPTION OF THE DRAWINGS
[0084] Figure 1 A schematic flow chart of an overall method in an embodiment of the present invention;
[0085] Figure 2 A schematic flow chart of the overall device in one embodiment of the present invention;
[0086] Figure numerals: 1. Data acquisition module; 2. Data cleaning module; 3. Data fusion module; 4. Decision fusion module; 5. Data clustering module. DETAILED DESCRIPTION
[0087] Example 1, as Figure 1 As shown, the present invention proposes a multi-level heterogeneous data processing method based on big data processing technology, including:
[0088] S1. Acquire multi-level heterogeneous data through multiple sensor devices, represent the multi-level heterogeneous data in a data table format, and obtain attributes of the heterogeneous data based on the type of sensor device;
[0089] S2. Calculate the similarity between multi-level heterogeneous data based on the attributes of the heterogeneous data, and clean the multi-level heterogeneous data using a trained neural network-based data cleaning model to obtain standard multi-level heterogeneous data;
[0090] S3. Performing data fusion on the standard multi-level heterogeneous data to obtain a first data fusion result, and performing feature fusion on the standard multi-level heterogeneous data to obtain a second data fusion result;
[0091] S4. Perform decision fusion on the standard multi-level heterogeneous data based on the first data fusion result and the second data fusion result to obtain a third data fusion result;
[0092] S5. Clustering the third data fusion result to obtain a plurality of clusters corresponding to the standard multi-level heterogeneous data, and dividing the standard multi-level heterogeneous data based on the plurality of clusters to obtain a plurality of pattern data sets.
[0093] In the present invention, multi-level heterogeneous data refers to a collection of data from different sources, formats, and structures; data table format refers to structuring data into a tabular form for easy processing and analysis; attributes refer to the characteristics or descriptions of data, such as temperature sensor readings, timestamps, etc.; similarity calculation refers to measuring the degree of similarity between data to determine the relevance or consistency of data; pattern data set refers to a grouped data set generated based on clustering results, each set represents a pattern or category; clustering refers to grouping data according to similarity to identify natural classifications or patterns in the data; data fusion refers to integrating data from different sources into a unified data set to enhance the overall value of the data.
[0094] In the second embodiment, the present invention proposes a multi-level heterogeneous data processing method based on big data processing technology. Compared with the first embodiment, this embodiment further includes: the similarity calculation formula between the multi-level heterogeneous data is as follows:
[0095] ;
[0096] in, Representing heterogeneous data and heterogeneous data The similarity between Representing heterogeneous data and heterogeneous data The normalized Manhattan distance between , Represents heterogeneous data dimensions, Represents the attributes of heterogeneous data, and Represent heterogeneous data The corresponding data table's attributes and heterogeneous data The corresponding attributes of the corresponding data table.
[0097] In this embodiment, Manhattan distance refers to a method for calculating the distance between two points in a multidimensional space, which is equal to the sum of the absolute values of their differences in each dimension; normalization refers to adjusting data to a standard range to eliminate the impact of different data sizes and facilitate comparison; dimension refers to the number of features of the data; for example, the dimensions of a data point may include different measurements such as temperature and humidity; attributes refer to columns in a data table, representing different characteristics of the data, such as sensor reading type, timestamp, etc.
[0098] In an optional embodiment, a neural network-based data cleaning model includes an input layer, multiple hidden layers and an output layer, the number of neurons in the input layer corresponds to the number of attributes of the multi-level heterogeneous data, the node weights between the input layer and the hidden layer correspond to the similarity between the multi-level heterogeneous data, and the output layer is used to output standard multi-level heterogeneous data.
[0099] It should be noted that a neural network refers to a computational model that mimics the neuronal structure of the human brain and is used to process and analyze complex data.
[0100] In an optional embodiment, a training method for a neural network-based data cleaning model includes:
[0101] A1. Construct a node weight matrix between the layers of the neural network-based data cleaning model. The node weight matrix is as follows:
[0102] ;
[0103] in, represents the node weight matrix, Represents the input layer of the data cleaning model based on neural network nodes and The weight vector between hidden layers, Indicates the number of hidden layers in the neural network-based data cleaning model;
[0104] A2. Based on the node weight matrix, the node energy function of multi-level heterogeneous data is obtained. The node energy function is as follows:
[0105] ;
[0106] in, Indicates the first Node data, Indicates the first Node data, Indicates the first The bias coefficient of the node data, Indicates the first The bias coefficient of the node data, Indicates the first Node data and Node energy between node data;
[0107] A3. Obtain the data distribution probability of the nodes in the input layer and the nodes in the hidden layer based on the node energy function. The data distribution probability calculation formula is as follows:
[0108] ;
[0109] in, represents the distribution factor, Represents the probability of data distribution;
[0110] A4. After activating the nodes in the hidden layer of the neural network-based data cleaning model based on the above operations, the node data in the input layer performs back propagation on the neural network-based data cleaning model to obtain a trained neural network-based data cleaning model.
[0111] In an optional embodiment, performing data fusion on standard multi-level heterogeneous data to obtain a first data fusion result includes:
[0112] B1. Calculate the total mean square error of the sensor device. The total mean square error calculation formula is as follows:
[0113] ;
[0114] in, represents the total mean square error, represents an unbiased estimate of multi-level heterogeneous data, express The corresponding weight, express The corresponding weight, Indicates the Metadata collected by sensor devices, Indicates the Metadata collected by sensor devices;
[0115] B2. Calculate the optimal weighting factor of the sensor device based on the total mean square error of the standard multi-level heterogeneous data. The optimal weighting factor calculation formula of the sensor device is as follows:
[0116] ;
[0117] in, Indicates the The optimal weighting factor for each sensor device;
[0118] B3. Calculate the first data fusion result based on the historical data mean of the sensor device, the optimal weighting factor, and the minimum total mean square error. The minimum total mean square error calculation formula is as follows:
[0119] ;
[0120] in, represents the minimum total mean square error;
[0121] The calculation formula for the first data fusion result is as follows:
[0122] ;
[0123] in, represents the first data fusion result, Indicates the The average value of historical data of sensor devices.
[0124] In an optional embodiment, feature fusion is performed on the standard multi-level heterogeneous data to obtain a second data fusion result, including:
[0125] C1. Based on the variance maximization criterion, the feature matrix composed of the p-dimensional feature vectors of the standard multi-level heterogeneous data is linearly combined by the principal component analysis method to obtain the p principal components of the feature. The p principal components of the feature are expressed as follows:
[0126] ;
[0127] in, represents the pth principal component of the feature, represents the pth eigenvector of the feature matrix, represents the coefficients of the feature matrix;
[0128] C2. Calculate the projection of the p principal components of the feature in space. The projection calculation formula in space is as follows:
[0129] ;
[0130] in, and They represent the projection in space, and They represent the vectors corresponding to the projections in space, Indicates the feature vectors, Indicates the feature vectors;
[0131] C3. Obtain feature components based on the projection of the p principal components of the features in space, fuse the feature vectors and feature components to obtain a second data fusion result. The calculation formula for the second data fusion result is as follows:
[0132] ;
[0133] in, represents the second data fusion result, and Respectively represent The eigenvector and The characteristic components of the eigenvector.
[0134] It should be noted that the variance maximization criterion is used as the standard for selecting principal components. The purpose is to select the direction that maximizes the variance of the data, so that as much data variation information as possible can be retained. In principal component analysis (PCA), the selection of principal components is based on this criterion; principal component analysis (PCA) refers to a dimensionality reduction technique that transforms data from the original feature space to a new feature space through linear transformation, where the new features (principal components) are linear combinations of the original features, and these principal components are arranged according to the size of the variance; the goal of PCA is to reduce the dimension of the data while retaining the original information of the data as much as possible; principal components refer to the new features obtained by PCA transformation, which represent the information with the largest variance in the data, and each principal component is a linear combination of certain original features in the feature matrix; projection in space refers to the projection of data points in the direction of the principal component, that is, mapping data from the original space to the principal component space, and projection is the performance of data in the direction of the principal component.
[0135] In an optional embodiment, the third data fusion result calculation formula is as follows:
[0136] ;
[0137] in, Represents the third data fusion result.
[0138] In an optional embodiment, clustering is performed on the third data fusion result to obtain a plurality of clusters corresponding to the standard multi-level heterogeneous data, including:
[0139] D1. Analyze the separation degree of the third data fusion result under the q-dimensional feature through the data clustering silhouette coefficient. The data clustering silhouette coefficient is as follows:
[0140] ;
[0141] in, represents the interval of clustering stability in q-dimensional features, represents the interval of cluster transition state in q-dimensional features, The number of cluster sequences representing the stable state, It represents the data clustering silhouette coefficient of the heterogeneous data points in the q-dimensional feature in the third data fusion result. The total number of cluster sequences representing stable states, The total dimension of the features representing the third data fusion result;
[0142] D2. Set the sliding window and calculate the average value of the sliding window feature value. The sliding window setting formula is as follows:
[0143] ;
[0144] in, Represents the internal difference of p-dimensional features, represents a sliding window, Indicates the width of the sliding window;
[0145] The calculation formula for the average value of the sliding window feature quantity is as follows:
[0146] ;
[0147] in, Represents the average value of the sliding window feature;
[0148] D3. Limit the data range within the sliding window based on the data clustering silhouette coefficient. The data range expression is as follows:
[0149] ;
[0150] in, Represents the range value of the p-dimensional feature within the sliding window.
[0151] In an optional embodiment, clustering the third data fusion result to obtain a plurality of clusters corresponding to the standard multi-level heterogeneous data further includes:
[0152] D4. Calculate the feature difference based on the average value of the sliding window feature value, and set the cluster center based on the feature difference. The feature difference calculation formula is as follows:
[0153] ;
[0154] in, represents the eigenvalue of the initial cluster p-point sliding window, Represents the feature difference between sliding windows;
[0155] D5. Set the clustering conditions that the cluster center must meet based on the feature difference and the data range within the sliding window. The clustering conditions are as follows:
[0156] ;
[0157] in, represents the data clustering standard threshold;
[0158] D6. Calculate the clustering results based on the clustering conditions. The clustering result calculation formula is as follows:
[0159] ;
[0160] in, Indicates data space category To The Euclidean distance of the data sample points, represents the data sample point, represents the sample point density function, represents the clustering results, Indicates the effective radius of the cluster center.
[0161] Example 3, as Figure 2 As shown, a multi-stage heterogeneous data processing device based on big data processing technology is applicable to the above-mentioned multi-stage heterogeneous data processing method based on big data processing technology, including:
[0162] Data acquisition module 1, which is used to acquire multi-level heterogeneous data through multiple sensor devices, represent the data in a data table format, and acquire attributes of the heterogeneous data based on the type of sensor device;
[0163] Data cleaning module 2 is used to calculate the similarity between multi-level heterogeneous data based on the attributes of the heterogeneous data, and to clean the multi-level heterogeneous data using a trained neural network-based data cleaning model to obtain standard multi-level heterogeneous data;
[0164] The data fusion module 3 is used to perform data fusion on the standard multi-level heterogeneous data to obtain a first data fusion result, and perform feature fusion on the standard multi-level heterogeneous data to obtain a second data fusion result;
[0165] The decision fusion module 4 is used to perform decision fusion on the standard multi-level heterogeneous data based on the first data fusion result and the second data fusion result to obtain a third data fusion result;
[0166] The data clustering module 5 is used to cluster the third data fusion result to obtain multiple clusters corresponding to the standard multi-level heterogeneous data, and divide the standard multi-level heterogeneous data based on the multiple clusters to obtain multiple pattern data sets.
[0167] The embodiments of the present invention are described in detail above with reference to the accompanying drawings, but the present invention is not limited thereto. Various changes can be made within the scope of knowledge possessed by those skilled in the art without departing from the spirit of the present invention.
Claims
1. A multi-level heterogeneous data processing method based on big data processing technology, characterized in that: include: Acquire multi-level heterogeneous data through multiple sensor devices, represent the multi-level heterogeneous data in a data table format, and acquire attributes of the heterogeneous data based on the type of the sensor device; Calculating similarities between the multi-level heterogeneous data based on the attributes of the heterogeneous data, and performing data cleaning on the multi-level heterogeneous data using a trained neural network-based data cleaning model to obtain standard multi-level heterogeneous data; Performing data fusion on the standard multi-level heterogeneous data to obtain a first data fusion result, and performing feature fusion on the standard multi-level heterogeneous data to obtain a second data fusion result; performing decision fusion on the standard multi-level heterogeneous data based on the first data fusion result and the second data fusion result to obtain a third data fusion result; The third data fusion result is clustered to obtain a plurality of clusters corresponding to the standard multi-level heterogeneous data, and the standard multi-level heterogeneous data is divided based on the plurality of clusters to obtain a plurality of pattern data sets.
2. The multi-level heterogeneous data processing method based on big data processing technology according to claim 1, characterized in that: The similarity calculation formula between the multi-level heterogeneous data is as follows: ; in, Representing heterogeneous data and heterogeneous data The similarity between Representing heterogeneous data and heterogeneous data The normalized Manhattan distance between , Represents heterogeneous data dimensions, Represents the attributes of heterogeneous data, and Represent heterogeneous data The corresponding data table's attributes and heterogeneous data The corresponding attributes of the corresponding data table.
3. The multi-level heterogeneous data processing method based on big data processing technology according to claim 2, characterized in that: The neural network-based data cleaning model includes an input layer, multiple hidden layers and an output layer. The number of neurons in the input layer corresponds to the number of attributes of the multi-level heterogeneous data, the node weights between the input layer and the hidden layer correspond to the similarity between the multi-level heterogeneous data, and the output layer is used to output standard multi-level heterogeneous data.
4. The multi-level heterogeneous data processing method based on big data processing technology according to claim 3, characterized in that: The training method of the neural network-based data cleaning model includes: Construct a node weight matrix between the levels of the neural network-based data cleaning model. The node weight matrix is as follows: ; in, represents the node weight matrix, Represents the input layer of the data cleaning model based on neural network nodes and The weight vector between hidden layers, Indicates the number of hidden layers in the neural network-based data cleaning model; Based on the node weight matrix, a node energy function of the multi-level heterogeneous data is obtained, and the node energy function is as follows: ; in, Indicates the first Node data, Indicates the first Node data, Indicates the first The bias coefficient of the node data, Indicates the first The bias coefficient of the node data, Indicates the first Node data and Node energy between node data; The data distribution probability of the nodes in the input layer and the nodes in the hidden layer is obtained based on the node energy function. The data distribution probability calculation formula is as follows: ; in, represents the distribution factor, Represents the probability of data distribution; After activating the nodes in the hidden layer of the neural network-based data cleaning model based on the above operations, the node data in the input layer performs back propagation on the neural network-based data cleaning model to obtain a trained neural network-based data cleaning model.
5. The multi-level heterogeneous data processing method based on big data processing technology according to claim 1, characterized in that: Performing data fusion on the standard multi-level heterogeneous data to obtain a first data fusion result includes: The total mean square error (TMSE) of the sensor device is calculated. The total mean square error calculation formula is as follows: ; in, represents the total mean square error, represents an unbiased estimate of multi-level heterogeneous data, express The corresponding weight, express The corresponding weight, Indicates the Metadata collected by sensor devices, Indicates the Metadata collected by sensor devices; The optimal weighting factor of the sensor device is calculated based on the total mean square error of the standard multi-level heterogeneous data. The optimal weighting factor calculation formula of the sensor device is as follows: ; in, Indicates the The optimal weighting factor for each sensor device; The first data fusion result is calculated based on the historical data mean of the sensor device, the optimal weighting factor, and the minimum total mean square error. The minimum total mean square error calculation formula is as follows: ; in, represents the minimum total mean square error; The calculation formula of the first data fusion result is as follows: ; in, represents the first data fusion result, Indicates the The average value of historical data of sensor devices.
6. The multi-level heterogeneous data processing method based on big data processing technology according to claim 5, characterized in that: Performing feature fusion on the standard multi-level heterogeneous data to obtain a second data fusion result includes: Based on the variance maximization criterion, the feature matrix composed of the p-dimensional feature vectors of the standard multi-level heterogeneous data is linearly combined by the principal component analysis method to obtain p principal components of the feature. The p principal components of the feature are expressed as follows: ; in, represents the pth principal component of the feature, represents the pth eigenvector of the feature matrix, represents the coefficients of the feature matrix; Calculate the projection of the p principal components of the feature in space. The projection calculation formula in space is as follows: ; in, and They represent the projection in space, and They represent the vectors corresponding to the projections in space, Indicates the feature vectors, Indicates the feature vectors; A feature component is obtained based on the projection of the p principal components of the feature in space, and the feature vector and the feature component are fused to obtain a second data fusion result. The calculation formula of the second data fusion result is as follows: ; in, represents the second data fusion result, and Respectively represent The eigenvector and The characteristic components of the eigenvector.
7. The multi-level heterogeneous data processing method based on big data processing technology according to claim 6, characterized in that: The calculation formula of the third data fusion result is as follows: ; in, Represents the third data fusion result.
8. The multi-level heterogeneous data processing method based on big data processing technology according to claim 1, characterized in that: Clustering the third data fusion result to obtain a plurality of clusters corresponding to the standard multi-level heterogeneous data, including: The separation degree of the third data fusion result under the q-dimensional feature is analyzed by the data clustering silhouette coefficient, and the data clustering silhouette coefficient is as follows: ; in, represents the interval of clustering stability in q-dimensional features, represents the interval of cluster transition state in q-dimensional features, The number of cluster sequences representing the stable state, It represents the data clustering silhouette coefficient of the heterogeneous data points in the q-dimensional feature in the third data fusion result. The total number of cluster sequences representing stable states, The total dimension of the features representing the third data fusion result; Set a sliding window and calculate the average value of the sliding window feature value. The sliding window setting formula is as follows: ; in, Represents the internal difference of p-dimensional features, represents a sliding window, Indicates the width of the sliding window; The calculation formula for the average value of the sliding window feature value is as follows: ; in, Represents the average value of the sliding window feature; The data range within the sliding window is limited based on the data clustering silhouette coefficient. The data range expression is as follows: ; in, Indicates the range value of the p-dimensional feature within the sliding window.
9. The multi-level heterogeneous data processing method based on big data processing technology according to claim 8, characterized in that: Clustering the third data fusion result to obtain a plurality of clusters corresponding to the standard multi-level heterogeneous data, further comprising: The feature difference is calculated based on the average value of the sliding window feature value, and the cluster center is set based on the feature difference. The feature difference calculation formula is as follows: ; in, represents the eigenvalue of the initial cluster p-point sliding window, Represents the feature difference between sliding windows; The clustering conditions that the cluster centers need to meet are set based on the feature difference and the data range in the sliding window. The clustering conditions are as follows: ; in, represents the data clustering standard threshold; The clustering result is calculated based on the clustering condition. The clustering result calculation formula is as follows: ; in, Indicates data space category To The Euclidean distance of the data sample points, represents the data sample point, represents the sample point density function, represents the clustering results, Indicates the effective radius of the cluster center.
10. A multi-stage heterogeneous data processing device based on big data processing technology, which is applicable to the multi-stage heterogeneous data processing method based on big data processing technology according to any one of claims 1 to 9, characterized in that: include: A data acquisition module (1), the data acquisition module (1) is used to acquire multi-level heterogeneous data through multiple sensor devices, represent the data in a data table format, and acquire attributes of the heterogeneous data based on the type of the sensor device; A data cleaning module (2), the data cleaning module (2) is used to calculate the similarity between the multi-level heterogeneous data based on the attributes of the heterogeneous data, and to clean the multi-level heterogeneous data using a trained neural network-based data cleaning model to obtain standard multi-level heterogeneous data; A data fusion module (3), the data fusion module (3) is used to perform data fusion on the standard multi-level heterogeneous data to obtain a first data fusion result, and perform feature fusion on the standard multi-level heterogeneous data to obtain a second data fusion result; A decision fusion module (4), the decision fusion module (4) is used to perform decision fusion on the standard multi-level heterogeneous data based on the first data fusion result and the second data fusion result to obtain a third data fusion result; A data clustering module (5) is used to cluster the third data fusion result to obtain a plurality of clusters corresponding to the standard multi-level heterogeneous data, and to divide the standard multi-level heterogeneous data based on the plurality of clusters to obtain a plurality of pattern data sets.
Citation Information
Patent Citations
Multi-source heterogeneous big data fusion method and system in aerospace manufacturing process
CN114841252A
Fault diagnosis method and device based on transformer oil chromatographic analysis
CN115563563A