A multi-source heterogeneous data cleaning method based on generative adversarial network

By generating an adversarial network model to generate missing data filling matrix, the fusion and cleaning of multi-source heterogeneous data in intelligent manufacturing equipment is solved, the accuracy and effectiveness of data cleaning are improved, and data quality is ensured.

CN116166650BActive Publication Date: 2025-08-15CHONGQING UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310137600.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-20
Publication Date
2025-08-15
Estimated Expiration
2043-02-20

AI Technical Summary

Technical Problem

The prior art cannot effectively handle the fusion and cleaning of multi-source heterogeneous data in intelligent manufacturing equipment, resulting in a decline in data quality and affecting production scheduling and equipment management.

Method used

A method based on generative adversarial network is adopted, and redundant and abnormal data are identified and deleted through clustering analysis algorithms, and a missing data filling matrix is generated using the generative adversarial network model to realize the fusion and cleaning of multi-source heterogeneous data.

Benefits of technology

It improves the accuracy and effectiveness of multi-source heterogeneous data cleaning, ensures the quality of intelligent production line data, and reduces the complexity of data processing and calculation difficulty.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116166650B_ABST
    Figure CN116166650B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of intelligent manufacturing technology, and specifically to a multi-source heterogeneous data cleaning method based on a generative adversarial network, comprising: obtaining multi-source heterogeneous data of an intelligent production line, and synthesizing each multi-source heterogeneous data into a corresponding multi-source heterogeneous data fusion table; analyzing the redundant data, abnormal data, and missing data remaining in the multi-source heterogeneous data fusion table through a clustering analysis algorithm, and then determining the multi-source heterogeneous data with missing data; inputting the multi-source heterogeneous data with missing data into a trained generative adversarial network model, and outputting a corresponding missing data filling matrix; filling the multi-source heterogeneous data with missing data through the missing data filling matrix to achieve fusion and cleaning of the multi-source heterogeneous data. The present invention can divide the multi-source heterogeneous data with missing data, and can fill the multi-source heterogeneous data with missing data to achieve fusion and cleaning of the multi-source heterogeneous data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent manufacturing technology, and in particular to a multi-source heterogeneous data cleaning method based on a generative adversarial network. Background Art

[0002] As modern manufacturing rapidly develops toward automation, informatization, and intelligence, the production process generates a vast amount of multi-source, heterogeneous data. Multi-source, heterogeneous data refers to the collection of data from multiple sources and stored in different ways within a manufacturing environment. Due to inherent hardware limitations and the influence of factors such as environmental noise, sensor devices inevitably experience omissions, overreading, and misreading, resulting in reduced sensor data quality. However, data is a carrier of information, and its ability to accurately reflect the real world is crucial to its effectiveness in higher-level applications. Effective processing of multi-source, heterogeneous data can provide manufacturers with more effective strategies for production scheduling and equipment management, thereby improving production quality and efficiency.

[0003] Currently, there are numerous analytical methods for multi-source heterogeneous data from intelligent manufacturing equipment, but few technologies exist for cleaning and filling multi-source heterogeneous data. To address this issue, Chinese patent application number CN112347093A discloses a method for facilitating the cleaning, integration, and storage of massive multi-source heterogeneous data. The method involves constructing a data source set, traversing the set, recording the type and data protocol, performing data access and protocol adaptation, forming first-order data, pushing a cache queue, pulling first-order data, performing a cleaning step, forming second-order data, pushing a cache queue, pulling second-order data, performing an active-passive hybrid mode conversion and integration step, forming third-order data, pushing a cache queue, and finally pulling third-order data, and performing distributed storage to complete the final storage.

[0004] The existing solutions mentioned above primarily reduce the coupling between the cleaning, integration, and storage of massive amounts of heterogeneous, multi-source data, but fail to truly achieve the fusion and cleaning of heterogeneous, multi-source data in intelligent production lines. Therefore, designing a method that can effectively fuse and cleanse heterogeneous, multi-source data in intelligent production lines is a pressing technical challenge. Summary of the Invention

[0005] In view of the above-mentioned deficiencies in the prior art, the technical problem to be solved by the present invention is: how to provide a multi-source heterogeneous data cleaning method based on a generative adversarial network, which can divide multi-source heterogeneous data with missing data, and can fill in multi-source heterogeneous data with missing data to achieve the fusion and cleaning of multi-source heterogeneous data, thereby improving the accuracy and effectiveness of multi-source heterogeneous data cleaning, and providing a more practical solution for ensuring the data quality of multi-source heterogeneous data of intelligent production lines.

[0006] In order to solve the above technical problems, the present invention adopts the following technical solutions:

[0007] Multi-source heterogeneous data cleaning methods based on generative adversarial networks include:

[0008] S1: Obtain multi-source heterogeneous data of the intelligent production line and synthesize each multi-source heterogeneous data into a corresponding multi-source heterogeneous data fusion table;

[0009] S2: Analyze the redundant data, abnormal data, and missing data remaining in the multi-source heterogeneous data fusion table through cluster analysis algorithms, and then determine the multi-source heterogeneous data with missing data;

[0010] S3: Input the multi-source heterogeneous data with missing data into the trained generative adversarial network model and output the corresponding missing data filling matrix;

[0011] S4: Fill in the missing data with the missing data matrix to achieve the fusion and cleaning of multi-source heterogeneous data.

[0012] Preferably, in step S1, the data tables corresponding to the various multi-source heterogeneous data are imported into the data warehouse to synthesize the corresponding multi-source heterogeneous data fusion table; during the synthesis process, abnormal data with association rule errors are detected and deleted.

[0013] Preferably, in step S2, an improved K-means algorithm is selected as the cluster analysis algorithm; the improved K-means algorithm uses the edit distance as a measure of similarity between data, and then uses the characteristic that the farthest data do not belong to the same category to automatically determine the cluster center and the number of clusters through the maximum distance between data.

[0014] Preferably, in step S2, the working logic of the improved K-means algorithm is as follows:

[0015] S201: Convert the record set of the multi-source heterogeneous data fusion table into the corresponding string set A = {a1, a2, ..., a n}, and calculate the edit distance between each two strings in the string set, and generate the corresponding edit distance result set G;

[0016] S202: Select the two data objects with the largest distance in the edit distance result set G. and As the cluster centers of the initial two clusters S1 and S2, and The distance between them is recorded as d1, that is

[0017] Divide the string set A by and Data objects other than and To classify the cluster centers, if a i ∈A and a i and Edit distance Less than Then a i Divide into S1, otherwise a i Divided into S2;

[0018] S203: Get all data objects in S1 to The maximum edit distance d 11 ,Right now Get all data objects in S2 to The maximum edit distance d 22 ,Right now

[0019] Take d 11 and d 22 The maximum value d2, record d2 = max{d 11 ,d 22};

[0020] S204: If d2>k*d1, then take the corresponding data object as the cluster center of the third cluster S3

[0021] by and As the cluster center, clustering is performed to form three clusters S1, S2 and S3. That is, if a i ∈A and Then a i Divided into S1, if a i ∈A and Then a i Divide into S2, otherwise a i Divided into S3;

[0022] S205: Get all data objects in S1 to S1 * The maximum edit distance d 11 ,Right now Get all data objects in S2 to The maximum edit distance d 22 ,Right now Get all data objects in S3 to The maximum edit distance d 33 ,Right now

[0023] Take d 11 d 22 and d 33The maximum value d3, record d3 = max{d 11 ,d 22 ,d 33};

[0024] S206: If d3>k*(d2+d1)÷2, then take the corresponding data object as the cluster center of the fourth cluster S4 by and Clustering is performed as the cluster center to form four clusters S1, S2, S3 and S4;

[0025] S207: Repeat steps S204 to S206 until the clustering condition is no longer met, and output the corresponding data clustering result.

[0026] Preferably, in step S201, the edit distance between two character strings is calculated using the following formula:

[0027]

[0028] Where: Str1 and Str2 represent two character strings, L1 represents the length of Str1, and L2 represents the length of Str2; Indicates the result after randomly deleting a character from Str1; Indicates the result after randomly deleting a character from Str2; The last operation of converting Str1 to Str2 is the edit distance of inserting one character into Str1; Indicates that the last operation in converting Str1 to Str2 is the edit distance of deleting one character from Str2; Indicates that the last operation in converting Str1 to Str2 is the edit distance of replacing one character in Str1; Dis(Str1,Str2) means taking and The minimum value of , that is, the edit distance between Str1 and Str2.

[0029] Preferably, the improved K-means algorithm calculates the edit distance between data objects in the same cluster: if the edit distance between data objects is less than the set empirical value η, then one of the data objects is evaluated as redundant data; if the edit distance between data objects is greater than the set empirical value Φ, then one of the data objects is evaluated as abnormal data; if there is a missing attribute or missing identifier in the data object, then the data object is evaluated as missing data;

[0030] Redundant data is deleted, and fields with abnormal attributes in abnormal data are deleted and converted into missing data, thereby determining multi-source heterogeneous data with missing data.

[0031] Preferably, in step S3, a generative adversarial network model is constructed and trained by the following steps:

[0032] S301: Select a multi-layer perceptron neural network to construct a generative model and a discriminant model of a generative adversarial network model, and initialize model parameters of the generative model and the discriminant model;

[0033] The discriminant model consists of m discriminant units corresponding to attribute features, where m represents the number of attribute features. Each discriminant unit takes an attribute feature as input and outputs the missing probability of the corresponding attribute feature.

[0034] The generation model consists of m generation units corresponding to attribute features. Each generation unit takes the attribute feature field to be generated as output and other attribute feature fields related to the output as input.

[0035] S302: Construct a real data training set for the generative model, train the generative model to simulate the mapping relationship between various attribute features of the real data, and train the discriminative model to learn the mapping relationship between the data and the probability of missing data;

[0036] S303: Generate a missing data filling matrix using the trained generative model, and determine the data missing probability of the missing data filling matrix using the trained discriminative model;

[0037] S304: Determine whether the generated data result of the generative model and the discriminant model have reached Nash equilibrium: if not, update and iterate the parameters of the generative model and the discriminant model; otherwise, terminate the training;

[0038] S305: Evaluate the data cleaning performance of the generative adversarial network model.

[0039] Preferably, the cost function of the constructed generative adversarial network model is expressed by the following formula:

[0040] 1) Cost function of the discriminant model:

[0041]

[0042] Where: D and G represent the discriminant model function and the generative model function respectively; J (D) (D, G) represents the cost function of the discriminant model; E represents expectation; D and G represent the discriminant model function and the generative model function respectively; z represents noise; Indicates that the discriminant model determines that the input x is a real data sample; Indicates that the discriminant model determines that the data sample is generated;

[0043] 2) Cost function of the generative model:

[0044] J (G) (D,G)=-J (D) (D,G);

[0045] Where: J (G) (D,G) represents the cost function of the generative model;

[0046] The maximum and minimum value functions of the constructed generative adversarial network model are expressed by the following formula:

[0047]

[0048] Where: P data (x) represents the real data distribution, P z (z) represents the generated data distribution.

[0049] Preferably, the loss function of the constructed generative adversarial network model is expressed by the following formula:

[0050] 1) Loss function of the generative model:

[0051]

[0052] Where: Γ G Represents the loss function of the generative model; for normal heterogeneous data objects without missing data, that is, real data: a i Indicates the value of attribute feature field i, Represents other attribute feature fields; n represents the number of training samples; δ represents the regularization parameter; Indicates the learning weights of generative model training; Represents the attribute feature data a calculated by the generative model i The missing probability of m is the number of attribute features;

[0053] 2) Loss function of the discriminant model:

[0054]

[0055] Where: Γ D Represents the loss function of the discriminant model; a i Represents the attribute characteristics of the discriminant model input; D(a i )∈(0,1) represents the missing probability of the corresponding attribute feature output by the discriminant model; Represents the discriminant model training learning weight; M i is the corresponding data value of the missing matrix M; D i (a i ) represents the attribute feature data a calculated by the discriminant model i The missing probability of M ijIndicates the value of the i-th row and j-th column of the missing matrix M;

[0056] The iterative formula of the constructed generative adversarial network model is expressed by the following formula;

[0057] 1) Iterative formula of the discriminant model:

[0058]

[0059] Where: Γ D represents the loss function of the discriminant model; α is the iteration step size, that is, the learning rate; D i represents the discriminant model matrix; Denotes derivation (optional); D i (a (j) ) represents the output of the discriminant model; D i : represents the matrix after the discriminant model iteration;

[0060] 2) Iterative formula for generating the model:

[0061]

[0062] Where: Γ G Represents the loss function of the generative model; α is the iteration step size, i.e., the learning rate; G i represents the generative model matrix; G i Represents the matrix after the generation model iteration; Indicates other attribute feature fields calculated by the generative model The missing probability of a i (j) Indicates the value of attribute feature field i at the jth iteration;

[0063] For the generative model and the discriminative model: when the missing data of the new input is filled in the matrix A new And the missing data input in the previous iteration fills the matrix A old When the difference starts to increase, it means that the output result has converged and the iteration process stops;

[0064] The data filling matrix A is calculated by the following formula new and A old The difference:

[0065]

[0066] Where: Δ represents the missing data filling matrix A new and A old The difference.

[0067] Preferably, in step 305, the data cleaning performance of the generative adversarial network model is evaluated by the following formula:

[0068]

[0069] Where: RMSE represents the root mean square error; items represents the number of missing attribute features of heterogeneous data objects; A miss Represents the attribute feature matrix containing missing data; A new Represents the missing data filling matrix of the new input.

[0070] Compared with the existing technology, the multi-source heterogeneous data cleaning method based on generative adversarial network in this invention has the following beneficial effects:

[0071] The present invention synthesizes each multi-source heterogeneous data into a multi-source heterogeneous data fusion table, and analyzes the redundant data, abnormal data and missing data remaining in the multi-source heterogeneous data fusion table through a cluster analysis algorithm, thereby determining the multi-source heterogeneous data with missing data, so that the multi-source heterogeneous data with missing data can be accurately and effectively divided, thereby improving the accuracy of multi-source heterogeneous data cleaning. At the same time, the present invention generates a missing data filling matrix for multi-source heterogeneous data with missing data through a trained generative adversarial network model, and then can fill the multi-source heterogeneous data with missing data through the missing data filling matrix to achieve the fusion and cleaning of multi-source heterogeneous data, thereby improving the effectiveness of multi-source heterogeneous data cleaning, and providing a relatively practical solution for ensuring the data quality of multi-source heterogeneous data of intelligent production lines.

[0072] In the present invention, an improved K-means algorithm is used as a clustering analysis algorithm, wherein the improved K-means algorithm uses the edit distance as a measure of the similarity between data. The edit distance has good robustness and compatibility in processing string similarity, making it possible to more accurately and effectively divide multi-source heterogeneous data with missing data, thereby further improving the accuracy of multi-source heterogeneous data cleaning. At the same time, the traditional K-means clustering algorithm requires the number of clusters to be defined in advance, and different numbers of clusters affect the clustering results. The improved K-means clustering algorithm of the present invention uses the characteristic that the data with the greatest distance do not belong to the same category to automatically determine the cluster center and the number of clusters through the maximum distance between the data, thereby further improving the clustering effect of the data.

[0073] The present invention generates a missing data filling matrix through a generative adversarial network. Compared with traditional explicit density generator models such as maximum likelihood estimation method, Markov chain method, and approximation method, the generative adversarial network is an implicit density generator model of adversarial training. The model does not explicitly give the distribution density function of the data sample, which reduces the complex calculation of fitting the distribution of real data samples and the difficulty of expanding to high dimensions. The applicability is wider than that of traditional models. In addition, due to the powerful data fitting performance of the neural network, the generative adversarial network has more obvious advantages in data generation problems, and can better generate missing data filling matrices, and thus can better realize the fusion and cleaning of multi-source heterogeneous data, thereby further improving the effectiveness of multi-source heterogeneous data cleaning. BRIEF DESCRIPTION OF THE DRAWINGS

[0074] In order to make the purpose, technical solutions and advantages of the invention more clear, the present invention will be further described in detail below with reference to the accompanying drawings, in which:

[0075] Figure 1 This is a logical block diagram of a multi-source heterogeneous data cleaning method based on a generative adversarial network.

[0076] Figure 2 Flowchart for training a generative adversarial network;

[0077] Figure 3 (a), (b), and (c) are schematic diagrams of data A, data B, and data C in Table 1, respectively;

[0078] Figure 4 The following are the classification results of four groups of normal data samples;

[0079] Figure 5 This is the classification result diagram of missing data samples;

[0080] Figure 6 This is the result graph of 200 iterations of the generative adversarial network model;

[0081] Figure 7 This is the result of 2400 iterations of the generative adversarial network model;

[0082] Figure 8 Filling effect diagram for missing data sample set;

[0083] Figure 9 To compare the generated data with the real data;

[0084] Figure 10 This is a comparison of the filling error rates of different algorithms under different sample missing rates (the sample size is 5000);

[0085] Figure 11 The figure shows the comparison of filling error rates of different algorithms under different sample capacities (sample missing rate is 0.2). DETAILED DESCRIPTION

[0086] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. The components of the embodiments of the present invention generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the invention claimed for protection, but only represents selected embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0087] It should be noted that similar reference numerals and letters denote similar items in the following figures. Therefore, once an item is defined in one figure, it does not require further definition or explanation in subsequent figures. In the description of the present invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer" indicate positions or relationships based on the positions or relationships shown in the figures, or the positions or relationships in which the inventive product is typically placed when in use. These terms are intended solely to facilitate the description of the present invention and simplify the description. They do not indicate or imply that the devices or components referred to must have a specific orientation, be constructed, or operate in a specific orientation, and are therefore not to be construed as limiting the present invention. Furthermore, the terms "first," "second," and "third," etc., are used solely to distinguish descriptions and are not to be construed as indicating or implying relative importance. Furthermore, terms such as "horizontal" and "vertical" do not imply that a component is absolutely horizontal or overhanging, but rather may be slightly tilted. For example, "horizontal" simply means that its direction is more horizontal than "vertical," and does not mean that the structure must be completely horizontal, but rather may be slightly tilted. In the description of the present invention, it should also be noted that, unless otherwise expressly specified or limited, the terms "disposed," "installed," "connected," and "connected" should be understood in a broad sense. For example, they may refer to fixed connections, detachable connections, or integral connections; they may refer to mechanical connections or electrical connections; they may refer to direct connections or indirect connections through an intermediate medium; and they may refer to internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on the specific circumstances.

[0088] The following is a further detailed description through specific implementation methods:

[0089] Example:

[0090] This embodiment discloses a multi-source heterogeneous data cleaning method based on a generative adversarial network.

[0091] like Figure 1 As shown in Figure 2, the multi-source heterogeneous data cleaning method based on the generative adversarial network includes:

[0092] S1: Obtain multi-source heterogeneous data of the intelligent production line and synthesize each multi-source heterogeneous data into a corresponding multi-source heterogeneous data fusion table;

[0093] S2: Analyze the redundant data, abnormal data, and missing data remaining in the multi-source heterogeneous data fusion table through cluster analysis algorithms, and then determine the multi-source heterogeneous data with missing data;

[0094] S3: Input the multi-source heterogeneous data with missing data into the trained generative adversarial network model and output the corresponding missing data filling matrix;

[0095] S4: Fill in the missing data with the missing data matrix to achieve the fusion and cleaning of multi-source heterogeneous data.

[0096] The present invention synthesizes each multi-source heterogeneous data into a multi-source heterogeneous data fusion table, and analyzes the redundant data, abnormal data and missing data remaining in the multi-source heterogeneous data fusion table through a cluster analysis algorithm, thereby determining the multi-source heterogeneous data with missing data, so that the multi-source heterogeneous data with missing data can be accurately and effectively divided, thereby improving the accuracy of multi-source heterogeneous data cleaning. At the same time, the present invention generates a missing data filling matrix for multi-source heterogeneous data with missing data through a trained generative adversarial network model, and then can fill the multi-source heterogeneous data with missing data through the missing data filling matrix to achieve the fusion and cleaning of multi-source heterogeneous data, thereby improving the effectiveness of multi-source heterogeneous data cleaning, and providing a relatively practical solution for ensuring the data quality of multi-source heterogeneous data of intelligent production lines.

[0097] During the specific implementation process, the data tables corresponding to each multi-source heterogeneous data are imported into the data warehouse to synthesize the corresponding multi-source heterogeneous data fusion table; during the synthesis process, abnormal data with association rule errors are detected and deleted.

[0098] The present invention can reduce the possibility of local optimal solutions in subsequent cluster analysis by detecting and deleting abnormal data with association rule errors during the synthesis process and initializing the classification of multi-source heterogeneous data of the gearbox intelligent production line, thereby improving the effectiveness of multi-source heterogeneous data cleaning.

[0099] During implementation, an improved K-means algorithm was selected as the cluster analysis algorithm. This algorithm uses edit distance as a measure of similarity between data points, leveraging the fact that the furthest-distance data points do not belong to the same category. This algorithm automatically determines the cluster center and number of clusters based on the maximum distance between data points. The edit distance, also known as the Levenshtein distance, is the minimum number of character transformations required to transform a source string into a target string through deletion, substitution, and addition. The smaller the edit distance between two data points, the greater their similarity.

[0100] The working logic of the improved K-means algorithm is as follows:

[0101] S201: Convert the record set of the multi-source heterogeneous data fusion table into the corresponding string set A = {a1, a2, ..., a n}, and calculate the edit distance between each pair of strings in the string set to generate the corresponding edit distance result set G; the multi-source heterogeneous data fusion table is mostly of string types, and converting it into a string set facilitates the calculation of the edit distance (the edit distance has good robustness and compatibility in processing string similarity).

[0102] S202: Select the two data objects with the largest distance in the edit distance result set G. and As the cluster centers of the initial two clusters S1 and S2, and The distance between them is recorded as d1, that is

[0103] Divide the string set A by and Data objects other than and To classify the cluster centers, if a i ∈A and a i and Edit distance Less than Then a i Divide into S1, otherwise a i Divided into S2;

[0104] S203: Get all data objects in S1 to The maximum edit distance d 11 ,Right now Get all data objects in S2 to The maximum edit distance d 22 ,Right now

[0105] Take d 11 and d22 The maximum value d2, record d2 = max{d 11 ,d 22};

[0106] S204: If d2>k*d1, then take the corresponding data object as the cluster center of the third cluster S3

[0107] In this embodiment, k is an empirical parameter, 0≤k≤1, determined according to the change trend of the cluster center of the data object, and generally k is 0.5.

[0108] by and As the cluster center, clustering is performed to form three clusters S1, S2 and S3. That is, if a i ∈A and Then a i Divided into S1, if a i ∈A and Then a i Divide into S2, otherwise a i Divided into S3;

[0109] S205: Get all data objects in S1 to The maximum edit distance d 11 ,Right now Get all data objects in S2 to The maximum edit distance d 22 ,Right now Get all data objects in S3 to The maximum edit distance d 33 ,Right now

[0110] Take d 11 d 22 and d 33 The maximum value d3, record d3 = max{d 11 ,d 22 ,d 33};

[0111] S206: If d3>k*(d2+d1)÷2, then take the corresponding data object as the cluster center of the fourth cluster S4 by and Clustering is performed as the cluster center to form four clusters S1, S2, S3 and S4;

[0112] S207: Repeat steps S204 to S206 until the clustering condition is no longer met, and output the corresponding data clustering result.

[0113] In the specific implementation process, in step S201, the edit distance between two character strings is calculated using the following formula:

[0114]

[0115] Where: Str1 and Str2 represent two character strings, L1 represents the length of Str1, and L2 represents the length of Str2; Str1 * Indicates the result after randomly deleting a character from Str1; Indicates the result after deleting a character randomly from Str2; Dis(Str1 * ,Str2)+1 means that the last operation of transforming Str1 into Str2 is the edit distance of inserting one character into Str1; Indicates that the last operation in converting Str1 to Str2 is the edit distance of deleting a character from Str2; Dis(Str1 * ,Str2 * )+1 means that the last operation in converting Str1 to Str2 is the edit distance of replacing one character in Str1; Dis(Str1,Str2) means taking Dis(Str1 * ,Str2)+1, and The minimum value of , that is, the edit distance between Str1 and Str2.

[0116] The improved K-means algorithm calculates the edit distance between data objects in the same cluster: if the edit distance between data objects is less than the set empirical value η, then one of the data objects is evaluated as redundant data; if the edit distance between data objects is greater than the set empirical value Φ, then one of the data objects is evaluated as abnormal data; if there is a missing attribute or missing identifier in the data object, then the data object is evaluated as missing data;

[0117] In this embodiment, η and Φ are empirical values determined according to the degree of discreteness of the data set.

[0118] Redundant data is deleted, and fields with abnormal attributes in abnormal data are deleted and converted into missing data, thereby determining multi-source heterogeneous data with missing data.

[0119] In the present invention, an improved K-means algorithm is used as a clustering analysis algorithm, wherein the improved K-means algorithm uses the edit distance as a measure of the similarity between data. The edit distance has good robustness and compatibility in processing string similarity, making it possible to more accurately and effectively divide multi-source heterogeneous data with missing data, thereby further improving the accuracy of multi-source heterogeneous data cleaning. At the same time, the traditional K-means clustering algorithm requires the number of clusters to be defined in advance, and different numbers of clusters affect the clustering results. The improved K-means clustering algorithm of the present invention uses the characteristic that the data with the greatest distance do not belong to the same category to automatically determine the cluster center and the number of clusters through the maximum distance between the data, thereby further improving the clustering effect of the data.

[0120] During the specific implementation process, the generative adversarial network consists of a generative model and a discriminative model. The generative model analyzes the underlying distribution patterns of real data samples and generates new data samples. The discriminative model is a binary classifier that determines whether the input data sample is a real data sample or a generated data sample. If the discriminative model fails to make a judgment, the discriminative model parameters are optimized to improve the accuracy of the discriminative model; if the discriminative model succeeds, the generative model parameters are optimized to make the generated data samples more consistent with the real data samples. The cyclic iterative optimization process of the generative model and the discriminative model is actually a maximum-minimum game problem. The final game result of the generative model and the discriminative model is that the generative model and the discriminative model reach a non-cooperative game equilibrium, that is, the execution strategy of each model aims to achieve the maximum expected benefit.

[0121] Combine Figure 2 As shown in Figure 2, the following steps are used to build and train the generative adversarial network model:

[0122] S301: Select a multi-layer perceptron neural network to construct a generative model and a discriminant model of a generative adversarial network model, and initialize model parameters of the generative model and the discriminant model;

[0123] The discriminant model consists of m discriminant units corresponding to attribute features, where m represents the number of attribute features. Each discriminant unit takes an attribute feature as input and outputs the missing probability of the corresponding attribute feature.

[0124] The generation model consists of m generation units corresponding to attribute features. Each generation unit takes the attribute feature field to be generated as output and other attribute feature fields related to the output as input.

[0125] S302: Construct a real data training set for the generative model, train the generative model to simulate the mapping relationship between various attribute features of the real data, and train the discriminative model to learn the mapping relationship between the data and the probability of missing data;

[0126] S303: Generate a missing data filling matrix using the trained generative model, and determine the data missing probability of the missing data filling matrix using the trained discriminative model;

[0127] S304: Determine whether the generated data result of the generative model and the discriminant model have reached Nash equilibrium: if not, update and iterate the parameters of the generative model and the discriminant model; otherwise, terminate the training;

[0128] In this embodiment, the Nash equilibrium is an existing concept, that is, a strategy combination that satisfies the following property: any player will not increase his or her own profit if he or she unilaterally changes his or her strategy under this strategy combination (while other players' strategies remain unchanged).

[0129] S305: Evaluate the data cleaning performance of the generative adversarial network model.

[0130] In a generative adversarial network, the generative model and the discriminative model can be any differentiable function. The iterative confrontation between the generative model and the discriminative model is a minimax game problem. The generative model and the discriminative model are closely related, and the game relationship between the two is a zero-sum game, that is, the combined cost function of the two is zero. The cost function of the constructed generative adversarial network model is expressed by the following formula:

[0131] 1) Cost function of the discriminant model:

[0132]

[0133] Where: D and G represent the discriminant model function and the generative model function respectively; J (D) (D, G) represents the cost function of the discriminant model; E represents expectation; D and G represent the discriminant model function and the generative model function respectively; z represents noise; Indicates that the discriminant model determines that the input x is a real data sample; Indicates that the discriminant model determines that the data sample is generated;

[0134] 2) Cost function of the generative model:

[0135] J (G) (D,G)=-J (D) (D,G);

[0136] Where: J (G) (D,G) represents the cost function of the generative model;

[0137] The value function for mini-batch gradient descent optimization of the discriminant model of the generative adversarial network:

[0138]

[0139] The value function optimized for the generative model of the Generative Adversarial Network is:

[0140]

[0141] The goal of optimizing the discriminant model is to maximize the results obtained from real data sample inputs and minimize the results obtained from generated data sample inputs. Therefore, we have 1-D(G(z)), and the result is maximized. Optimizing the generative model is to maximize the results D(G(z)) obtained from generated sample inputs, and minimize the results. The maximum and minimum value functions of the constructed generative adversarial network model are expressed as follows:

[0142]

[0143] Where: P data (x) represents the real data distribution, P z (z) represents the generated data distribution.

[0144] The discriminant model D consists of m discriminant units corresponding to attribute features, where m is the number of attribute features. Each discriminant unit uses attribute feature data a i As input, the output corresponding attribute feature missing probability D(a i )∈(0,1), D(a i ) is 1-M i , M i is the corresponding data value of the missing matrix M, and the loss function Γ D represents D(a i ) and 1-M i The generation model G consists of m generation units corresponding to attribute features. Each generation unit takes the attribute feature field to be generated as output and other attribute feature fields related to the output as input. The loss function Γ G Represents the difference between the generated attribute feature field and the true attribute feature field minus the loss of the discriminant model.

[0145] Assume that a heterogeneous data object A contains missing data, and the value of its missing attribute feature field i is Other attribute feature fields are For normal heterogeneous data objects in the same cluster that do not contain missing data, the value of the attribute feature field i is a i , other attribute feature fields are Generative Model G Simulation to a i The mapping is trained, and the generated model after training is based on generate Fill in the missing data. The loss function of the constructed generative adversarial network model is expressed by the following formula:

[0146] 1) Loss function of the generative model:

[0147]

[0148] Where: Γ G Represents the loss function of the generative model; for normal heterogeneous data objects without missing data, that is, real data: a i Indicates the value of attribute feature field i, Represents other attribute feature fields; n represents the number of training samples; δ represents the regularization parameter; Indicates the learning weights of generative model training; Represents the attribute feature data a calculated by the generative model i The missing probability of m is the number of attribute features;

[0149] 2) Loss function of the discriminant model:

[0150]

[0151] Where: Γ D Represents the loss function of the discriminant model; a i Represents the attribute characteristics of the discriminant model input; D(a i )∈(0,1) represents the missing probability of the corresponding attribute feature output by the discriminant model; Represents the discriminant model training learning weight; M i is the corresponding data value of the missing matrix M; D i (a i ) represents the attribute feature data a calculated by the discriminant model i The probability of missing Indicates the value of the i-th row and j-th column of the missing matrix M;

[0152] The adversarial game between the generative model and the discriminative model is a cyclical, iterative learning and optimization process, where the optimization method directly determines the quality of the optimization results. Traditional iterative learning optimization methods, such as the least squares method, are computationally difficult and produce unsatisfactory optimization results when faced with large sample sizes and complex models, such as those found in multi-source heterogeneous data from intelligent production lines. Therefore, the present invention adopts a small-batch gradient descent optimization method, using a small batch of random sample instances for optimization calculations in each iteration. This method offers advantages in computational time and convergence over batch gradient descent optimization methods and stochastic gradient descent methods.

[0153] A small sample with a sample size of b is selected for gradient descent optimization. The iterative formula of the constructed generative adversarial network model is expressed by the following formula;

[0154] 1) Iterative formula of the discriminant model:

[0155]

[0156] Where: Γ D represents the loss function of the discriminant model; α is the iteration step size, that is, the learning rate; D i represents the discriminant model matrix; Denotes derivation (optional); D i (a (j) ) represents the output of the discriminant model; D i : represents the matrix after the discriminant model iteration;

[0157] 2) Iterative formula for generating the model:

[0158]

[0159] Where: Γ G Represents the loss function of the generative model; α is the iteration step size, i.e., the learning rate; G i represents the generative model matrix; G i Represents the matrix after the generation model iteration; Indicates other attribute feature fields calculated by the generative model The missing probability of a i (j) Indicates the value of attribute feature field i at the jth iteration;

[0160] The generation effect of the generative model G is determined by the discriminative model D. Therefore, the iterative effect of the generative model G needs to consider the discriminative effect of the discriminative model D, that is, the loss function of the discriminative model D is added to the iterative formula of the generative model G.

[0161] For the generative model and the discriminative model: when the missing data of the new input is filled in the matrix A new And the missing data input in the previous iteration fills the matrix A old When the difference starts to increase, it means that the output result has converged and the iteration process stops;

[0162] The data filling matrix A is calculated by the following formula new and A old The difference:

[0163]

[0164] Where: Δ represents the missing data filling matrix A new and A old The difference.

[0165] In the specific implementation process, in step 305, the data cleaning performance of the generative adversarial network model is evaluated by the following formula:

[0166]

[0167] Where: RMSE represents the root mean square error; item s represents the number of missing attribute features of heterogeneous data objects; A miss Represents the attribute feature matrix containing missing data; A new Represents the missing data filling matrix of the new input.

[0168] The present invention generates a missing data filling matrix through a generative adversarial network. Compared with traditional explicit density generator models such as maximum likelihood estimation method, Markov chain method, and approximation method, the generative adversarial network is an implicit density generator model of adversarial training. The model does not explicitly give the distribution density function of the data sample, which reduces the complex calculation of fitting the distribution of real data samples and the difficulty of expanding to high dimensions. The applicability is wider than that of traditional models. In addition, due to the powerful data fitting performance of the neural network, the generative adversarial network has more obvious advantages in data generation problems, and can better generate missing data filling matrices, and thus can better realize the fusion and cleaning of multi-source heterogeneous data, thereby further improving the effectiveness of multi-source heterogeneous data cleaning.

[0169] At the same time, the present invention uses the root mean square error as a parameter indicator to evaluate the data cleaning performance of the generative adversarial network model, so that the data cleaning performance of the generative adversarial network model can be effectively evaluated, thereby further improving the effectiveness of multi-source heterogeneous data cleaning.

[0170] In order to better illustrate the advantages of the technical solution of the present invention, the following experiments are disclosed in this embodiment.

[0171] 1. Experimental Preparation

[0172] In order to verify the effectiveness and feasibility of the multi-source heterogeneous data cleaning method proposed in this invention, this experiment uses a multi-source heterogeneous data set of a gearbox box intelligent production line for data cleaning verification, and selects the common data cleaning methods singular value decomposition algorithm (SVD) and nearest neighbor node method (KNN) as comparison.

[0173] This experiment was conducted on a computer with an Intel i5 processor, 16G memory, and Windows 10 operating system. The basic information of the dataset is shown in Table 1.

[0174] Table 1 Basic information of a multi-source heterogeneous dataset of a gearbox box intelligent production line

[0175]

[0176] 2. Experimental Process

[0177] This experiment constructs an attribute feature matrix based on the data of a multi-source heterogeneous data set of a certain gearbox case intelligent production line, and initializes the generation model and the discrimination model as a multi-layer perception model. Experimental data samples are randomly selected from 9000 groups of data sets, and the spindle vibration attribute characteristics of 200 groups of data samples are randomly selected from the selected 9000 groups of experimental data samples to perform data missing and data anomaly processing. After processing, the experimental data samples containing missing data and abnormal data are evaluated, outlier processed and classified by the multi-source heterogeneous data cleaning evaluation model of the gearbox case intelligent production line constructed by the present invention, and clustered to form four groups of normal data distributions and two groups of data distributions containing missing data. The clustering results are as follows: Figure 4 and Figure 5 shown.

[0178] In this experiment, four groups of normal data samples are used as training sample data for the generative model, and the main shaft vibration attribute characteristics are used as a i , side wear, processing materials and other attribute feature fields as Training a generative model to learn fitting to a i The mapping relationship between sample data and missing probability is trained to train the discriminant model. 150 groups of data samples are randomly selected from the two groups of data samples with missing data each time, and the generative model and discriminant model of small batch data samples are iteratively gamed. Repeated selection is performed to avoid the randomness of the iterative game. The iterative results are as follows: Figure 6 and Figure 7 shown.

[0179] The stopping condition of the iterative game is that if the difference between the generated data of the generative model and the data generated in the previous iterative game increases, the iterative game between the generative model and the discriminant model is stopped. The generative model inputs a data sample set containing missing data, and the generative model fills in the missing data in the data sample set. The filling effect is as follows: Figure 8 As shown, the comparison between the filled data and the real data is Figure 9 As shown in Figure 3, the root mean square error (RMSE) between the generated data and the real data is 3.75388.

[0180] The experimental results demonstrate that the multi-source heterogeneous data cleaning and evaluation model for intelligent transmission case production lines, constructed based on the method developed in this invention, demonstrates feasibility in evaluating, classifying, analyzing, and processing heterogeneous data from processing categories such as intelligent transmission case production lines. This means that the multi-source heterogeneous data cleaning method based on a generative adversarial network (GAN) in this invention demonstrates a certain degree of reliability and accuracy in filling gaps and correcting anomalies in multi-source heterogeneous data from processing categories such as intelligent transmission case production lines.

[0181] 3. Algorithm Comparative Analysis

[0182] This experiment selects the common nearest neighbor method (KNN) and singular value decomposition algorithm (SVD) as comparison algorithms for the multi-source heterogeneous data cleaning method based on generative adversarial networks. The KNN algorithm analyzes the k data feature attributes that are most correlated with the missing attributes, and cyclically merges and calculates the filling value of the missing attributes; the SVD algorithm converges through multiple iterations, calculates the approximate low-rank matrix of the data feature attribute matrix, and replaces the missing attribute values. This experiment compares the error rate of data filling between different algorithms from the two dimensions of sample missing rate and sample capacity. The comparison results are shown in the figure below. Figure 10 and Figure 11 shown.

[0183] From the comparison results, it can be seen that the multi-source heterogeneous data cleaning method for the intelligent production line of the gearbox case proposed based on the method of the present invention has certain advantages in the error rate of filling values compared with the common KNN data cleaning algorithm and SVD data cleaning algorithm under different sample capacities and different sample missing values.

[0184] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit the technical solutions. Those skilled in the art should understand that modifications or equivalent replacements of the technical solutions of the present invention that do not depart from the purpose and scope of the technical solutions of the present invention should be included in the scope of the claims of the present invention.

Claims

1. A multi-source heterogeneous data cleaning method based on generative adversarial networks, characterized by: include: S1: Obtain multi-source heterogeneous data from the gearbox case intelligent production line and synthesize each multi-source heterogeneous data into a corresponding multi-source heterogeneous data fusion table; S2: Analyze the redundant data, abnormal data, and missing data remaining in the multi-source heterogeneous data fusion table through cluster analysis algorithms, and then determine the multi-source heterogeneous data with missing data; S3: Input the multi-source heterogeneous data with missing data into the trained generative adversarial network model and output the corresponding missing data filling matrix; In step S3, the generative adversarial network model is constructed and trained through the following steps: S301: Select a multi-layer perceptron neural network to construct a generative model and a discriminant model of a generative adversarial network model, and initialize model parameters of the generative model and the discriminant model; The discriminant model consists of m discriminant units corresponding to attribute features, where m represents the number of attribute features. Each discriminant unit takes an attribute feature as input and outputs the missing probability of the corresponding attribute feature. The generation model consists of m generation units corresponding to attribute features. Each generation unit takes the attribute feature field to be generated as output and other attribute feature fields related to the output as input. S302: Construct a real data training set for the generative model, train the generative model to simulate the mapping relationship between various attribute features of the real data, and train the discriminative model to learn the mapping relationship between the data and the probability of missing data; S303: Generate a missing data filling matrix using the trained generative model, and determine the data missing probability of the missing data filling matrix using the trained discriminative model; S304: Determine whether the generated data result of the generative model and the discriminant model have reached a Nash equilibrium: if not, update and iterate the parameters of the generative model and the discriminant model; Otherwise, the training ends; S305: Evaluate the data cleaning performance of the generative adversarial network model; S4: Fill missing data in multi-source heterogeneous data through the missing data filling matrix to achieve the fusion and cleaning of multi-source heterogeneous data; The experiment constructs an attribute feature matrix based on the data of a multi-source heterogeneous dataset of the gearbox case intelligent production line, and initializes the generative model and discriminant model as a multi-layer perception model. 9,000 sets of experimental data samples are randomly selected from the dataset, and 200 sets of spindle vibration attribute features are randomly selected from the 9,000 sets of experimental data samples to process missing data and data anomaly. The attribute characteristics include flank wear, experimental time, cutting depth, processing material, AC spindle motor current, DC spindle motor current, chuck vibration, spindle vibration, chuck acoustic emission, and spindle acoustic emission.

2. The multi-source heterogeneous data cleaning method based on a generative adversarial network according to claim 1, characterized in that: In step S1, the data tables corresponding to the various multi-source heterogeneous data are imported into the data warehouse to synthesize the corresponding multi-source heterogeneous data fusion table; during the synthesis process, abnormal data with association rule errors are detected and deleted.

3. The multi-source heterogeneous data cleaning method based on generative adversarial network according to claim 1, characterized in that: In step S2, the improved K-means algorithm is selected as the cluster analysis algorithm; The improved K-means algorithm uses the edit distance as a measure of similarity between data, and then uses the characteristic that the farthest data do not belong to the same category to automatically determine the cluster center and the number of clusters through the maximum distance between data.

4. The multi-source heterogeneous data cleaning method based on generative adversarial network according to claim 3 is characterized in that: In step S2, the working logic of the improved K-means algorithm is as follows: S201: Convert the record set of the multi-source heterogeneous data fusion table into the corresponding string set A = {a1, a2, ..., a n }, and calculate the edit distance between each two strings in the string set, and generate the corresponding edit distance result set G; S202: Select the two data objects with the largest distance in the edit distance result set G. and As the cluster centers of the initial two clusters S1 and S2, and The distance between them is recorded as d1, that is Divide the string set A by and Data objects other than and To classify the cluster centers, if a i ∈A and a i and Edit distance Less than Then a i Divide into S1, otherwise a i Divided into S2; S203: Get all data objects in S1 to The maximum edit distance d 11 ,Right now Get all data objects in S2 to The maximum edit distance d 22 ,Right now Take d 11 and d 22 The maximum value d2, record d2 = max{d 11 ,d 22 }; S204: If d2>k*d1, then take the corresponding data object as the cluster center of the third cluster S3 by and As the cluster center, clustering is performed to form three clusters S1, S2 and S3. That is, if a i ∈A and Then a i Divided into S1, if a i ∈A and Then a i Divide into S2, otherwise a i Divided into S3; S205: Get all data objects in S1 to The maximum edit distance d 11 ,Right now Get all data objects in S2 to The maximum edit distance d 22 ,Right now Get all data objects in S3 to The maximum edit distance d 33 ,Right now Take d 11 d 22 and d 33 The maximum value d3, record d3 = max{d 11 ,d 22 ,d 33 }; S206: If d3>k*(d2+d1)÷2, then take the corresponding data object as the cluster center of the fourth cluster S4 by and Clustering is performed as the cluster center to form four clusters S1, S2, S3 and S4; S207: Repeat steps S204 to S206 until the clustering condition is no longer met, and output the corresponding data clustering result.

5. The multi-source heterogeneous data cleaning method based on generative adversarial network according to claim 4 is characterized in that: In step S201, the edit distance between two character strings is calculated using the following formula: Where: Str1 and Str2 represent two character strings, L1 represents the length of Str1, and L2 represents the length of Str2; Str1 * Indicates the result after randomly deleting a character from Str1; Indicates the result after deleting a character randomly from Str2; Dis(Str1 * ,Str2)+1 means that the last operation of transforming Str1 into Str2 is the edit distance of inserting one character into Str1; Indicates that the last operation in converting Str1 to Str2 is the edit distance of deleting one character from Str2; Indicates that the last operation in converting Str1 to Str2 is the edit distance of replacing one character in Str1; Dis(Str1,Str2) means taking Dis(Str1 * ,Str2)+1, and The minimum value of , that is, the edit distance between Str1 and Str2.

6. The multi-source heterogeneous data cleaning method based on generative adversarial network according to claim 4, characterized in that: The improved K-means algorithm calculates the edit distance between data objects in the same cluster: if the edit distance between data objects is less than the set empirical value η, then one of the data objects is evaluated as redundant data; if the edit distance between data objects is greater than the set empirical value Φ, then one of the data objects is evaluated as abnormal data; if there is a missing attribute or missing identifier in the data object, then the data object is evaluated as missing data; Redundant data is deleted, and fields with abnormal attributes in abnormal data are deleted and converted into missing data, thereby determining multi-source heterogeneous data with missing data.

7. The multi-source heterogeneous data cleaning method based on generative adversarial network according to claim 1, characterized in that: The cost function of the constructed generative adversarial network model is expressed by the following formula; 1) Cost function of the discriminant model: Where: D and G represent the discriminant model function and the generative model function respectively; J (D) (D, G) represents the cost function of the discriminant model; E represents expectation; D and G represent the discriminant model function and the generative model function respectively; z represents noise; Indicates that the discriminant model determines that the input x is a real data sample; Indicates that the discriminant model determines that the data sample is generated; 2) Cost function of the generative model: J (G) (D,G)=-J (D) (D,G); Where: J (G) (D,G) represents the cost function of the generative model; The maximum and minimum value functions of the constructed generative adversarial network model are expressed by the following formula: Where: P data (x) represents the real data distribution, P z (z) represents the generated data distribution.

8. The multi-source heterogeneous data cleaning method based on generative adversarial network according to claim 7, characterized in that: The loss function of the constructed generative adversarial network model is expressed by the following formula; 1) Loss function of the generative model: Where: Γ G Represents the loss function of the generative model; for normal heterogeneous data objects without missing data, that is, real data: a i Indicates the value of attribute feature field i, Represents other attribute feature fields; n represents the number of training samples; δ represents the regularization parameter; Indicates the learning weights of generative model training; Represents the attribute feature data a calculated by the generative model i The missing probability of m is the number of attribute features; 2) Loss function of the discriminant model: Where: Γ D Represents the loss function of the discriminant model; a i Represents the attribute characteristics of the discriminant model input; D(a i )∈(0,1) represents the missing probability of the corresponding attribute feature output by the discriminant model; Represents the discriminant model training learning weight; M i is the corresponding data value of the missing matrix M; D i (a i ) represents the attribute feature data a calculated by the discriminant model i The missing probability of M ij Indicates the value of the i-th row and j-th column of the missing matrix M; The iterative formula of the constructed generative adversarial network model is expressed by the following formula; 1) Iterative formula of the discriminant model: Where: Γ D represents the loss function of the discriminant model; α is the iteration step size, that is, the learning rate; D i represents the discriminant model matrix; Denotes derivative; D i (a (j) ) represents the output of the discriminant model; D i : represents the matrix after the discriminant model iteration; 2) Iterative formula for generating the model: Where: Γ G Represents the loss function of the generative model; α is the iteration step size, i.e., the learning rate; G i represents the generative model matrix; G i Represents the matrix after the generation model iteration; Indicates other attribute feature fields calculated by the generative model The missing probability of a i (j) Indicates the value of attribute feature field i at the jth iteration; For the generative model and the discriminative model: when the missing data of the new input is filled in the matrix A new And the missing data input in the previous iteration fills the matrix A old When the difference starts to increase, it means that the output result has converged and the iteration process stops; The data filling matrix A is calculated by the following formula new and A old The difference: Where: Δ represents the missing data filling matrix A new and A old The difference.

9. The multi-source heterogeneous data cleaning method based on generative adversarial network according to claim 1, characterized in that: In step 305, the data cleaning performance of the generative adversarial network model is evaluated using the following formula: Where: RMSE represents the root mean square error; items represents the number of missing attribute features of heterogeneous data objects; A miss Represents the attribute feature matrix containing missing data; A new Represents the missing data filling matrix of the new input.

Citation Information

Patent Citations

  • Method convenient for cleaning, integrating and storing massive multi-source heterogeneous data

    CN112347093A